SOTAVerified

Red Teaming

Papers

Showing 101–125 of 251 papers

TitleStatusHype
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red TeamingCode0
Steering Without Side Effects: Improving Post-Deployment Control of Language ModelsCode0
Red Teaming Language Models for Processing Contradictory DialoguesCode0
Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal ModelsCode0
Overriding Safety protections of Open-source ModelsCode0
Automated Progressive Red TeamingCode0
Look Before You Leap: Enhancing Attention and Vigilance Regarding Harmful Content with GuidelineLLMCode0
ASTPrompter: Weakly Supervised Automated Language Model Red-Teaming to Identify Low-Perplexity Toxic PromptsCode0
No Offense Taken: Eliciting Offensiveness from Language ModelsCode0
RabakBench: Scaling Human Annotations to Construct Localized Multilingual Safety Benchmarks for Low-Resource LanguagesCode0
Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges—0
CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring—0
JBFuzz: Jailbreaking LLMs Efficiently and Effectively Using Fuzzing—0
Conversational Complexity for Assessing Risk in Large Language Models—0
A Safe Harbor for AI Evaluation and Red Teaming—0
Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency—0
Jailbreaking Large Language Models with Symbolic Mathematics—0
Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters—0
Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts—0
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming—0
JAB: Joint Adversarial Prompting and Belief Augmentation—0
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs—0
Computational Red Teaming in a Sudoku Solving Context: Neural Network Based Skill Representation and Acquisition—0
IterAlign: Iterative Constitutional Alignment of Large Language Models—0
Investigating Bias Representations in Llama 2 Chat via Activation Steering—0
Show:102550
← PrevPage 5 of 11Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1SUDOAttack Success Rate41—Unverified