SOTAVerified

Red Teaming

Papers

Showing 101–125 of 251 papers

TitleStatusHype
Offensive Security for AI Systems: Concepts, Practices, and Applications—0
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents—0
Safety by Measurement: A Systematic Literature Review of AI Safety Evaluation Methods—0
DMRL: Data- and Model-aware Reward Learning for Data Extraction—0
Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs—0
Red Teaming Large Language Models for Healthcare—0
When Testing AI Tests Us: Safeguarding Mental Health on the Digital Frontlines—0
SAGE: A Generic Framework for LLM Safety EvaluationCode0
Understanding and Mitigating Risks of Generative AI in Financial Services—0
RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models—0
ELAB: Extensive LLM Alignment Benchmark in Persian Language—0
X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents—0
The Structural Safety Generalization ProblemCode0
Multi-lingual Multi-turn Automated Red Teaming for LLMs—0
Strategize Globally, Adapt Locally: A Multi-Turn Red Teaming Agent with Dual-Level Learning—0
Red Teaming with Artificial Intelligence-Driven Cyberattacks: A Scoping Review—0
AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack Integration—0
MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models—0
Making Every Step Effective: Jailbreaking Large Vision-Language Models Through Hierarchical KV Equalization—0
A Framework for Evaluating Emerging Cyberattack Capabilities of AI—0
Red Teaming Contemporary AI Models: Insights from Spanish and Basque Perspectives—0
JBFuzz: Jailbreaking LLMs Efficiently and Effectively Using Fuzzing—0
MAD-MAX: Modular And Diverse Malicious Attack MiXtures for Automated LLM Red Teaming—0
Reinforced Diffuser for Red Teaming Large Vision-Language Models—0
Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges—0
Show:102550
← PrevPage 5 of 11Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1SUDOAttack Success Rate41—Unverified