SOTAVerified

Red Teaming

Papers

Showing 76–100 of 251 papers

TitleStatusHype
Insights and Current Gaps in Open-Source LLM Vulnerability Scanners: A Comparative Analysis—0
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming—0
Computational Red Teaming in a Sudoku Solving Context: Neural Network Based Skill Representation and Acquisition—0
CELL your Model: Contrastive Explanations for Large Language Models—0
Arondight: Red Teaming Large Vision Language Models with Auto-generated Multi-modal Jailbreak Prompts—0
Investigating Bias Representations in Llama 2 Chat via Activation Steering—0
A Reward-driven Automated Webshell Malicious-code Generator for Red-teaming—0
Can Large Language Models Change User Preference Adversarially?—0
A Red Teaming Roadmap Towards System-Level Safety—0
GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models—0
Can Large Language Models Automatically Jailbreak GPT-4V?—0
Can Language Models be Instructed to Protect Personal Information?—0
A Red Teaming Framework for Securing AI in Maritime Autonomous Systems—0
Building Safe GenAI Applications: An End-to-End Overview of Red Teaming for Large Language Models—0
Breaking the Global North Stereotype: A Global South-centric Benchmark Dataset for Auditing and Mitigating Biases in Facial Recognition Systems—0
A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management—0
In-Context Experience Replay Facilitates Safety Red-Teaming of Text-to-Image Diffusion Models—0
IterAlign: Iterative Constitutional Alignment of Large Language Models—0
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols—0
LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs—0
FLIRT: Feedback Loop In-context Red Teaming—0
GhostPrompt: Jailbreaking Text-to-image Generative Models based on Dynamic Optimization—0
A Multi-Disciplinary Review of Knowledge Acquisition Methods: From Human to Autonomous Eliciting Agents—0
Hidden Ghost Hand: Unveiling Backdoor Vulnerabilities in MLLM-Powered Mobile GUI Agents—0
Finding Safety Neurons in Large Language Models—0
Show:102550
← PrevPage 4 of 11Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1SUDOAttack Success Rate41—Unverified