SOTAVerified

LifeTox: Unveiling Implicit Toxicity in Life Advice

2023-11-16Code Available0· sign in to hype

Minbeom Kim, Jahyun Koo, Hwanhee Lee, Joonsuk Park, Hwaran Lee, Kyomin Jung

Code Available — Be the first to reproduce this paper.

Reproduce

Code

Abstract

As large language models become increasingly integrated into daily life, detecting implicit toxicity across diverse contexts is crucial. To this end, we introduce LifeTox, a dataset designed for identifying implicit toxicity within a broad range of advice-seeking scenarios. Unlike existing safety datasets, LifeTox comprises diverse contexts derived from personal experiences through open-ended questions. Experiments demonstrate that RoBERTa fine-tuned on LifeTox matches or surpasses the zero-shot performance of large language models in toxicity classification tasks. These results underscore the efficacy of LifeTox in addressing the complex challenges inherent in implicit toxicity. We open-sourced the dataset https://huggingface.co/datasets/mbkim/LifeTox and the LifeTox moderator family; 350M, 7B, and 13B.

Reproductions