ShieldGemma: Generative AI Content Moderation Based on Gemma (2407.21772v2)

Published 31 Jul 2024 in cs.CL and cs.LG

Abstract: We present ShieldGemma, a comprehensive suite of LLM-based safety content moderation models built upon Gemma2. These models provide robust, state-of-the-art predictions of safety risks across key harm types (sexually explicit, dangerous content, harassment, hate speech) in both user input and LLM-generated output. By evaluating on both public and internal benchmarks, we demonstrate superior performance compared to existing models, such as Llama Guard (+10.8\% AU-PRC on public benchmarks) and WildCard (+4.3\%). Additionally, we present a novel LLM-based data curation pipeline, adaptable to a variety of safety-related tasks and beyond. We have shown strong generalization performance for model trained mainly on synthetic data. By releasing ShieldGemma, we provide a valuable resource to the research community, advancing LLM safety and enabling the creation of more effective content moderation solutions for developers.

References (36)

Citations (15)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Tweets

https://twitter.com/fly51fly/status/1819854718866465194

https://twitter.com/arxivsanitybot/status/1819002601569996946

https://twitter.com/javaeeeee1/status/1819146211573535231

https://twitter.com/osanpochuudayo/status/1820670930882089214

https://twitter.com/osanpochuudayo/status/1820672261399192051

YouTube

Show All Videos

ShieldGemma: Generative AI Content Moderation Based on Gemma (2407.21772v2)

Summary

Related Papers

Tweets

YouTube