ShieldGemma: Generative AI Content Moderation Based on Gemma
Wenjun Zeng, Yuchi Liu, Ryan Mullins, Ludovic Peran, Joe, Fernandez, Hamza Harkous, Karthik Narasimhan, Drew Proud, Piyush, Kumar, Bhaktipriya Radharapu, Olivia Sturman, Oscar Wahltinez

TL;DR
ShieldGemma introduces advanced LLM-based safety content moderation models that outperform existing solutions and includes a novel data curation pipeline, enhancing safety and generalization in AI-generated content moderation.
Contribution
The paper presents ShieldGemma, a new suite of LLM-based safety moderation models with superior performance and a novel data curation pipeline, advancing LLM safety research.
Findings
Outperforms Llama Guard (+10.8% AU-PRC) and WildCard (+4.3%) on benchmarks.
Demonstrates strong generalization with synthetic data training.
Provides a new resource for the research community.
Abstract
We present ShieldGemma, a comprehensive suite of LLM-based safety content moderation models built upon Gemma2. These models provide robust, state-of-the-art predictions of safety risks across key harm types (sexually explicit, dangerous content, harassment, hate speech) in both user input and LLM-generated output. By evaluating on both public and internal benchmarks, we demonstrate superior performance compared to existing models, such as Llama Guard (+10.8\% AU-PRC on public benchmarks) and WildCard (+4.3\%). Additionally, we present a novel LLM-based data curation pipeline, adaptable to a variety of safety-related tasks and beyond. We have shown strong generalization performance for model trained mainly on synthetic data. By releasing ShieldGemma, we provide a valuable resource to the research community, advancing LLM safety and enabling the creation of more effective content…
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Code & Models
- 🤗google/shieldgemma-2bmodel· 7.2k dl· ♡ 1117.2k dl♡ 111
- 🤗google/shieldgemma-9bmodel· 1.5k dl· ♡ 261.5k dl♡ 26
- 🤗google/shieldgemma-27bmodel· 194 dl· ♡ 27194 dl♡ 27
- 🤗QuantFactory/shieldgemma-2b-GGUFmodel· 219 dl· ♡ 2219 dl♡ 2
- 🤗QuantFactory/shieldgemma-9b-GGUFmodel· 248 dl· ♡ 2248 dl♡ 2
- 🤗LiteLLMs/shieldgemma-2b-GGUFmodel· 32 dl32 dl
- 🤗LiteLLMs/shieldgemma-9b-GGUFmodel· 29 dl29 dl
- 🤗RichardErkhov/google_-_shieldgemma-2b-ggufmodel· 102 dl102 dl
- 🤗RichardErkhov/google_-_shieldgemma-9b-ggufmodel· 133 dl133 dl
- 🤗RichardErkhov/google_-_shieldgemma-27b-ggufmodel· 150 dl150 dl
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
Taxonomy
TopicsHate Speech and Cyberbullying Detection
MethodsLLaMA
