Can Safety Emerge from Weak Supervision? A Systematic Analysis of Small Language Models

Punyajoy Saha; Sudipta Halder; Debjyoti Mondal; Subhadarshi Panda

arXiv:2603.07017·cs.CL·March 10, 2026

Can Safety Emerge from Weak Supervision? A Systematic Analysis of Small Language Models

Punyajoy Saha, Sudipta Halder, Debjyoti Mondal, Subhadarshi Panda

PDF

Open Access

TL;DR

This paper presents Self-MOA, an automated framework that uses weak supervision to align small language models for safety and helpfulness, reducing reliance on costly human annotations.

Contribution

Introduction of Self-MOA, a fully automated, weak supervision-based method for aligning small language models with improved safety and maintained helpfulness.

Findings

01

12.41% safety improvement across benchmarks

02

Uses 11 times less data than human supervision

03

Effective in resource-constrained settings

Abstract

Safety alignment is critical for deploying large language models (LLMs) in real-world applications, yet most existing approaches rely on large human-annotated datasets and static red-teaming benchmarks that are costly, difficult to scale, and slow to adapt to evolving model behaviors. Moreover, overly conservative safety mechanisms can reduce model usefulness by rejecting sensitive but legitimate queries. We introduce Self-MOA (Self Multi-Objective Alignment), a fully automated framework for aligning small language models using weak supervision from automated evaluator models. Self-MOA operates as a closed loop that dynamically generates model-specific red team prompts, constructs preference data from model-generated responses, and aligns models via multi-objective preference optimization to jointly optimize for safety and helpfulness. Across multiple small language models and safety…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling · Adversarial Robustness in Machine Learning · Explainable Artificial Intelligence (XAI)