Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models

Martin Pawelczyk; Lillian Sun; Zhenting Qi; Aounon Kumar and; Himabindu Lakkaraju

arXiv:2501.00418·cs.LG·January 3, 2025

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models

Martin Pawelczyk, Lillian Sun, Zhenting Qi, Aounon Kumar and, Himabindu Lakkaraju

PDF

Open Access

TL;DR

This paper investigates whether trustworthiness properties like robustness and fairness can transfer from weaker to stronger language models through fine-tuning, revealing mixed results and highlighting the potential and limits of weak-to-strong trustworthiness generalization.

Contribution

It introduces novel training strategies for enhancing trustworthiness transfer from weak to strong models and provides the first empirical study on this form of generalization.

Findings

01

Fairness and robustness improve with regularization transfer.

02

Privacy does not significantly benefit from weak-to-strong transfer.

03

Some trustworthiness properties are more transferable than others.

Abstract

The rapid proliferation of generative AI, especially large language models, has led to their integration into a variety of applications. A key phenomenon known as weak-to-strong generalization - where a strong model trained on a weak model's outputs surpasses the weak model in task performance - has gained significant attention. Yet, whether critical trustworthiness properties such as robustness, fairness, and privacy can generalize similarly remains an open question. In this work, we study this question by examining if a stronger model can inherit trustworthiness properties when fine-tuned on a weaker model's outputs, a process we term weak-to-strong trustworthiness generalization. To address this, we introduce two foundational training strategies: 1) Weak Trustworthiness Finetuning (Weak TFT), which leverages trustworthiness regularization during the fine-tuning of the weak model, and…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAccess Control and Trust