To trust or not to trust: Attention-based Trust Management for LLM Multi-Agent Systems

Pengfei He; Zhenwei Dai; Xianfeng Tang; Yue Xing; Hui Liu; Jingying Zeng; Qiankun Peng; Shrivats Agrawal; Samarth Varshney; Suhang Wang; Jiliang Tang; Qi He

arXiv:2506.02546·cs.CR·April 15, 2026

To trust or not to trust: Attention-based Trust Management for LLM Multi-Agent Systems

Pengfei He, Zhenwei Dai, Xianfeng Tang, Yue Xing, Hui Liu, Jingying Zeng, Qiankun Peng, Shrivats Agrawal, Samarth Varshney, Suhang Wang, Jiliang Tang, Qi He

PDF

TL;DR

This paper introduces a comprehensive trust management system for LLM-based multi-agent systems, using an attention-based score to evaluate message trustworthiness and improve robustness against malicious inputs.

Contribution

It proposes a novel holistic trustworthiness framework with six dimensions and an attention-based trust score, enhancing message and agent trust assessments in LLM-MAS.

Findings

01

The trust management system improves robustness against malicious messages.

02

The Attention Trust Score effectively evaluates message trustworthiness.

03

Experiments demonstrate significant performance gains across diverse tasks.

Abstract

Large Language Model-based Multi-Agent Systems (LLM-MAS) have demonstrated strong capabilities in solving complex tasks but remain vulnerable when agents receive unreliable messages. This vulnerability stems from a fundamental gap: LLM agents treat all incoming messages equally without evaluating their trustworthiness. While some existing studies approach trustworthiness, they focus on a single type of harmfulness rather than analyze it in a holistic approach from multiple trustworthiness perspectives. We address this gap by proposing a comprehensive definition of trustworthiness inspired by human communication theory (Grice, 1975). Our definition identifies six orthogonal trust dimensions that provide interpretable measures of trustworthiness. Building on this definition, we introduce the Attention Trust Score (A -Trust), a lightweight, attention-based method for evaluating the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.