The Reciprocity Gradient

Yue Lin; Pascal Poupart; Shuhui Zhu; Dan Qiao; Wenhao Li; Yuan Liu; Hongyuan Zha; Baoxiang Wang

arXiv:2605.08323·cs.LG·May 12, 2026

The Reciprocity Gradient

Yue Lin, Pascal Poupart, Shuhui Zhu, Dan Qiao, Wenhao Li, Yuan Liu, Hongyuan Zha, Baoxiang Wang

PDF

TL;DR

The paper introduces the reciprocity gradient, a method for learning context-sensitive policies by backpropagating reward gradients through reputation chains, improving over sample-based baselines.

Contribution

It formulates the influence attribution problem and proposes the reciprocity gradient to explicitly optimize actions and signals in strategic interactions.

Findings

01

Recovers near-optimal context-sensitive policies

02

Sample-based baselines collapse into constant policies

03

Gradient flows through reputation chains analytically

Abstract

Communication is fundamental to sustaining reciprocity and cooperation in strategic interactions. We identify and formulate the influence attribution problem as the central optimization difficulty inherent in such dynamics for a learning agent: any action or signal the agent emits reshapes the reputations of many third parties along combinatorially branching paths before feeding back into its own future rewards, forcing the agent to account for all of these indirect channels at once when choosing every action. To address this, we introduce the reciprocity gradient, which explicitly backpropagates reward gradients through private estimators of opponents' policies trained from public observations. The gradient flows through the reputation chain itself analytically, rather than being estimated from sampled returns. It jointly optimizes actions and evaluative signals without intrinsic…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.