TL;DR
This paper provides a theoretical understanding of Lifelong Normalization in sequential model editing, revealing its role in stability and proposing improvements for long-term performance.
Contribution
It offers the first theoretical analysis of Lifelong Normalization, explaining its stabilizing effects and introducing StableEdit to enhance long-horizon stability.
Findings
Removing LN causes performance collapse.
Early edits can positively influence future edits.
StableEdit improves long-term stability with minimal overhead.
Abstract
Lifelong Model Editing aims to continuously update evolving facts in Large Language Models while preserving unrelated knowledge and general capabilities, yet it remains plagued by catastrophic forgetting and model collapse. Empirically, we find that recent editors resilient over long horizons share the same core strategy: Lifelong Normalization (LN), which normalizes value gradients using running statistics. Removing LN causes immediate performance collapse, and we observe a counter-intuitive positive cumulative effect where early edits can promote the success of future edits. Yet the mechanism of LN remains a "black box", leaving its precise role in lifelong stability poorly understood. In this work, we provide the first theoretical account of LN in the lifelong regime. Our analysis reveals a self-reinforcing stability loop and proves that, when combined with ridge-regularized…
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Code & Models
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
