LoRA-Augmented Generation (LAG) for Knowledge-Intensive Language Tasks

William Fleshman; Benjamin Van Durme

arXiv:2507.05346·cs.CL·August 19, 2025

LoRA-Augmented Generation (LAG) for Knowledge-Intensive Language Tasks

William Fleshman, Benjamin Van Durme

PDF

Open Access

TL;DR

LoRA-Augmented Generation (LAG) is a method that efficiently leverages knowledge and task-specific adapters in large language models without additional training, improving performance on knowledge-intensive tasks.

Contribution

LAG introduces a data-free, efficient approach to select and combine knowledge experts in language models, compatible with retrieval-based methods.

Findings

01

LAG outperforms existing data-free methods on knowledge-intensive tasks.

02

LAG is compatible with retrieval-augmented generation approaches.

03

LAG requires no additional training or data access.

Abstract

The proliferation of fine-tuned language model experts for specific tasks and domains signals the need for efficient selection and combination methods. We propose LoRA-Augmented Generation (LAG) for leveraging large libraries of knowledge and task-specific LoRA adapters. LAG requires no additional training or access to data, and efficiently filters, retrieves, and applies experts on a per-token and layer basis. We evaluate LAG on various knowledge-intensive tasks, achieving superior performance over existing data-free methods. We explore scenarios where additional data is available, demonstrating LAG's compatibility with alternative solutions such as retrieval-augmented generation (RAG).

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling · Multimodal Machine Learning Applications · Artificial Intelligence in Healthcare and Education