Probing the Knowledge Boundary: An Interactive Agentic Framework for Deep Knowledge Extraction

Yuheng Yang; Siqi Zhu; Tao Feng; Ge Liu; Jiaxuan You

arXiv:2602.00959·cs.LG·February 3, 2026

Probing the Knowledge Boundary: An Interactive Agentic Framework for Deep Knowledge Extraction

Yuheng Yang, Siqi Zhu, Tao Feng, Ge Liu, Jiaxuan You

PDF

Open Access

TL;DR

This paper introduces an interactive framework for systematically probing and quantifying the knowledge contained in large language models, revealing insights into their knowledge boundaries, scaling laws, and differences across model types.

Contribution

The paper presents a novel interactive probing method with adaptive policies and a multi-stage knowledge processing pipeline, advancing systematic knowledge extraction from LLMs.

Findings

01

Recursive taxonomy is the most effective exploration strategy.

02

Larger models extract more knowledge, following a scaling law.

03

Domain-specialized models have higher initial accuracy but degrade faster.

Abstract

Large Language Models (LLMs) can be seen as compressed knowledge bases, but it remains unclear what knowledge they truly contain and how far their knowledge boundaries extend. Existing benchmarks are mostly static and provide limited support for systematic knowledge probing. In this paper, we propose an interactive agentic framework to systematically extract and quantify the knowledge of LLMs. Our method includes four adaptive exploration policies to probe knowledge at different granularities. To ensure the quality of extracted knowledge, we introduce a three-stage knowledge processing pipeline that combines vector-based filtering to remove exact duplicates, LLM-based adjudication to resolve ambiguous semantic overlaps, and domain-relevance auditing to retain valid knowledge units. Through extensive experiments, we find that recursive taxonomy is the most effective exploration strategy.…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling · Advanced Graph Neural Networks · Natural Language Processing Techniques