Characterizing Model-Native Skills

Feiyang Kang; Mahavir Dabas; Myeongseob Ko; Ruoxi Jia

arXiv:2604.17614·cs.AI·April 21, 2026

Characterizing Model-Native Skills

Feiyang Kang, Mahavir Dabas, Myeongseob Ko, Ruoxi Jia

PDF

1 Repo

TL;DR

This paper proposes a model-native approach to skill characterization by recovering an interpretable basis from model activations, enabling more effective data selection and steering for improving model performance and safety.

Contribution

It introduces a novel method to recover an internal, interpretable basis from model activations that captures behavioral axes without relying on external ontologies.

Findings

01

Data selection along model-native directions improves reasoning accuracy.

02

Inference-time steering using these directions enhances model performance.

03

Model-native skill coverage leads to more sample-efficient safety alignment.

Abstract

Skills are a natural unit for describing what a language model can do and how its behavior can be changed. However, existing characterizations rely on human-written taxonomies, textual descriptions, or manual profiling pipelines--all external hypotheses about what matters that need not align with the model's internal representations. We argue that when the goal is to intervene on model behavior, skill characterization should be *model-native*: grounded in the model's own representations rather than imposed through external ontologies. We instantiate this view by recovering a compact orthogonal basis from sequence-level activations. The resulting basis is semantically interpretable but need not correspond to any predefined human ontology; instead, it captures axes of behavioral variation that the model itself organizes around. We validate this characterization on reasoning post-training,…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

null
github

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.