Training Language Models to Use Prolog as a Tool

Niklas Mellgren; Peter Schneider-Kamp; Lukas Galke Poech

arXiv:2512.07407·cs.CL·April 21, 2026

Training Language Models to Use Prolog as a Tool

Niklas Mellgren, Peter Schneider-Kamp, Lukas Galke Poech

PDF

1 Repo 1 Datasets

TL;DR

This paper explores fine-tuning language models to utilize Prolog for symbolic reasoning, balancing accuracy and auditability, with a new reinforcement learning approach outperforming supervised methods.

Contribution

It introduces a reinforcement learning method for training models to use Prolog, revealing a trade-off between correctness and interpretability in reasoning tasks.

Findings

01

RL fine-tuning outperforms supervised fine-tuning on GSM8K.

02

3B model achieves competitive zero-shot performance on MMLU benchmarks.

03

Identifies an accuracy--auditability trade-off influenced by reward design.

Abstract

Language models frequently produce plausible yet incorrect reasoning traces that are difficult to verify. We investigate fine-tuning models to use Prolog as an external symbolic reasoning tool, training Qwen2.5-3B-Instruct with Group Relative Policy Optimization (GRPO) on a cleaned version of GSM8K (which we release as gsm8k-prolog-prover). We systematically vary prompt structure, reward composition (execution, syntax, semantics, structure), and inference protocol (single-try, multiple-try, and two agentic modes). Our reinforcement learning approach outperforms supervised fine-tuning on GSM8K, and the resulting 3B model achieves zero-shot performance on MMLU-STEM and MMLU-Pro competitive with 7B few-shot baselines. Most importantly, we identify an accuracy--auditability trade-off: configurations tuned for correctness alone learn to delegate reasoning to natural language and use Prolog…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

aisilab/Prolog-as-a-Tool
github

Datasets

niklasm222/gsm8k-prolog-prover
dataset· 70 dl
70 dl

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.