Evolving Scientific Discovery by Unifying Data and Background Knowledge   with AI Hilbert

Ryan Cory-Wright; Cristina Cornelio; Sanjeeb Dash; Bachir El Khadir,; Lior Horesh

arXiv:2308.09474·cs.AI·March 24, 2025·1 cites

Evolving Scientific Discovery by Unifying Data and Background Knowledge with AI Hilbert

Ryan Cory-Wright, Cristina Cornelio, Sanjeeb Dash, Bachir El Khadir,, Lior Horesh

PDF

Open Access 1 Repo

TL;DR

This paper introduces a method that unifies data and background knowledge to automatically discover scientific laws, using polynomial optimization and logical constraints, demonstrated on classical laws like Kepler's Third Law.

Contribution

It presents a novel approach combining polynomial equalities, mixed-integer optimization, and Positivstellensatz certificates to derive scientific laws from data and background theory.

Findings

01

Successfully derived Kepler's Third Law from data and axioms.

02

Derived Hagen-Poiseuille Equation using the proposed method.

03

Validated the approach on gravitational wave power equation.

Abstract

The discovery of scientific formulae that parsimoniously explain natural phenomena and align with existing background theory is a key goal in science. Historically, scientists have derived natural laws by manipulating equations based on existing knowledge, forming new equations, and verifying them experimentally. In recent years, data-driven scientific discovery has emerged as a viable competitor in settings with large amounts of experimental data. Unfortunately, data-driven methods often fail to discover valid laws when data is noisy or scarce. Accordingly, recent works combine regression and reasoning to eliminate formulae inconsistent with background theory. However, the problem of searching over the space of formulae consistent with background theory to find one that best fits the data is not well-solved. We propose a solution to this problem when all axioms and scientific laws are…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

IBM/AI-Hilbert
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsData Management and Algorithms · Constraint Satisfaction and Optimization · Bayesian Modeling and Causal Inference

Methodsfail · ALIGN