Physics-Audited Agentic Discovery in Scientific Machine Learning
06:00 · July 9, 2026 · arXiv cs.AI RSS

In agentic scientific machine learning (SciML), large language model (LLM) agents can discover surrogate models and select one by an automated score, typically an error metric. A low error, however, does not establish that the predicted fields satisfy the physics that matter for mechanics, such as boundary conditions, superposition, stiffness scaling, or causality. We introduce Physics-Audited Agentic SciML (PA-SciML), a verification-first workflow for agentic SciML discovery. The workflow fixes a scoring evaluator before search, derives reviewable machine-checkable physics requirements, checks each trained candidate on its outputs, and separately searches prescribed input ranges or measured load-history spans for high-violation cases without reference solution fields. A surrogate is reported as verified only under the stated checks. When enabled, the workflow also adds advisory numerical probes before training and tests one modeling change at a time to record which isolated edits are associated with score gains before reuse. In the reported computational-solid-mechanics numerical examples, the static elasticity run selects a surrogate with lower validation error than the error-only baseline while both selected models pass the common linear-elastic checks. In the transient elastodynamics run, an error-only baseline with similar mean error fails a stricter causality check by responding to future parts of the loading history, while the selected surrogate passes the stated checks. The main distinction is per-candidate physics evidence on predicted fields, not a richer aggregate score.
Summary
In agentic scientific machine learning, large language model agents autonomously propose and refine surrogate models for physical systems, typically selecting finalists according to an aggregate error metric on held-out data. The paper argues that such metrics alone cannot confirm whether a model respects the underlying mechanics, such as satisfaction of boundary conditions, adherence to superposition, correct stiffness scaling, or strict causality in time-dependent problems.
PA-SciML therefore inverts the usual workflow. Before any search begins, the authors fix a scoring evaluator and translate the relevant physical laws into machine-checkable requirements that can be evaluated directly on a candidate’s predicted fields. Each trained surrogate is then audited against these requirements. In addition, the workflow systematically probes prescribed ranges of inputs or measured load histories to locate high-violation regimes, even when no reference solution is available. Only models that survive every stated check are accepted. Optional safeguards include advisory numerical probes run before training and controlled, one-at-a-time edits whose effect on the score is recorded before reuse.
The approach is illustrated on two computational-solid-mechanics tasks. In a static linear-elasticity example the physics-audited procedure returns a surrogate whose validation error is lower than that of an error-only baseline, while both models satisfy the common linear-elastic checks. In a transient elastodynamics case an error-only model with comparable mean error violates a causality constraint by responding to future portions of the loading history; the PA-SciML surrogate passes the same check. The decisive difference therefore lies not in a richer aggregate score but in explicit, per-candidate evidence that the predicted fields obey the stated physical requirements.
Why it matters
This research is highly relevant for the Dutch AI market, particularly for its strong high-tech engineering and manufacturing sectors that rely heavily on scientific machine learning and digital twins. The focus on verifiable, physics-compliant AI aligns with the EU's emphasis on trustworthy AI and provides actionable methodologies for researchers at Dutch technical universities.




