Interventional Grounding Audits: Black-Box Premise-Dependency Tests for LLM Chain-of-Thought via Predicate Substitution
06:00 · July 16, 2026 · arXiv cs.AI RSS

Large language models produce chain-of-thought (CoT) reasoning that appears logically sound yet may not genuinely depend on its stated premises. We introduce interventional grounding audits, a black-box, step-level test of premise dependency: we intervene on a single premise by substituting its target predicate with a fresh symbol, re-run the model, and check whether each reasoning step's normalized conclusion (canonical predicate form) changes. We evaluate on ProntoQA, a synthetic multi-hop deductive reasoning benchmark with gold proof trees, where step-level premise dependencies are known. Applied to 50 ProntoQA problems with GPT-4o, our method achieves F1 = 0.806 on detecting proof-tree dependencies (F1 = 0.885 on predicate-determining dependencies; Recall = 100%), significantly outperforming a self-consistency baseline (F1 = 0.343; 95% bootstrap CIs non-overlapping). We further identify that 66% of correctly-solved problems contain at least one aligned step insensitive to a direct proof-tree dependency under consistent substitution -- all involving entity-introduction premises, a documented blind spot of the consistent-substitution evaluator -- a "right answer, wrong reasoning" signal invisible to passive methods. All audit certificates, raw outputs, and reproduction scripts are available in a public GitHub repository, and we discuss scope limits beyond formal, parsable benchmarks.
Summary
Interventional grounding audits provide a black-box technique for verifying whether individual steps in an LLM’s chain-of-thought genuinely depend on the premises they cite. The method works by substituting a target predicate in a selected premise with a fresh symbol, re-running the model, and checking whether the normalized conclusion of each subsequent reasoning step changes. Two substitution regimes are used: consistent replacement across all premises, which isolates predicate-determining dependencies, and local replacement confined to a single premise, which additionally surfaces transitive and structural dependencies. A cascade filter then removes propagation artifacts that would otherwise inflate false positives downstream.
Evaluated on 50 ProntoQA problems with GPT-4o, the approach yields an F1 of 0.806 for recovering proof-tree dependencies and 0.885 when restricted to predicate-determining cases, compared with 0.343 for a self-consistency baseline. Recall on predicate-determining dependencies reaches 100 percent. The same audits flag “right answer, wrong reasoning” patterns—correct final answers accompanied by at least one step insensitive to a direct proof-tree premise—in 66 percent of solved problems, all traceable to entity-introduction premises that passive consistency checks overlook. Every audit certificate, raw trace, and reproduction script is released with SHA-256 verification in a public repository, enabling independent validation of the reported metrics.
Why it matters
Directly supports trustworthy and transparent AI evaluation, a core Dutch/EU priority. The black-box protocol and open artifacts are actionable for Dutch researchers and advanced practitioners auditing LLM reliability.






