AI News selected for Professionals and Decision Makers
Primary Research Stream

From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents

06:00 · July 11, 2026 · arXiv cs.AI RSS

From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents

Enterprise large language model (LLM) applications often begin as prototypes whose behavior is carried by prompts and retrieval context. Productization adds requirements for source boundaries, entity routing, answer contracts, and reproducible traces. We present a harness-engineering approach that reconstructs this pattern into a traceable, auditable LLM-agent architecture: deterministic behavior moves into code, manifests, schemas, and validation artifacts around a replaceable composition boundary, while source-backed claims remain the authority for runtime answers. We instantiate it on a public-data slice of five Korean corporate groups (25 listed companies) and evaluate three research questions. (1) The harness preserves its source-grounding, entity-routing, trace, output-hygiene, and recommendation-language contracts across the fixed validation scenarios; a fault-injection control confirms the validators flag deliberately broken contracts. (2) The checks the harness enforces held under model substitution: across three hosted models, they passed on all 270 composition-boundary runs; failures were confined to the model-composed side and were caught and recorded. (3) The code-owned guarantees are load-bearing, not reproducible by prompting alone: holding the model fixed and varying only the enforcement layer, prompt instructions alone let recommendation-language and internal-trace-leakage violations reach the reader, which the harness blocks entirely. A bolt-on external guardrail prevents such violations too but over-refuses, dropping utility to 88/120 where the harness preserves full utility (120/120); in this ablation, only code-owned enforcement preserves both safety and utility. The result is a reusable engineering pattern for turning exploratory prototypes into auditable applications with versioned source, control, and validation artifacts.

Summary

The paper addresses the common pattern in which enterprise LLM applications begin as prompt-dominant prototypes whose core behavior—source selection, entity routing, claim eligibility, and output constraints—resides chiefly in natural-language instructions and retrieval context. Productization exposes the limits of this approach: prompts cannot reliably enforce traceability to bounded sources, versioned contracts, or reproducible traces once answers reach users. The authors therefore reconstruct the prototype as a harness-engineered architecture in which deterministic control moves into code, manifests, schemas, and validation artifacts that surround a replaceable composition boundary. Within this boundary the language model handles only phrasing; source-backed claims, routing metadata, answer contracts, and trace generation remain authoritative and auditable.

The method is instantiated on a public-data slice covering five Korean corporate groups and 25 listed companies, yielding 113 source-backed runtime claims. A source-to-claim pipeline separates raw documents, evidence records, runtime-eligible claims, maintained wiki context, and reader-facing answers, while manifests define permissible sources and routing rules bind questions to entities. Validation checks source grounding, entity routing, trace completeness, output hygiene, and recommendation-language constraints. Three research questions guide the evaluation. The first confirms that the harness preserves its contracts across fixed validation scenarios, with fault injection demonstrating that validators correctly flag deliberate violations. The second shows that the same guarantees hold under model substitution: across three hosted models and 270 composition-boundary runs, all harness-enforced checks passed while model-side failures were isolated and recorded.

The third question tests whether code-owned enforcement is load-bearing. With the model fixed, prompt instructions alone permitted recommendation-language and internal-trace violations to reach the reader; the harness blocked every such case. A bolt-on external guardrail also prevented violations yet reduced utility through over-refusal, whereas the harness maintained full utility. The resulting pattern therefore supplies versioned source, control, and validation artifacts that turn exploratory prototypes into auditable enterprise agents without tying guarantees to any single model or prompt formulation.

Why it matters

Directly actionable for Dutch/EU teams building compliant LLM agents; aligns with Netherlands emphasis on ethical, transparent AI and EU regulatory needs for auditability. Offers novel, technically rigorous methodology with high reproducibility for researchers and advanced practitioners.

More in this beat
agent-safetyai-agentsdeployment-readinessharness-engineeringllm-agentsmlops-deploymentpaper-key-findings
From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents

06:00 · July 9, 2026

From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents

This research is highly relevant for Dutch AI researchers and enterprise developers building autonomous agents, as it offers a novel method to reduce reasoning overhead and API costs while improving reliability. The transition from static tools to self-evolving SOPs aligns well with the Dutch market's focus on scalable, efficient AI automation for SMEs.

Relevance 85 · Audience 95

Agentic Data Environments

06:00 · July 9, 2026

Agentic Data Environments

This research is highly relevant for Dutch AI researchers and engineers focusing on safe and reliable AI deployment. By proposing a framework that enforces safety guarantees for autonomous agents, it aligns strongly with the EU AI Act's emphasis on risk management and the Netherlands' strategic focus on ethical AI.

Relevance 85 · Audience 90

NVIDIA Nemotron Achieves Benchmark-Leading Performance With LangChain Deep Agents Harness

17:00 · July 8, 2026

NVIDIA Nemotron Achieves Benchmark-Leading Performance With LangChain Deep Agents Harness

This development is highly relevant as it offers a cost-effective, open-source alternative to closed AI models, which is crucial for driving AI adoption among Dutch SMEs. Furthermore, the ability to run these agents on proprietary infrastructure aligns perfectly with European data sovereignty and strict AI governance requirements.

Relevance 85 · Audience 75

How NVIDIA’s Inference Software Stack Powers the Lowest Token Cost

17:00 · June 30, 2026

How NVIDIA’s Inference Software Stack Powers the Lowest Token Cost

This article is relevant because it addresses a critical bottleneck in AI adoption: inference costs. For Dutch enterprises and SMEs scaling AI from pilots to production, understanding how software optimizations lower the cost per token is essential for sustainable AI deployment.

Relevance 75 · Audience 65

SkillHarness: Harnessing Safe Skills for Computer-Use Agents

06:00 · June 23, 2026

SkillHarness: Harnessing Safe Skills for Computer-Use Agents

This research directly supports the Dutch and EU strategic focus on safe, ethical, and reliable AI deployment. For researchers and advanced practitioners in the Netherlands, it provides actionable methodologies to build autonomous agents that comply with stringent safety constraints in dynamic environments.

Relevance 85 · Audience 90

Specifying AI-SDLC Processes: A Protocol Language for Human-Agent Boundaries

06:00 · June 23, 2026

Specifying AI-SDLC Processes: A Protocol Language for Human-Agent Boundaries

This research is highly relevant for the Dutch AI market due to its strong alignment with EU AI Act requirements for human oversight and governance. By providing a formal language to enforce human-agent boundaries, it offers researchers and enterprises a rigorous method to build compliant, transparent, and safe multi-agent systems.

Relevance 85 · Audience 95

Build your own vulnerability harness

19:59 · June 18, 2026

Build your own vulnerability harness

Directly actionable for Dutch security teams building or adapting AI security pipelines; addresses core security risks of AI agents in vulnerability discovery and offers concrete mitigation patterns relevant under EU contexts.

Relevance 85 · Audience 90

Scaling Managed Agents: Decoupling the brain from the hands

02:00 · April 8, 2026

Scaling Managed Agents: Decoupling the brain from the hands

Highly actionable for Product Teams and Builders implementing agent workflows with Claude, including code-level interface patterns, security mitigations, and performance gains like reduced TTFT. Directly addresses model updates, harness evolution, and production observability.

Relevance 80 · Audience 85