The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI
06:00 · July 9, 2026 · arXiv cs.AI RSS

Agentic AI development today runs on token maxing: buying capability with tokens -- longer reasoning traces, more turns, wider tool payloads, bigger replayed contexts -- so tokens per task grow faster than task value. Falling per-token prices mask the pattern; total spend rises anyway. We argue the decisive lever against token maxing is the harness: the orchestration layer that assembles context, exposes tools, sequences turns, delegates work, and carries enterprise observability and governance. We isolate it with a controlled swap: 22 locked evaluation tasks, six foundation models (Claude Sonnet 4.6, Gemini 3.1, Gemini Flash 3.5, Qwen 3.6, GLM 5.1, Palmyra X6), changing only the orchestration layer -- a frozen conventional production loop versus the Writer Agent Harness. Holding models constant, the harness cuts blended cost per task 41% ($0.21->$0.12), median wall-clock 44% (48s->27s), and tokens per task 38% (14.2k->8.8k), with task-completion quality at parity (0.78->0.81, directional at this sample size). Efficiency is model-invariant -- every model gets cheaper (33-61%) -- while quality gains are capability-dependent: a model's gain correlates almost perfectly with its baseline strength (r=0.99, n=6), a phenomenon we term harness leverage. Quality per dollar rises 82%; task-completions per million tokens rise from 54.9 to 92.0. On this workload the orchestration layer moved cost per task more than the full spread of the model menu did. We formalize token economics at the orchestration layer (including effective input price under prompt caching), detail the six mechanism families behind the effect -- cache-shape discipline to failure-spend governance -- compare six widely used agent systems on the same axes, and argue the harness is the one component whose efficiency multiplies across every model an organization runs -- present and future.
Summary
The paper frames current agentic AI practice as “token maxing,” in which developers purchase additional capability by expanding reasoning traces, tool payloads, context replay and turn counts, causing token consumption to outpace delivered value even as per-token prices decline. The decisive countermeasure, the authors argue, lies not in model selection but in the orchestration layer they term the harness: the component that assembles context, exposes tools, sequences turns, delegates subtasks and embeds enterprise observability and governance controls.
To isolate the harness effect they performed a controlled swap across 22 fixed evaluation tasks and six foundation models, holding model weights and prompts constant while exchanging only the orchestration layer. Replacing a conventional production loop with the Writer Agent Harness reduced blended cost per task by 41 percent, median wall-clock latency by 44 percent and tokens per task by 38 percent, while task-completion quality remained statistically indistinguishable. Efficiency gains proved model-invariant, with every model registering cost reductions between 33 and 61 percent; quality improvements, by contrast, scaled with each model’s baseline capability (Pearson r = 0.99).
The work formalizes token economics at the orchestration layer, incorporating effective input pricing under prompt caching, and enumerates six mechanism families that drive the observed savings, ranging from cache-shape discipline to failure-spend governance. On the tested workload the orchestration layer shifted cost per task more than the entire spread among the six evaluated models. The authors further note that the harness multiplies efficiency across every model an organization deploys, present and future, and therefore constitutes a high-leverage target for enterprise governance and cost control.
Why it matters
Provides actionable, reproducible insights into orchestration design that Dutch AI teams and enterprises can directly apply to lower costs of agentic systems; presents novel empirical findings and formalization suitable for advanced researchers.






