TraceCoder: Explainable and Auditable Code Generation with Position-Key Snippet Versioning
06:00 · July 30, 2026 · arXiv cs.AI RSS

Contemporary LLM-based coding agents produce code as black-box outputs: the rationale behind each line is hidden, the evolution of the code through benchmark-driven repair is ephemeral, and post-hoc auditing is impossible. We present a code generation concept that addresses these shortcomings through three complementary mechanisms: (i) a relational snippet-history schema that records, per repair event, the benchmark reference, round number, failure text, and LLM explanation, enabling full provenance queries; (ii) a browser-based visualisation tool that renders this history as heat-mapped, hover-annotated source code; and (iii) a competitive fractional position-key indexing scheme with tree-node delimiters that assigns stable, lexicographically-ordered identifiers to each code snippet, enabling fine-grained tracking without disrupting surrounding lines. We evaluate TraceCoder on 30 algorithmic programming tasks spanning string processing, mathematical computation, and data-structure manipulation, across two provider configurations. Of these, 10 exhaust the 6-iteration budget on tasks with subtle edge-case behaviour. Mean Chg% reaches 30%, three in ten code snippets carry a traceable repair-event row, compared to 21% when using Gemini 2.0 Flash as sole provider on a 20-task subset. Three detailed case studies demonstrate how the system explains which specific benchmark failures shaped each line of the final program. The proposed mechanism makes the internal "narrative" of automated code generation auditable and replayable, a property essential for trust and accountability in production deployments.
Summary
TraceCoder addresses the opacity of LLM-driven coding agents, which typically discard the causal history behind each generated line and treat the final program as an atomic artefact. The system captures that history at snippet granularity by maintaining a relational database that logs every repair event. For each modification triggered by a benchmark run, it stores the benchmark identifier, iteration round, failure message, and the LLM’s accompanying explanation, creating a queryable provenance trail without overwriting prior rows.
A second component is a fractional position-key indexing scheme that assigns stable, lexicographically ordered identifiers to individual snippets. Inspired by collaborative-editing techniques, the scheme uses tree-node delimiters and competitive fractional keys so that insertions, deletions, or in-place edits preserve surrounding line order without rebalancing or custom comparators. This enables fine-grained tracking while leaving the source code itself unchanged for the developer.
A browser-based viewer renders the stored history directly over the code, applying heat-map colouring to indicate change intensity and pop-up panels that display the linked benchmark failures and explanations on hover. In an evaluation across 30 algorithmic tasks covering string processing, mathematics, and data structures, the approach maintained traceable repair records for roughly three in ten snippets—higher than the 21 percent observed with a single-provider baseline—while incurring only modest storage overhead. Three case studies illustrate how specific test failures can be directly linked to the lines they ultimately shaped, supporting post-hoc auditing and replay of the generation process.
Why it matters
Directly supports EU-aligned requirements for transparent, auditable AI systems; actionable for Dutch researchers and SMEs building compliant coding agents.







