AI News selected for Professionals and Decision Makers
Primary Research Stream

SkillTrace: Multi-Trace Provenance Auditing for LLM-Agent Skill Reuse

06:00 · August 7, 2026 · arXiv cs.AI RSS

SkillTrace: Multi-Trace Provenance Auditing for LLM-Agent Skill Reuse

LLM-agent ecosystems are rapidly growing around reusable skills: mixed-modality packages of metadata, natural-language instructions, code, tools, references, and operational workflows. As skills become marketplace artifacts, auditing their reuse is no longer the same problem as ordinary code clone detection. Existing detectors target single-modality source code or whole-package similarity, yet skill reuse evidence is distributed across authored text, implementation fragments, and operational structure. As a result, they can miss reuse that preserves only one part of a skill. We present SKILLTRACE, a multi-trace provenance auditing framework for LLM-agent skill reuse. SKILLTRACE extracts three provenance traces: Expression, Implementation, and Operational. It represents the Operational Trace as a Skill Operational Graph (SOG) that captures activation, procedure, and resource-flow structure. An LLM assists only the Operational-trace extraction, once at ingestion; at audit time SKILLTRACE compares cached traces deterministically, calibrates each trace against same-function strict negatives, and reports which trace supports a reuse decision. On SKILLTRACE-BENCH, with 820 transformed reuse positives over 100 marketplace anchors and 751 negative controls, SKILLTRACE achieves AUROC 0.938 and F1 0.898. A 36,446-skill wild audit further shows that trace-attributed evidence surfaces actionable reuse review queues beyond repository-level baselines.

Summary

SkillTrace addresses the growing challenge of auditing reuse within LLM-agent skill marketplaces, where skills function as mixed-modality packages containing natural-language instructions, executable code, tool interfaces, and operational workflows. Unlike conventional code-clone detectors that operate on single-modality source or whole-package similarity, the framework recognizes that reuse evidence can survive in only one of these layers after transformation. It therefore extracts three distinct provenance traces from each skill: an Expression trace for authored text, an Implementation trace for scripts and API patterns, and an Operational trace that records activation logic, task procedures, and resource flows.

The Operational trace is encoded as a Skill Operational Graph (SOG) whose three views capture the structural elements most likely to persist under rewriting. Extraction of the SOG relies on an LLM only once, at ingestion time; subsequent audits compare cached traces deterministically and calibrate each modality against same-function negative examples to reduce false attribution. The resulting report indicates which trace, if any, supports a reuse finding, allowing reviewers to distinguish inherited artifacts from independent implementations that merely solve the same task.

Evaluation on SkillTrace-Bench, comprising 820 transformed reuse cases derived from 100 marketplace anchors together with 751 strict negative controls, yields an AUROC of 0.938 and an F1 score of 0.898. A separate audit of 36,446 publicly available skills demonstrates that the trace-attributed evidence produces review queues containing reuse relations that repository-level fingerprinting methods overlook or rank too low. By separating provenance signals from functional similarity, SkillTrace supplies the visibility required for governance, deduplication, and remediation tasks in expanding LLM-agent ecosystems.

Why it matters

This research is highly relevant for Dutch AI researchers and enterprises focused on AI governance, IP protection, and compliance with EU transparency regulations. It provides a rigorous, actionable methodology for auditing LLM-agent ecosystems, which is crucial for maintaining ethical and transparent AI marketplaces.

More in this beat
agent-skillsdata-provenancedocument-deduplicationevaluation-benchmarksllm-agentsSkillTrace
AI Tool Discovery at Scale: All You Need is DNS

06:00 · July 22, 2026

AI Tool Discovery at Scale: All You Need is DNS

This research is highly relevant for Dutch AI infrastructure developers and researchers building multi-agent systems. Its decentralized governance model aligns well with European data sovereignty and transparent AI goals, offering a scalable alternative to centralized tool registries.

Relevance 85 · Audience 95

Self-Evolving Agents as Dynamic Graph Transformation: A Survey and New Perspective

06:00 · August 20, 2026

Self-Evolving Agents as Dynamic Graph Transformation: A Survey and New Perspective

The paper provides foundational research on making autonomous AI agents auditable, safe, and transparent through dynamic graph modeling. This aligns strongly with the Dutch and EU focus on ethical AI and regulatory compliance, offering advanced researchers actionable frameworks for building governable agentic systems.

Relevance 85 · Audience 95

FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

06:00 · August 20, 2026

FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

This research is highly relevant for Dutch AI researchers and the strong local fintech and banking sector exploring customer-facing LLM agents. It provides a rigorous, reproducible framework to test agent compliance and security against fraud, aligning with strict EU financial and AI regulations.

Relevance 85 · Audience 95

Harnessing agent memory to build lifelong AI partners for materials scientists

06:00 · August 13, 2026

Harnessing agent memory to build lifelong AI partners for materials scientists

This research is highly relevant for Dutch AI researchers and high-tech materials enterprises looking to deploy autonomous AI agents for R&D. The proposed model-agnostic memory framework addresses critical challenges in AI reproducibility and workflow efficiency, offering actionable methodologies for advanced scientific computing.

Relevance 85 · Audience 95

SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents

06:00 · August 7, 2026

SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents

This research is highly relevant for Dutch AI researchers and developers building autonomous agents, as it provides a robust framework for auditing and debugging complex AI behaviors. Furthermore, its focus on transparency and error attribution aligns strongly with EU AI Act requirements for reliable and accountable AI systems.

Relevance 85 · Audience 95

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

06:00 · August 4, 2026

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

This research is highly relevant for Dutch AI researchers developing autonomous LLM agents, providing a rigorous framework for evaluating continuous learning in realistic deployment scenarios. Understanding how model capabilities gate self-evolution is crucial for building robust and reliable AI systems.

Relevance 85 · Audience 95

Household Movement Detection in Mixed-Format Occupancy Data Using LLM-Based Entity Resolution

06:00 · July 27, 2026

Household Movement Detection in Mixed-Format Occupancy Data Using LLM-Based Entity Resolution

The methodology is highly actionable for Dutch AI researchers and data scientists working with administrative registries, census data, or customer databases. Its focus on handling noisy data without explicit identifiers aligns well with EU GDPR constraints, offering a robust approach for privacy-preserving entity resolution in public and private sectors.

Relevance 85 · Audience 95