AI News selected for Professionals and Decision Makers
Primary Research Stream

SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text

06:00 · July 27, 2026 · arXiv cs.AI RSS

SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text

Schema graphs are an upstream bottleneck of schema-grounded information extraction and knowledge graph construction, yet most extraction systems assume the schema is already available. We introduce SCOPE (Schema Construction and Ontology-induction Pipeline Evaluation), a train-text-only benchmark for corpus-to-schema induction and optional schema fusion from raw text, built from 24 public information extraction sources (15 RE and 9 EE) normalized into evaluation-only gold schema graphs; its core event-extraction target covers event types and within-event argument roles, with inter-event links reported separately. We present SCION (Schema Construction and Induction with Ontology Normalization), an auditable reference pipeline rather than a new extraction architecture; it constructs candidate spaces from train text and restricts naming, merging, filtering, validation, and conservative fusion to candidate-linked evidence under strict JSON contracts. On the SCOPE core suite, SCION-lite attains the highest F1 among released source-schema references, Text2Onto-style, LLM-only, and matched extract-then-aggregate baselines under Literal, Fuzzy, Continuous, and Graph schema-graph metrics, while the compact open-model SCION-RL variant reduces reliance on proprietary LLM schema engineers. These results are reported against normalized typed-edge targets rather than as claims that induced schemas surpass human ontology design; the release includes evidence-linked outputs, parse/fallback logs, candidate retention/merging logs, run manifests, code, and benchmark packages at https://github.com/wandugu/paper_scion.

Summary

Schema graphs serve as a foundational but often overlooked prerequisite for schema-grounded information extraction and knowledge-graph construction. Most existing systems presuppose that an appropriate schema already exists, leaving the upstream task of inducing one from raw text comparatively underexplored. To address this gap, the authors present SCOPE, a train-text-only benchmark constructed from 24 publicly available information-extraction datasets—15 relation-extraction and 9 event-extraction sources—whose schemas have been normalized into evaluation-only gold graphs. The benchmark’s primary focus is event extraction, specifically event types and the argument roles that appear within individual events; inter-event relations are tracked separately.

Alongside the benchmark they release SCION, an auditable reference pipeline rather than a novel extraction architecture. SCION first assembles candidate schema elements directly from the training text, then applies naming, merging, filtering, validation, and conservative fusion steps exclusively to evidence-linked candidates. All operations are constrained by strict JSON contracts that record every decision, enabling inspection of parse logs, fallback behavior, candidate retention, and merging history. On the SCOPE core suite, the SCION-lite configuration records the highest F1 scores among published baselines—including source-schema references, Text2Onto-style pipelines, LLM-only approaches, and extract-then-aggregate methods—across Literal, Fuzzy, Continuous, and Graph schema-graph metrics.

A compact open-model variant, SCION-RL, achieves comparable results while eliminating dependence on proprietary large-language-model schema engineers. The authors emphasize that performance figures are measured against the normalized typed-edge targets supplied by SCOPE and do not constitute claims of superiority over human-designed ontologies. The full release comprises evidence-linked outputs, run manifests, benchmark packages, and source code, supporting reproducible experimentation in schema induction for knowledge-graph construction.

Why it matters

This research is highly relevant for Dutch AI researchers and enterprises developing knowledge graphs and information extraction systems. Its emphasis on an auditable pipeline and open models aligns perfectly with the Netherlands' strategic focus on transparent, ethical, and sovereign AI solutions.

More in this beat
evaluation-benchmarksgenerative-ontology-inductionknowledge-graphsreproducibility-assetsscionscope
Featuring Every Eval Ever Results on Hugging Face Model Pages

02:00 · June 30, 2026

Featuring Every Eval Ever Results on Hugging Face Model Pages

This article provides ML Engineers with a concrete, actionable MLOps tool to standardize model evaluation and benchmarking. By adopting the EEE schema, Dutch AI teams can ensure reproducibility and transparency in their model deployments, which is increasingly important for compliance with EU AI regulations and building trust in enterprise AI solutions.

Relevance 85 · Audience 95

Self-Evolving Agents as Dynamic Graph Transformation: A Survey and New Perspective

06:00 · August 20, 2026

Self-Evolving Agents as Dynamic Graph Transformation: A Survey and New Perspective

The paper provides foundational research on making autonomous AI agents auditable, safe, and transparent through dynamic graph modeling. This aligns strongly with the Dutch and EU focus on ethical AI and regulatory compliance, offering advanced researchers actionable frameworks for building governable agentic systems.

Relevance 85 · Audience 95

FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

06:00 · August 20, 2026

FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

This research is highly relevant for Dutch AI researchers and the strong local fintech and banking sector exploring customer-facing LLM agents. It provides a rigorous, reproducible framework to test agent compliance and security against fraud, aligning with strict EU financial and AI regulations.

Relevance 85 · Audience 95

Position: Behavioral Systems Require Behavioral Tests

06:00 · August 20, 2026

Position: Behavioral Systems Require Behavioral Tests

The article is highly relevant for Dutch AI researchers and practitioners focused on ethical and transparent AI. By proposing behavioral tests to evaluate AI alignment, safety, and decision-making processes, it provides a crucial methodological framework that supports compliance with EU regulations like the AI Act and advances responsible AI deployment.

Relevance 85 · Audience 95

ASI-Bench: At the Dawn of Artificial Superintelligence

06:00 · August 19, 2026

ASI-Bench: At the Dawn of Artificial Superintelligence

Offers a novel, high-depth evaluation framework that Dutch AI researchers and advanced labs can directly apply to measure progress toward autonomous scientific agents, aligning with the Netherlands' strengths in ethical AI and SME-driven innovation.

Relevance 62 · Audience 88

FLOPs vs Real Work: The Importance of Replication in AI Efficiency Assessment

06:00 · August 18, 2026

FLOPs vs Real Work: The Importance of Replication in AI Efficiency Assessment

Directly relevant for Dutch AI researchers and advanced practitioners working on Green AI, model optimization, and reproducible efficiency metrics; authors are local, findings address EU energy concerns, and results are actionable for accurate cost assessment on modern GPUs.

Relevance 85 · Audience 90

Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture

06:00 · August 18, 2026

Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture

This article is highly relevant for Dutch AI researchers and practitioners focused on ethical AI, aligning strongly with the Netherlands' and EU's emphasis on transparent and trustworthy AI systems. It provides a critical framework for advancing LLM evaluation beyond simple value alignment toward robust normative reasoning.

Relevance 85 · Audience 95

Measuring Cross-Task Behavioral Consistency in Language Model Agents

06:00 · August 17, 2026

Measuring Cross-Task Behavioral Consistency in Language Model Agents

The article provides a novel, quantifiable method for assessing the reliability and behavioral consistency of AI agents, which is crucial for compliance with EU AI regulations and the Dutch focus on transparent AI. Researchers can directly apply the open-source BCM framework to evaluate and improve the predictability of enterprise AI deployments.

Relevance 85 · Audience 95