AI News selected for Professionals and Decision Makers
Primary Research Stream

DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization

06:00 · August 19, 2026 · arXiv cs.AI RSS

DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization

As text-to-image generative models advance, they raise critical safety concerns, particularly the generation of Not-Safe-For-Work (NSFW) content such as violence and nudity, further exacerbated by red-teaming adversarial attacks. Existing defenses predominantly operate under white-box assumptions, relying on text encoder optimization, weight editing, or inference-time intervention, and fundamentally cannot scale to proprietary models. Black-box alternatives based on LLM prompt rewriting offer broader applicability, yet fail in a critical regime we identify as the \textit{benign adversarial} problem: prompts that are linguistically safe but still trigger harmful generation due to the model's learned data distribution. We propose DiSCO, a zero-shot, strictly black-box defense that operates entirely at the prompt level as a plug-and-play module, requiring no model retraining, fine-tuning, or access to model internals. DiSCO performs distribution-guided suffix expansion via beam search, optimized through contrastive scoring over safe and unsafe image pools generated by the target model itself, with iterative adaptive feedback until safe content is produced. We demonstrate that DiSCO consistently enhances the safety of both undefended and defended models on the I2P benchmark under multiple red-teaming attacks, achieving 37.7% and 25.13% ASR reduction, respectively, while maintaining semantic fidelity and improving image coherence. As a black-box, architecture-agnostic module, DiSCO can be readily applied to any text-to-image system without necessitating any changes to the model itself.

Summary

DiSCO addresses a persistent safety gap in text-to-image generation by targeting the benign adversarial regime, in which linguistically innocuous prompts still elicit unsafe outputs because they align with unsafe regions of the model’s learned image distribution. Existing white-box defenses require internal access or retraining and therefore cannot protect proprietary systems, while black-box LLM rewriting methods often leave these distribution-driven failures untouched.

The method operates entirely at the prompt level as a plug-and-play module. It constructs model-specific reference pools of safe and unsafe images by sampling the target generator on the I2P dataset and filtering outputs with classifier consensus. From an incoming prompt, DiSCO then performs distribution-guided suffix expansion through beam search in CLIP embedding space, scoring candidate suffixes by their contrastive alignment with the safe versus unsafe pools. An iterative feedback loop adjusts the optimization objective according to the remaining severity of harmful content until the generated image falls within the safe region.

Evaluations across 32 system–attack combinations and five random seeds show consistent reductions in attack success rate on the I2P benchmark, both for undefended models and for models already equipped with prior defenses. The approach preserves semantic fidelity and perceptual quality while remaining strictly black-box and architecture-agnostic. By formalizing the benign adversarial problem and demonstrating that prompt-level distributional steering can close the gap left by text-only sanitizers, the work supplies a practical, training-free layer that can be added to any text-to-image pipeline.

Why it matters

Directly addresses ethical AI safety and regulatory compliance needs in the Netherlands/EU; the plug-and-play black-box design is immediately actionable for Dutch SMEs and researchers working on generative models.

More in this beat
DiSCOI2PNSFW contentred-teamingsafety-alignmenttext-to-image
Stability AI’s Annual Integrity Transparency Report

02:00 · September 17, 2025

Stability AI’s Annual Integrity Transparency Report

This report is highly relevant for Dutch product teams and builders as it details the safety mechanisms, API filters, and C2PA provenance standards implemented in Stability AI models. Understanding these safeguards is crucial for building compliant, ethical AI applications that align with stringent EU and Dutch regulations.

Relevance 85 · Audience 80

Army Cyber training AI agents in cyber ‘work roles’ alongside human counterparts

15:57 · August 20, 2026

Army Cyber training AI agents in cyber ‘work roles’ alongside human counterparts

This article provides critical insights into how a leading NATO ally is operationalizing agentic AI in cyber warfare, directly informing Dutch and European doctrine developers and defense technologists. It highlights practical human-machine teaming models and ethical guardrails that align with the Netherlands' focus on responsible military AI.

Relevance 85 · Audience 95

FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

06:00 · August 20, 2026

FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

This research is highly relevant for Dutch AI researchers and the strong local fintech and banking sector exploring customer-facing LLM agents. It provides a rigorous, reproducible framework to test agent compliance and security against fraud, aligning with strict EU financial and AI regulations.

Relevance 85 · Audience 95

Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese

06:00 · August 15, 2026

Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese

This study is highly relevant for Dutch AI researchers and policymakers focused on ethical AI and EU AI Act compliance, as it demonstrates that safety guardrails can behave unpredictably across different languages. It underscores the necessity for multilingual safety evaluations, which is critical for Dutch enterprises deploying LLMs.

Relevance 85 · Audience 95

WorldClaw: Agentic 3D Open-World Generation at Scale

06:00 · August 7, 2026

WorldClaw: Agentic 3D Open-World Generation at Scale

This research is highly relevant for Dutch AI researchers and practitioners in the creative industries, gaming (e.g., Guerrilla Games), and digital twin sectors. It provides a novel, scalable approach to 3D environment generation using LLM agents and foundation models, offering actionable methodologies for advanced simulation development.

Relevance 85 · Audience 95

Auto mode is now the default in Claude Code for Pro, Max, and Team plans

02:00 · August 7, 2026

Auto mode is now the default in Claude Code for Pro, Max, and Team plans

Provides actionable implementation details, safety data, and configuration steps for an AI coding tool update directly usable by product teams and builders. Addresses workflow automation, risk mitigation, and observability in long-running AI tasks with specific model references.

Relevance 85 · Audience 90

When Vibe Hacking Turns AI into the Junior Hacker Every Adversary Always Wanted

13:30 · August 4, 2026

When Vibe Hacking Turns AI into the Junior Hacker Every Adversary Always Wanted

This article is highly relevant for security professionals as it highlights the evolving AI-driven threat landscape where the technical barrier to entry for attackers is significantly lowered. It provides actionable strategic advice on shifting from point-in-time security assessments to continuous threat exposure management, which is crucial for Dutch enterprises defending against AI-assisted cyberattacks.

Relevance 85 · Audience 90

Horizon3 Secures $250 Million to Lead AI-Versus-AI Cyber Defense

22:21 · August 3, 2026

Horizon3 Secures $250 Million to Lead AI-Versus-AI Cyber Defense

Article covers dual-use AI cyber defense with military/security relevance and EMEA growth plans that include the Netherlands; directly addresses AI/ML applications in proactive defense for industry and government users.

Relevance 68 · Audience 72