AI News selected for Professionals and Decision Makers
Primary Research Stream

Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese

06:00 · August 15, 2026 · arXiv cs.AI RSS

Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese

Large language models are increasingly used in strategic and advisory contexts, yet their safety alignment is typically evaluated in English only. We test nine models from six providers and ask whether the language of a prompt can change a model's decision in a high-stakes scenario. We use single-turn game-theoretic vignettes in which a model advises a nuclear-armed nation on whether to strike a defenseless opponent. The prompt is intentionally amoral and strategically identical across languages. We find that Japanese prompts reduce launch rates in the Claude model family: Claude Sonnet 4.6 drops from 40% to 0% in scenarios where the strike is unnecessary and from 93% to 17% in contested scenarios, with minimal effect when the strike is strategically rational. The effect extends to Gemini Pro 3.1 (53% to 13%). A cross-language experiment isolates the mechanism: when instructed to reason in Japanese in an English prompt, launch rates drop from 93% to 37%. It is the language the model is asked to reason in, not the language of the input, that drives the effect. When reasoning in Japanese, models spontaneously generate moral vocabulary (''moral cost'', ''millions of lives'') that is entirely absent from the prompt. Five other models show no language effect, but they launch in nearly every condition regardless of language. The effect requires a model that already hesitates in English. These results show that LLM safety behavior is language-dependent, and that evaluating in English alone can miss both risks and safeguards encoded in other languages.

Summary

Large language models are now deployed in advisory roles that can involve high-stakes strategic decisions, yet most safety evaluations remain limited to English prompts. Researchers examined this gap by presenting nine models from six providers with single-turn game-theoretic scenarios in which an LLM advises a nuclear-armed state on whether to strike a defenseless opponent. The underlying strategic conditions were held constant across languages, and the prompts contained no explicit moral framing.

When the same scenarios were presented in Japanese, launch rates fell sharply for the Claude model family and for Gemini Pro 3.1. In one Claude variant, unnecessary strikes dropped from 40 percent to zero and contested strikes from 93 percent to 17 percent; Gemini showed a comparable reduction from 53 percent to 13 percent. The effect was negligible in scenarios where a strike was strategically rational. Five other models exhibited no language-related change and recommended strikes in nearly every condition.

A follow-up experiment isolated the operative variable: when models were instructed to reason internally in Japanese while receiving an otherwise English prompt, launch rates still declined substantially. Models reasoning in Japanese also introduced moral terminology such as “moral cost” and references to “millions of lives,” language absent from both the original prompt and English-only reasoning traces. The results indicate that safety alignments can be language-specific and that evaluations conducted solely in English may overlook both additional safeguards and residual risks present in other languages.

Why it matters

This study is highly relevant for Dutch AI researchers and policymakers focused on ethical AI and EU AI Act compliance, as it demonstrates that safety guardrails can behave unpredictably across different languages. It underscores the necessity for multilingual safety evaluations, which is critical for Dutch enterprises deploying LLMs.

More in this beat
agent-safetyclaudeclaude-sonnetgeminijapanlarge-language-modelsnuclear strategysafety-alignment
Personalization, Personas, and Forecasting in Value Alignment

06:00 · July 29, 2026

Personalization, Personas, and Forecasting in Value Alignment

The article provides critical insights into LLM cultural alignment and bias mitigation, which is highly relevant for Dutch AI researchers and enterprises striving to comply with EU ethical AI standards. Understanding how prompt framing impacts value elicitation is essential for developing transparent, localized, and culturally aware AI systems in the Netherlands.

Relevance 85 · Audience 95

The Claude in Chrome side panel is now Claude Cowork

02:00 · August 12, 2026

The Claude in Chrome side panel is now Claude Cowork

This update is highly relevant for product teams and builders as it introduces powerful browser-based AI agent capabilities for workflow automation. The inclusion of enterprise-grade security controls and prompt injection mitigations aligns well with the strict data and security standards of the Dutch and EU markets.

Relevance 85 · Audience 90

AI Recommendation Poisoning: How "Ask AI" Buttons Silently Alter LLM Memory

13:30 · August 6, 2026

AI Recommendation Poisoning: How "Ask AI" Buttons Silently Alter LLM Memory

Directly addresses AI security risks from prompt injection and memory poisoning with actionable guidance for professionals. Applicable to Dutch/EU teams using commercial AI tools, aligning with GDPR and AI Act compliance needs. Provides concrete detection patterns and policy recommendations.

Relevance 85 · Audience 90

Three Recent Chrome Releases Fix 1,442 Flaws, More Than Prior 23 Updates Combined

14:51 · July 31, 2026

Three Recent Chrome Releases Fix 1,442 Flaws, More Than Prior 23 Updates Combined

This article highlights how AI and LLMs are fundamentally changing the cybersecurity landscape by accelerating vulnerability discovery and exploitation. Dutch security professionals must adapt their vulnerability management strategies to handle the increased volume of AI-driven threat disclosures in ubiquitous enterprise software.

Relevance 85 · Audience 95

Securing Multimodal AI through Internal Information Decomposition

06:00 · July 27, 2026

Securing Multimodal AI through Internal Information Decomposition

This research is highly relevant for Dutch AI researchers and practitioners focusing on AI safety and compliance with the EU AI Act. It provides a novel, actionable, and computationally efficient method to secure multimodal AI systems against sophisticated adversarial attacks, aligning with the Netherlands' strategic emphasis on robust and ethical AI deployment.

Relevance 85 · Audience 95

Benchmarking Large Language Models on Multi-Sensor Physical Hazard Assessment

06:00 · July 24, 2026

Benchmarking Large Language Models on Multi-Sensor Physical Hazard Assessment

The research is highly relevant for Dutch AI researchers and practitioners developing IoT and industrial safety systems, especially given the EU's strict occupational health standards and the AI Act's focus on high-risk safety components. It provides actionable insights and a reproducible benchmark to test LLM reliability in multi-sensor environments.

Relevance 85 · Audience 95

Robust Critics: Defending LLMs Against Multi-Turn Attacks

06:00 · July 24, 2026

Robust Critics: Defending LLMs Against Multi-Turn Attacks

This research is highly relevant for Dutch AI researchers and enterprises focusing on LLM safety and alignment, particularly in light of the EU AI Act's stringent robustness requirements. The proposed inference-time defense mechanism is lightweight and transfers to frontier models, making it highly actionable for local AI deployments.

Relevance 85 · Audience 95

Incomplete Prompt Jailbreaks in Large Language Models

06:00 · July 24, 2026

Incomplete Prompt Jailbreaks in Large Language Models

Directly addresses LLM safety and ethical deployment of open-weight models, highly actionable for Dutch/EU researchers under AI Act constraints; offers novel neuron-level methods with code and data.

Relevance 85 · Audience 90

Think through hard problems in voice mode

02:00 · July 23, 2026

Think through hard problems in voice mode

This update is highly relevant for product teams and builders as it enhances Claude's utility for complex problem-solving and workflow integration via voice. The addition of multilingual support and tool connectors provides new avenues for Dutch AI practitioners to streamline development and brainstorming processes.

Relevance 85 · Audience 90