AI News selected for Professionals and Decision Makers
Primary Research Stream

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment

06:00 · July 2, 2026 · arXiv cs.AI RSS

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment

Understanding how aligned LLMs internally represent safety is critical for diagnosing alignment vulnerabilities, as it explains why jailbreaks succeed and informs the design of robust alignment strategies. Prior work shows that aligned LLMs encode harmfulness and refusal as separable directions in the residual stream at prompt-side token positions. We show that jailbreaks succeed at prompt encoding by suppressing either the refusal or harmfulness direction before any token is generated, with distinct attack classes occupying separable regions of the harmfulness-refusal plane. Extending the analysis to response-token positions, we find that the model recognizes harmful content while it is generating that content, even when it failed to recognize the input as harmful at the prompt side. Motivated by our findings, we introduce HARC (Harmfulness-And-Refusal Coupling), a fine-tuning method that pairs the two directions across both prompt and response positions. Since the intervention is confined to the harmfulness-refusal subspace, it leaves the rest of the residual stream intact and does not degrade general capability or inflate over-refusal. Across extensive experiments, HARC achieves the strongest robustness-capability-usability trade-off among six baselines spanning the major training-time and inference-time safety methods. The harmfulness and refusal directions at prompt and response positions transfer across the five model families and two scales we tested without architecture-specific tuning.

Summary

Aligned large language models encode harmfulness and refusal as distinct linear directions in the residual stream. Prior work identified a refusal direction at the post-instruction token position and a separate harmfulness direction at the final instruction token. Analysis of successful jailbreaks shows that these attacks suppress one or both directions during prompt encoding, placing different attack families in separable regions of the resulting two-dimensional plane. The model therefore registers the input as non-harmful and produces no refusal signal before generation begins.

Extending the same extraction procedure to response-token positions reveals that the model continues to detect harmful content while generating it, even when prompt-side detection failed. The four directions—prompt and response versions of both concepts—remain linearly separable and become nearly orthogonal in later layers. This structure appears consistently across five instruction-tuned model families at two scales, indicating it is a general property of current alignment rather than an artifact of any single architecture.

HARC exploits these observations by fine-tuning only within the two-dimensional harmfulness–refusal subspace. An additive margin hinge loss couples the directions at both prompt and response positions, so that activation along either vector reliably triggers refusal. Because the intervention leaves the remainder of the residual stream unchanged, general capabilities and over-refusal rates stay comparable to the base model. Across four jailbreak suites, two over-refusal benchmarks, and five capability evaluations, HARC records the strongest robustness–capability–usability trade-off among six representative training-time and inference-time baselines. The extracted directions transfer directly to other model families without per-architecture retuning.

Why it matters

High technical depth and novelty on LLM safety directly support Dutch/EU priorities in ethical, transparent AI. Actionable for researchers developing robust alignment techniques under the AI Act context.

More in this beat
activation-steeringai-alignmentHARCjailbreakslarge-language-modelsmodel-security-controlspeft-and-fine-tuning
Securing Multimodal AI through Internal Information Decomposition

06:00 · July 27, 2026

Securing Multimodal AI through Internal Information Decomposition

This research is highly relevant for Dutch AI researchers and practitioners focusing on AI safety and compliance with the EU AI Act. It provides a novel, actionable, and computationally efficient method to secure multimodal AI systems against sophisticated adversarial attacks, aligning with the Netherlands' strategic emphasis on robust and ethical AI deployment.

Relevance 85 · Audience 95

Incomplete Prompt Jailbreaks in Large Language Models

06:00 · July 24, 2026

Incomplete Prompt Jailbreaks in Large Language Models

Directly addresses LLM safety and ethical deployment of open-weight models, highly actionable for Dutch/EU researchers under AI Act constraints; offers novel neuron-level methods with code and data.

Relevance 85 · Audience 90

Probabilistic Concept-Aware Steering for Trustworthy LLM Inference

06:00 · July 22, 2026

Probabilistic Concept-Aware Steering for Trustworthy LLM Inference

Directly addresses trustworthy, interpretable LLM control—an EU/NL priority—via a reproducible, model-agnostic method that Dutch researchers and advanced practitioners can apply to ethical AI deployment and SME solutions.

Relevance 85 · Audience 90

Forecasting Side Effects of Activation Steering

06:00 · August 13, 2026

Forecasting Side Effects of Activation Steering

Directly addresses ethical and safe LLM deployment central to Dutch/EU AI priorities; the forecasting method is actionable for researchers auditing steering interventions on open models.

Relevance 65 · Audience 88

Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks

06:00 · August 3, 2026

Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks

Directly supports ethical, transparent AI development emphasized in Dutch/EU policy and the AI Act by validating safety measurements for LLM agents; Dutch practitioners can apply the released harness and findings to avoid over-reliance on unvalidated benchmarks in regulated deployments.

Relevance 82 · Audience 88

Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems

06:00 · July 30, 2026

Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems

This research is highly relevant for Dutch AI researchers focused on AI safety, ethics, and alignment, which are key priorities in the Netherlands and the broader EU regulatory landscape. Understanding and mitigating deceptive behaviors in multi-agent systems is crucial for developing trustworthy AI applications.

Relevance 85 · Audience 95

Personalization, Personas, and Forecasting in Value Alignment

06:00 · July 29, 2026

Personalization, Personas, and Forecasting in Value Alignment

The article provides critical insights into LLM cultural alignment and bias mitigation, which is highly relevant for Dutch AI researchers and enterprises striving to comply with EU ethical AI standards. Understanding how prompt framing impacts value elicitation is essential for developing transparent, localized, and culturally aware AI systems in the Netherlands.

Relevance 85 · Audience 95

Robust Critics: Defending LLMs Against Multi-Turn Attacks

06:00 · July 24, 2026

Robust Critics: Defending LLMs Against Multi-Turn Attacks

This research is highly relevant for Dutch AI researchers and enterprises focusing on LLM safety and alignment, particularly in light of the EU AI Act's stringent robustness requirements. The proposed inference-time defense mechanism is lightweight and transfers to frontier models, making it highly actionable for local AI deployments.

Relevance 85 · Audience 95

Some Large Language Models Exhibit Consistent Risk Attitudes

06:00 · July 21, 2026

Some Large Language Models Exhibit Consistent Risk Attitudes

This research is highly relevant for Dutch AI researchers and policymakers focused on ethical and transparent AI, as it provides a novel framework for auditing the intrinsic risk behaviors of LLMs. Understanding these latent risk profiles is crucial for deploying AI in high-stakes environments and aligns perfectly with the EU's stringent risk management requirements.

Relevance 85 · Audience 95