AI News selected for Professionals and Decision Makers
Primary Research Stream

CreativityNeuro: Steering Language Model Weights to Improve Divergent Thinking and Reduce Mode Collapse

06:00 · July 3, 2026 · arXiv cs.AI RSS

CreativityNeuro: Steering Language Model Weights to Improve Divergent Thinking and Reduce Mode Collapse

Divergent thinking is a crucial aspect of creativity, yet large language models (LLMs) tend to consistently generate similar responses to open-ended questions, in what has been termed the artificial hivemind effect. Here, we introduce CreativityNeuro, a data-free method for enhancing divergent thinking in LLMs via contrastive weight steering. We evaluate our method across multiple creativity assessments and report several main findings. On the Divergent Association Task (DAT), a vocabulary-space creativity test, CreativityNeuro improves performance by up to 14 human percentile points. Next, in a large-scale human evaluation (N=720) on the Alternative Uses Test (AUT) and the Task Task, CreativityNeuro achieves significant improvements in originality, surprise, and creativity, transferring to longer-form and more open-ended tasks. Importantly, we find that across all three tasks, CreativityNeuro demonstrably reduces measures of mode collapse. Moreover, activation steering achieves comparable performance to CreativityNeuro on the DAT, but it does not transfer to the AUT and Task Task, demonstrating the effectiveness of weight-space steering in generalizing to unseen tasks. In conclusion, CreativityNeuro improves divergent thinking and reduces mode collapse without requiring behavioral data, re-training, or gradient-based fine-tuning, providing a straightforward way to enhance LLM performance in creative domains.

Summary

CreativityNeuro is a data-free technique that steers the weights of large language models to increase divergent thinking while curbing the artificial hivemind effect, in which models repeatedly produce similar answers to open-ended prompts. The approach relies on contrastive prompt sets—one set encouraging creative responses and the other favoring conventional ones—to compute per-weight importance scores across layers. It then isolates a sparse subset of creativity-specific parameters by subtracting those most active under non-creative prompts and applies a scaled multiplicative perturbation to those weights only.

Evaluations show consistent gains. On the Divergent Association Task, a vocabulary-based test of remote associations, the method lifts model performance by as much as 14 human percentile points. In a 720-participant human study using the Alternative Uses Test and the more open-ended Task Task, responses generated after steering received higher ratings for originality, surprise, and overall creativity. The same interventions measurably reduced indicators of mode collapse across all three tasks.

Unlike activation steering, which matches CreativityNeuro on the DAT but fails to transfer to longer-form tasks, weight-space adjustments generalize to unseen prompt distributions. The procedure requires no behavioral datasets, gradient updates, or retraining, distinguishing it from prompting frameworks, temperature tuning, and reinforcement-learning approaches that depend on labeled preference data.

Why it matters

This research provides Dutch AI researchers and practitioners with a resource-efficient method to improve LLM creativity and mitigate mode collapse. Its data-free, weight-steering approach aligns with the Netherlands' focus on sustainable, controllable, and innovative AI development.

More in this beat
activation-steeringCreativityNeuroDivergent Thinkingevaluation-benchmarkslarge-language-modelsmode-collapsenovel-methodologies
Woodpecker Distillation: Weak Models Diagnose Reasoning Bugs in Strong Models

06:00 · August 7, 2026

Woodpecker Distillation: Weak Models Diagnose Reasoning Bugs in Strong Models

This paper is highly relevant for AI researchers in the Netherlands focusing on LLM reasoning, alignment, and compute-efficient training. The proposed weak-to-strong distillation method offers actionable insights for Dutch AI labs aiming to enhance model performance without relying solely on massive scaling.

Relevance 85 · Audience 95

Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks

06:00 · July 14, 2026

Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks

This research is highly relevant for Dutch AI researchers and enterprises developing LLMs, as it offers a mathematically rigorous method to drastically reduce the computational cost and time required for model evaluation. This aligns with the European and Dutch focus on sustainable, resource-efficient AI development (Green AI).

Relevance 85 · Audience 95

L-MAD: A Systematic Evaluation of Multi-Agent Debate Structures in Legal Reasoning

06:00 · July 13, 2026

L-MAD: A Systematic Evaluation of Multi-Agent Debate Structures in Legal Reasoning

This research is highly relevant for Dutch AI researchers and LegalTech developers building multi-agent systems for high-stakes, regulatory, or compliance domains. It provides actionable insights into preventing hallucination and over-deliberation, aligning with the Netherlands' strong focus on transparent, ethical, and reliable AI.

Relevance 85 · Audience 95

Synthetic Consumer Insight Generation with Large Language Models

06:00 · July 8, 2026

Synthetic Consumer Insight Generation with Large Language Models

This article is highly relevant for researchers and advanced readers in the Dutch AI market as it addresses the growing need for synthetic data generation, which is crucial for navigating strict EU GDPR privacy regulations. The methodological insights into prompt engineering and model evaluation provide valuable frameworks for Dutch AI practitioners in marketing and consumer analytics.

Relevance 85 · Audience 95

Investigating Multi-Agent Deliberation in Law

06:00 · July 1, 2026

Investigating Multi-Agent Deliberation in Law

This research is highly relevant for Dutch AI researchers and legal tech practitioners, as it introduces novel multi-agent frameworks for legal reasoning. Given the Netherlands' strong emphasis on ethical AI and transparent legal applications, these law-inspired deliberation models offer actionable methodologies for developing robust AI systems in regulated domains.

Relevance 85 · Audience 95

BayesBench: Evaluating LLM Belief Trajectories Under Multi-Turn Evidence Accumulation

06:00 · July 1, 2026

BayesBench: Evaluating LLM Belief Trajectories Under Multi-Turn Evidence Accumulation

This research provides a rigorous framework for evaluating the reasoning and reliability of LLMs in dynamic, multi-turn environments. For Dutch AI researchers and developers, understanding and benchmarking these epistemic updates is crucial for building trustworthy, transparent AI systems that align with EU standards.

Relevance 85 · Audience 95

RoPoLL: Robust Panel of LLM Judges

06:00 · July 1, 2026

RoPoLL: Robust Panel of LLM Judges

Directly actionable for Dutch research teams and SMEs building LLM evaluation pipelines; aligns with EU emphasis on reliable and transparent AI; high technical depth and novelty for advanced readers.

Relevance 72 · Audience 88

NormAct: A Benchmark for Hidden Social Norm Compliance in Embodied Planning

06:00 · June 29, 2026

NormAct: A Benchmark for Hidden Social Norm Compliance in Embodied Planning

This research is highly relevant to the Dutch AI market's strong emphasis on ethical, transparent, and socially responsible AI. The benchmark provides Dutch researchers and enterprises with actionable tools to evaluate and improve the social compliance of embodied AI agents, aligning with EU regulatory frameworks for safe AI deployment.

Relevance 85 · Audience 95

PEAR: Permutation-Equivariant Adaptive Routing Multi-Agent Debate

06:00 · June 23, 2026

PEAR: Permutation-Equivariant Adaptive Routing Multi-Agent Debate

This research is highly relevant for Dutch AI researchers and advanced practitioners focusing on LLM reliability and multi-agent systems. The introduction of a dynamic, bias-reducing routing protocol aligns with the Netherlands' strategic emphasis on transparent, ethical, and robust AI development, offering actionable methodologies with open-source code.

Relevance 85 · Audience 95