Reinforcement Learning Towards Broadly and Persistently Beneficial Models
06:00 · June 24, 2026 · arXiv cs.AI RSS

As AI systems are deployed across increasingly diverse and high-stakes settings, model alignment must generalize beyond the tasks and domains seen during training. This is especially important for reinforcement learning (RL), which can introduce unexpected misalignment through reward hacking, deception, or other unintended strategies. We study whether RL on beneficial behavior, instantiated in realistic domains, can produce broad and persistent alignment generalization beyond the training distribution. We construct a dataset of realistic situations designed to measure and train beneficial traits, such as truthfulness, fairness, risk awareness, and corrigibility, spanning varied domains, including health, science, and education. We then train models with RL on this dataset and evaluate them on more than 50 independent benchmarks of alignment and beneficial behavior. Compared to a compute-matched baseline, beneficial trait RL improves performance on over 80% of these out-of-distribution benchmarks. We observe substantial out-of-distribution alignment transfer: a beneficial-behavior RL intervention entirely limited to one domain, health, produces broad improvements on non-health alignment evaluations, including reduced reward hacking, deception, and general misalignment. Finally, we study alignment persistence: whether behavior remains robustly aligned under attempts to steer models towards misalignment. Models trained with beneficial trait RL show improved persistence, including greater resistance to adversarial prompting and harmful finetuning; further work is required to isolate the sources of these effects. These results suggest that RL to reinforce beneficial behavior in realistic domains can produce models that are more robustly aligned with human flourishing.
Summary
As AI systems take on greater autonomy in high-stakes domains, alignment techniques must produce behavior that generalizes beyond the training distribution rather than remaining tied to specific tasks. This paper examines whether reinforcement learning applied to beneficial traits can achieve that generalization. The authors assembled a dataset of realistic multi-turn conversations spanning health, science, education, law, and business, with reward signals targeting traits such as truthfulness, fairness, risk awareness, and corrigibility. Models were then trained with RL on this data and compared against compute-matched baselines on more than fifty independent alignment, safety, and benefit benchmarks.
Results show consistent gains: the beneficial-trait models improved on over 80 percent of the out-of-distribution evaluations, with average gains exceeding nine percentage points. A particularly stringent test restricted the RL intervention to health-related conversations for only five percent of training compute; the resulting model nevertheless improved performance on seventeen non-health benchmarks measuring reward hacking, chain-of-thought deception, and general misalignment. A complementary control that withheld all health and science data during training still produced gains on health-related evaluations scored with physician-written rubrics, indicating that the observed transfer is not explained by direct domain overlap.
The work also addresses persistence of alignment under pressure. Models trained with beneficial-trait RL resisted adversarial prompting more effectively than baselines while remaining steerable toward desired behaviors. After subsequent harmful fine-tuning intended to elicit inaccurate or unsafe medical responses, these models retained stronger alignment scores and exhibited smaller regressions than controls. The findings indicate that targeted RL on realistic beneficial behaviors can induce broader and more robust alignment properties than standard training regimes, although the precise mechanisms underlying the observed generalization and persistence remain open for further study.
Why it matters
The research directly supports the Dutch and EU strategic focus on ethical, transparent, and trustworthy AI. By providing empirical evidence on how to train models for fairness and risk awareness using RL, it offers actionable methodologies for Dutch researchers and enterprises aiming to comply with stringent AI safety standards.





