Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction
06:00 · July 2, 2026 · arXiv cs.AI RSS

Most approaches to AI alignment treat human preferences as fixed targets to be inferred and optimized. This assumption conflicts with extensive empirical evidence showing that preferences are layered, dynamic, and constructed through interaction--particularly with adaptive technologies. As AI systems become more persistent, personalized, and socially embedded, they increasingly participate in shaping what people attend to, value, and endorse over time. We introduce Constructive Alignment, a paradigm that reframes alignment as a control problem over evolving human preference trajectories rather than static preference satisfaction. Drawing on behavioral economics, psychology, and constructivist social theory, we model preferences as layered state variables that evolve under interaction with AI systems. We formalize this view using a control-theoretic framework in which system actions and interaction design jointly influence both world states and human evaluative states. We argue that alignment is not primarily about controlling AI behavior, but about regulating how AI systems influence the evolution of human preferences--ensuring that value trajectories remain coherent, reflectively endorsed, epistemically grounded, bounded against manipulation, and empowering under uncertainty. Alignment thus becomes a problem of governing long-term value formation rather than simply satisfying static preferences.
Summary
Most current approaches to AI alignment treat human preferences as stable targets that systems should infer and optimize against. This assumption clashes with extensive evidence from psychology and behavioral economics that preferences are layered, context-dependent, and actively constructed through ongoing interaction, especially with adaptive technologies. As AI systems grow more persistent, personalized, and embedded in daily life, they participate in shaping what users attend to, value, and endorse over time, from recommendation-driven habit formation to shifts in attention and identity.
The paper introduces Constructive Alignment to address this dynamic. It reframes alignment not as the satisfaction of fixed preferences but as a control problem over evolving human preference trajectories. Drawing on constructivist accounts from psychology, decision research, and social theory, the authors model preferences as layered state variables—spanning immediate affective wants, instrumental goals, and longer-term identity commitments—that change under the influence of system actions and interaction design. A control-theoretic formulation captures how these actions jointly affect external world states and internal evaluative states.
The resulting framework shifts the focus from controlling AI behavior alone to regulating the processes through which AI influences long-term value formation. Alignment objectives therefore center on keeping preference trajectories coherent, reflectively endorsed by the individual, epistemically grounded, protected from undue manipulation, and supportive of agency under uncertainty. This view treats sustained human-AI interaction as an unavoidable site of value construction rather than a neutral channel for preference expression.
Why it matters
This research aligns perfectly with the Dutch AI market's strong strategic focus on ethical, transparent, and human-centric AI. It provides advanced researchers with a rigorous, control-theoretic framework to address the long-term societal impacts and potential manipulative risks of adaptive AI systems.




