Towards Inclusive Mobility Modeling: Characterizing and Evaluating Elderly Trajectory Patterns in Urban Systems
06:00 · July 1, 2026 · arXiv cs.AI RSS

The rapid advance of smart cities increasingly depends on trajectory data mining, yet underrepresented demographic groups, particularly the elderly, are often sparsely represented in public mobility datasets. This underrepresentation can introduce systematic bias into mobility modeling and downstream urban planning. Using the 2016-2020 Jersey City subset of the Citi Bike System Data, this study quantitatively examines how the absence of underrepresented subgroups' mobility signatures affects mobility modeling, using synthetic trajectory generation as a case study. The analysis reveals that elderly riders exhibit a structurally distinct mobility signature, including localized activity spaces (958 m vs. 1,189 m for young riders), lower mobility entropy (1.82 vs. 4.15), and asymmetric off-peak temporal patterns. To demonstrate that relying on majority-dominated training data yields biased synthetic outcomes, we further evaluate both a first-order Markov chain and a Qwen3-4B model fine-tuned with QLoRA across three demographic training settings: the full population, young riders only, and elderly riders only. Results show that models trained on majority-dominated populations systematically misrepresent elderly mobility behavior, particularly for spatial mobility metrics. The Markov model trained on the full population overestimates elderly step length by 4.5% and dwell time by 8.9%, whereas the elderly-specific model achieves substantially lower errors across most metrics. Comparisons between the Markov and LLM-based frameworks further show that higher-capability models do not necessarily improve subgroup-level fidelity under limited demographic data. These findings underscore the importance of demographic representation in mobility modeling and its downstream applications for underrepresented populations.
Summary
This study examines how demographic underrepresentation in mobility datasets introduces systematic bias into trajectory modeling, with a focus on elderly riders in the 2016–2020 Jersey City Citi Bike records. Analysis of the data shows that riders aged 65 and older display distinct patterns compared with younger adults aged 18–35: their activity spaces are more localized, with an average radius of gyration of 958 m versus 1,189 m, their mobility entropy is markedly lower at 1.82 versus 4.15, and their temporal usage skews toward off-peak hours rather than commuting peaks.
To test the downstream effects of training on majority-dominated data, the authors compare two generative approaches—a first-order Markov chain and a Qwen3-4B model fine-tuned with QLoRA—under three controlled demographic regimes: the full population, young riders only, and elderly riders only. When trained on the full or young-dominated sets, both models overestimate key elderly mobility metrics; the Markov model, for instance, inflates average step length by 4.5 % and dwell time by 8.9 %. Models trained exclusively on elderly trajectories reduce these errors substantially across spatial and temporal dimensions.
The experiments further indicate that increased model capacity does not automatically correct subgroup-level distortions when demographic coverage remains limited. These results point to the need for explicit demographic stratification in mobility datasets if synthetic trajectory generation is to support equitable urban planning rather than reinforce existing representational gaps.
Why it matters
The paper aligns strongly with the Dutch AI market's focus on ethical, transparent AI and smart city innovation. Its insights into mitigating demographic bias in mobility models are highly actionable for Dutch researchers working on urban planning and cycling infrastructure.









