Nicole Mitchell, Dhruv Agarwal, Maty Bohacek, Remi Denton, Roma Patel
A timely, well-synthesized agenda-setting position paper that crystallizes an emerging safety concern and could become a framing reference, but offers no empirical artifact and depends on others to execute its roadmap.
Language models have taken on the role of a very new type of technology, by virtue of their "human-ness" and rapid integration into users' daily lives. This combination of features can introduce longitudinal risks---cognitive, developmental and socio-affective changes in humans---that might not surface in short-term interactions, but can have lasting long-term effects on users. This forms the basis of a critical new mission for NLP: to pivot from static, short-term evaluations of text generations to long-term measurements of behavioral changes, towards a diachronic understanding of human-model interactions. In this work, we draw from measurements used in social science fields that are crucial to understand emergent phenomena in longitudinal data. We discuss how computational methods in the field of NLP need to be combined with such measurements, not only to understand long-term safety risks of human-model interactions, but to help steer model development towards positive rather than negative outcomes for users. This ability to model human behavioral shifts as a function of model interactions can facilitate online rather than post-hoc detection of problematic behaviors, and should be leveraged in alignment frameworks to mitigate long-term risks in users.
This is a position/agenda-setting paper that argues NLP should pivot from static, short-term evaluations of model outputs toward *longitudinal* measurement of how sustained human-AI interactions change users cognitively, socio-affectively, epistemically, and clinically. Its central conceptual move is reframing AI safety from a *synchronic* problem (harms visible at a single point in time) to a *diachronic* one (harms that accumulate or only surface over months of interaction). The paper's concrete deliverables are: (1) a taxonomy of longitudinal risks (Table 2), (2) a survey/mapping of validated behavioral-science measurement scales (psychometric, psychosocial, cognitive, value, well-being, utility) and how they might be "computationalized" for NLP, (3) a comparative analysis of three data-gathering paradigms (RCTs, field studies, user simulations) with their trade-offs (Table 1), and (4) a roadmap for integrating computational social science, cognitive models (RSA), and dynamic-systems methods into alignment pipelines. It solves no technical problem directly; rather it defines a research program and imports an interdisciplinary vocabulary.
As a position paper, there are no experiments or proofs to evaluate. The rigor lies in the quality of synthesis and argumentation, which is high. The taxonomy is well-organized and each risk category is grounded in cited empirical work (Fang et al. RCT, Cheng et al. Science paper on sycophancy, Moore et al. on delusional spirals). The Table 1 comparison of data paradigms is analytically sound, explicitly noting causal-control vs. ecological-validity vs. scalability trade-offs. Importantly, the authors demonstrate self-awareness of pitfalls—they devote real space to construct validity, psychometric bias, western-centrism of scales, and the distinction between applying scales to *humans* (their proposal) vs. to *models* (which they critique). This nuance strengthens credibility. However, the paper offers no proof-of-concept: no pilot measurement, no demonstration that any proposed metric actually detects a longitudinal shift. The proposals for integrating trajectory-level metrics into RL rewards remain speculative sketches.
The topic is highly timely. Concerns about AI companionship, emotional dependence, "AI psychosis," sycophancy, and cognitive deskilling have surged in both academic venues (CHI, Science) and public discourse. This paper is well-positioned to become a frequently-cited *framing reference* for the emerging subfield of longitudinal human-AI interaction safety. Its main value is orientational: giving researchers a shared taxonomy, a curated bridge to the behavioral-sciences measurement literature, and a vocabulary (synchronic/diachronic, longitudinal quiddity). Position papers of this type can accumulate substantial citations when they crystallize a nascent concern at the right moment. That said, its influence depends on others doing the hard empirical work it merely gestures at; the paper itself does not lower a technical barrier or provide a reusable artifact (dataset, benchmark, code).
Extremely timely. The paper explicitly capitalizes on the 3–4 year window since ChatGPT during which longitudinal effects have become observable. Its heavy reliance on very recent (2025–2026) citations shows it is riding the crest of an active wave. It addresses a genuine gap: most alignment/safety evaluation is single-session, and the field lacks frameworks for time-extended harm.
Strengths: (1) Excellent interdisciplinary synthesis bridging NLP, psychology, HCI, cognitive science, and dynamic-systems theory. (2) A clear, memorable conceptual framing (synchronic→diachronic). (3) Genuinely useful taxonomy and paradigm-comparison tables. (4) Intellectual honesty about limitations (construct validity, privacy, cultural bias, private-sector data control). (5) Comprehensive, current bibliography that itself serves as a resource.
Limitations: (1) No original empirical contribution—no data, metric validation, or demonstration. The paper is entirely prescriptive. (2) Many proposals (RL reward integration of psychometric features, dynamic-systems inflection detection on text) are underspecified and may face severe measurement-validity and identifiability challenges the authors acknowledge but don't resolve. (3) The data-access bottleneck (most longitudinal data is corporate-owned) is a fundamental barrier the paper cannot overcome. (4) Risk of being one of several concurrent "we should study long-term effects" position pieces, reducing distinctiveness. (5) Ethical/privacy tensions in continuous behavioral monitoring are noted but not resolved, and could limit real-world adoption of the proposed monitoring infrastructure.
The paper contributes no dataset or tool, so reproducibility in the empirical sense does not apply. Its foundational value is moderate—it may become a reference for the *framing* rather than a reusable *primitive*. The interdisciplinary reach is genuine and a real strength. The core ideas, while individually not novel (each risk is drawn from existing work), are novel in *synthesis and framing*—a well-read expert would recognize most components but would not necessarily have assembled them into this coherent longitudinal measurement program. Overall this is a solid, timely agenda-setting paper whose ultimate impact hinges on whether the community takes up its call with the empirical work it cannot itself provide.
Generated Aug 4, 2026
A timely, well-synthesized agenda-setting position paper that crystallizes an emerging safety concern and could become a framing reference, but offers no empirical artifact and depends on others to execute its roadmap.