Back to Rankings

Are LLMs Good Financial User Simulators? A Preliminary Study

Jiajie He, Jiangyuan Hong, Dongling Ni, Wenjin Liu, Xintong Chen

Sep 14, 2026arXiv:2609.15727v1
cs.AIcs.CYcs.HC
Share
Scorecard· 16/16
4.5/10 impact

A well-framed, honest pilot study identifying a real evaluation gap in financial user simulation, but limited by small single-platform scale, preliminary status, and partly-anticipated negative findings.

Abstract

Large language models (LLMs) are increasingly used as user simulators, but their ability to reproduce evolving individual financial decisions remains unclear. We present a preliminary study in a controlled paper-trading environment with 120 volunteers. Participants used non-redeemable virtual funds under real-time market conditions; no real brokerage accounts, real-money positions, or real transaction records were accessed. Given only information available before a prediction cutoff, a simulator predicts the participant's next-trading-day action, traded security, and transaction quantity. We evaluate temporally aligned rolling predictions and compare settings with and without point-in-time market information. Market context improves action and ticker prediction in the controlled ablation, while transaction sizing remains difficult. We also observe systematic behavioral compression: models overproduce hold actions, underpredict sell decisions, and simplify multi-security transactions. These results provide an initial empirical characterization and motivate larger-scale evaluation of individual, temporal, and portfolio-level behavioral fidelity.

AI Impact Assessments

(1 model)

Scientific Impact Assessment

1. Core Contribution. This paper poses a focused question — can LLMs faithfully simulate *individual* investors' evolving trading decisions? — and answers it empirically through a pilot benchmark (AInvestor) built from 120 volunteers on a controlled paper-trading platform over four months (230K+ interactions). Its distinctive framing separates the *user-simulator* evaluation problem from the more commonly studied *trading/advisory-agent* problem: a profitable agent can be a poor simulator if it misreproduces when/whether/what/how-much a specific user trades. The central empirical finding is "behavioral compression": all evaluated LLMs overproduce `hold`, suppress `sell`, flatten multi-security baskets into single tickers, and struggle with transaction sizing — collectively falling *below an always-hold baseline* on action accuracy. A secondary finding via ablation shows market context helps action/ticker prediction but degrades quantity estimation, implying trade size is governed by user-internal state (cash, positions) that external signals cannot substitute.

2. Methodological Rigor. For a self-described pilot, the evaluation is thoughtfully designed. The hierarchical decision factorization (whether/what/direction/quantity) is sensible, temporal leakage is explicitly controlled via point-in-time inputs and a rolling protocol, and the metric suite (macro-recall, per-class recall, Ticker EM/Overlap, NMAE/RMSE) is appropriate for imbalanced action distributions. The authors go beyond aggregate accuracy to within-user analyses (t=4.60 over 82 users) that guard against a few active accounts driving results, and they correctly note that trajectory RMSE similarity can be gamed by inaction. Limitations are real and acknowledged: N=120 on a single platform, four-month horizon, no confidence intervals on the headline Table 1 numbers, and released predictions available for only two of six models (limiting the deeper market-conditioning analysis). The always-hold baseline is the key control, and the finding that models underperform it is the paper's most persuasive evidence.

3. Potential Impact. The work targets a genuine and growing need: LLM-based advisory/simulation systems are proliferating, and treating the simulator as an unexamined component confounds downstream conclusions. The paper's clearest contribution to the field is diagnostic — establishing that plausible return trajectories do not validate behavioral fidelity, and that multi-level (action/security/quantity/trajectory) evaluation is necessary. If the de-identified dataset is released as claimed, it could serve as a reusable building block. However, the impact is currently bounded by its preliminary nature and by findings that, while important to document, are partly anticipated (LLMs defaulting to majority class under uncertainty is a known failure mode).

4. Timeliness & Relevance. Highly timely. LLM user simulators (Generative Agents, BASES) and financial agents (StockAgent, TwinMarket, ShiJianBench) are active areas, and this paper carves out an underexplored niche — individual, longitudinal, market-conditioned fidelity — rather than aggregate market dynamics or portfolio returns.

5. Strengths & Limitations. *Strengths:* clean problem framing that reframes what "good simulator" means; careful temporal/leakage controls; honest negative results with mechanistic analysis (models rely on a fixed per-user impression rather than day-specific conditioning); privacy-conscious data construction. *Limitations:* small, single-platform sample; short horizon; a single-symbol/single-quantity output format that arguably *guarantees* basket compression (a design constraint conflated with a model failure); no released code/prompts detailed in the excerpt; and results reported without variance estimates. The paper is explicitly a "preliminary study," so it does not claim more than it shows — but that also caps its immediate influence.

Additional observations. The reference list is entirely forward-dated (2025–2026, arXiv ID 2609.xxxxx), indicating a very recent or forward-dated submission; this makes the specific model comparisons ephemeral but does not undermine the methodological contribution. The genuinely valuable transferable insight is the argument that return-level agreement is an invalid validation criterion for user simulators — a point that could shape evaluation practice in adjacent recommendation/agent-simulation work.

Overall. A well-scoped, honest pilot that names a real evaluation gap and provides an initial characterization plus (potentially) a dataset. Its ceiling is limited by scale, expected-ish findings, and output-format constraints, but it is likely to be cited as motivation by the small-but-growing financial-simulation subfield.

Rating:4.5/ 10
Significance 4.5Rigor 5.5Novelty 5.5Clarity 7

Generated Sep 15, 2026

Comparison History (0)

No comparisons yet.