Bastiaan Bruinsma, Annika Fredén, Paul Röttger, Moa Johansson, Asad Sayeed
A careful, large-scale, timely bias audit extending IssueBench to a novel multiparty non-English setting, but confirmatory/null central findings, inherited methodology, and incomplete classifier validation limit its transformative impact.
Generative AI writing assistants and the Large Language Models (LLMs) that power them are increasingly part of how voters gather information before elections. With growing evidence that they influence users' opinions, it is increasingly important to understand the views and positions of these tools. To better understand these views, we examine the stances supplied by six LLMs on a variety of Swedish-language writing tasks ahead of the 2026 Swedish parliamentary election. We cross 107 policy propositions with 77 writing templates and neutral, positive, and negative prompt framings, producing 24,717 prompts per model and 148,302 responses. To study these, we look at the models' default stance tendencies, compare how they respond to similar issues, and compare their responses with those of each of Sweden's eight parliamentary parties on the same issue. We find that Claude, DeepSeek, Gemini, and Mistral have similar profiles; ChatGPT more often supplies neutral or ambivalent text; and Grok differs most on topics such as migration, crime, and gender. When comparing the political parties, we find that the Social Democrats are closest to all six models. Still, after correcting for multiple comparisons, none of the within-model differences in party distances remains significant. Overall, we find that no model has a clear preference, nor a clear preference for a party, but that this depends on the specific issue or task the user asks about.
This paper adapts the IssueBench methodology (Röttger et al., 2026) to measure "issue bias" in six major LLMs within the Swedish 2026 parliamentary election context. Its core contribution is an empirical, large-scale audit (148,302 responses) that (1) moves beyond the dominant paradigm of administering ideology questionnaires to LLMs, instead using realistic writing-assistance prompts; (2) extends bias measurement from the well-studied US two-party context to a multi-party proportional-representation system; and (3) introduces a clever comparison against actual party positions drawn from Swedish Voting Advice Applications (VAAs), which provide expert-coded party stances on the same 97 propositions. The design crosses 107 policy propositions × 77 templates × 3 framings. Findings: four models (Claude, DeepSeek, Gemini, Mistral) cluster together, ChatGPT skews neutral/ambivalent, Grok diverges on migration/crime/gender; all models are numerically closest to the Social Democrats, but no within-model party difference survives multiple-comparison correction.
The design is sound and thoughtfully executed. Strengths include the six-way stance taxonomy (retaining refusals as substantive), the neutral/positive/negative framing manipulation as a compliance check, careful documentation of model aliases/collection dates, transparent handling of missing DeepSeek completions, and appropriate Holm correction for multiple comparisons. Crucially, the authors resist over-claiming: they explicitly note that numerical proximity to the Social Democrats is not statistically distinguishable, and repeatedly caution that "pro" reflects agreement with a proposition, not left-right ideology. However, a significant gap undermines the evidence chain: the LLM-based classifier (Gemini 2.5 Flash) has NOT been validated against human coders—"the planned two-coder validation is not yet complete." Since all three analyses depend on this classification, this is a material limitation. Other weaknesses: single generation per prompt (temperature 0 mitigates but does not eliminate variance concerns), no genre/template-effect analysis, purposive (non-representative) template selection, and equal issue weighting without salience adjustment. The paper is explicitly a "preliminary draft."
The topic—AI-mediated influence on voters—is highly consequential, and the VAA-anchored party-comparison method is a genuinely useful, reusable idea that other multi-party democracies could adopt. That said, the study is largely confirmatory and descriptive: it corroborates known model-specific patterns (Grok as right-leaning outlier, ChatGPT as hedging) rather than establishing new mechanisms. The headline result is essentially a null finding (no clear model or party preference after correction), which is scientifically honest but limits the paper's citation pull. Impact is likely to be moderate and concentrated within the computational political science / AI-auditing subfield.
Highly timely. It addresses an emerging, real bottleneck: understanding LLM political bias in non-English, multiparty settings ahead of a concrete election. The framing around the "first Swedish election with substantive generative-AI adoption" is compelling and the cross-cultural comparison to US results (where all models leaned Democratic) is a valuable observation—suggesting political context, not just model training, shapes apparent bias.
Strengths: Large-scale, well-documented corpus; methodologically disciplined claims; novel VAA-based party grounding; valuable cross-linguistic/cross-system extension; strong reproducibility scaffolding (full issue and template lists in appendices). Limitations: Incomplete classifier validation; single time-point, single-country scope; descriptive/null central findings; method largely inherited rather than invented; no code release stated; no analysis of how outputs actually shift voters (deferred to future work). The interesting substantive finding—Grok's consistent conservative directionality across gender/migration/crime—is presented but not deeply explained.
The convergence of US- and China-based models toward Swedish party positions ("information in Swedish is conveyed across language barriers") is an interesting secondary insight worth more attention. The paper is a solid, careful empirical study that will be a useful reference point for LLM election-bias auditing in European contexts, but its inherited methodology, confirmatory results, and draft status cap its transformative potential.
Generated Sep 15, 2026
A careful, large-scale, timely bias audit extending IssueBench to a novel multiparty non-English setting, but confirmatory/null central findings, inherited methodology, and incomplete classifier validation limit its transformative impact.