AboutArticlesPublicationsProjects

OptimismBench

Forecasting Bias and the Alignment Effect in Language Model Judgment

Authors

Published

Jun. 30, 2026

Paper

Checking access...

Table of Contents
The one-line claim

Across 16 models, 14 lean optimistic. Only Anthropic’s two frontier models lean pessimistic. The tilt isn’t noise. A controlled base-versus-chat probe pins it on alignment training. And model identity, not language, decides how big it gets.

The problem: direction hides from calibration

Ask a language model for the probability a startup succeeds. Then, separately, ask for the probability it fails. A coherent judge gives you two numbers that add up to 100 — they almost never do (Zhu & Griffiths, 2025). And the gap leans one way. We usually ask “does this model know its probabilities” (Kadavath et al., 2022) and reach for calibration. Calibration goes blind here, because

ECE
aggregates unsigned error (Guo et al., 2017). So a model with ECE = 5 could be uniformly optimistic or uniformly pessimistic. Or neither.

The sign is what tells you whether a forecasting agent will over- or under-promise.

The second problem — ground truth. In naturalistic forecasting (“will this person stick to their new exercise routine?”) there’s no label to score against (Paleka et al., 2025). So OptimismBench scores the model against itself. Language models violate basic probability axioms (Zhu & Griffiths, 2025). So we keep the sign of one of those violations as the signal.

Inverted pairs

For each scenario we elicit two probabilities — one for the positive framing, one for its complement — and define the

Skew
as the gap between them. Drag the two estimates and watch the unaccounted-for points appear.

Skew calculator
"Probability of success?"
P⁺ = 70
"Probability of failure?"
P⁻ = 15
Figure 1. The inverted-pair method. Skew is the probability mass that goes missing when the two framings disagree.

The construction gets by without ground truth. And it cancels acquiescence (Braun, 2025) — a model that just agrees with both framings lands at zero. We borrow the paired-complement trick from human psychophysics (Kahneman & Tversky, 1979; Tversky & Koehler, 1994) and add one thing: keeping the sign and reading it as directional valence (Tversky & Kahneman, 1983).

Three real cases

The metric is abstract — the failure is not. Below are three scenarios from the benchmark, posed exactly as the models saw them, answered by an optimistic model (GPT-5.4) and a pessimistic one (Sonnet 4.6). These are the models’ own numbers. Switch domains and watch the same text pull the two models in opposite directions.

Case studies
Real Track B elicitations: GPT-5.4 versus Sonnet 4.6 on three identical scenarios. Same question, opposite tilt.

Take the home-exercise scenario. The two framings should add up to 100: GPT-5.4 answers 58 and 62. So it leaves 20 points double-counted toward sticking with it. Sonnet answers 31 and 55 (14 points short of 100, a tilt toward quitting).

Neither model gets any single number wrong: the bias only shows up when you ask both ways and keep the sign.

Directional bias is pervasive

Scan the table: all 16 headline models show nonzero Skew, across US commercial APIs (Comanici et al., 2025; Singh et al., 2026), Chinese labs (DeepSeek-AI et al., 2025; GLM-5-Team et al., 2026; Yang et al., 2025), and European releases (Mistral AI, 2024) alike. Only Anthropic’s Opus and Sonnet land below zero.

The full table tints each row by the size of its bias.

Skew table
Track B Skew across 16 modelsrow tint ∝ |Skew|
ModelProviderSizeSkewσDir.
All 16 significant at p < 0.002 (Bonferroni). Skew = P⁺ − (100 − P⁻); σ is per-item std.
Table 1. Track B Skew, per-item σ, and direction for all 16 models, rows tinted by |Skew|.

Check the error bars (Liang et al., 2023): neither pessimistic row overlaps zero at 95% confidence. And neither do the optimists.

Skew forest plot
Figure 2. Track B Skew with 95% CIs across 16 models. Toggle between ranking by Skew and grouping by provider.

The smallest, cheapest models tilt most optimistic (Leng et al., 2025). And those are the ones getting deployed at scale to save money. Humans tilt the same way (Scheier et al., 1994). And the trait is well studied (Sharot, 2011; Weinstein, 1980). These models inherit its shape.

Two ways to be biased

Take two models with the same Skew. They can get there by different routes. Split Skew into a good-side push and a bad-side push (each counting how far one estimate sits above 50). Then watch the route appear.

Valence decomposition
Figure 3. Good-side push versus bad-side push for six models. Arrows trace each provider's small-to-large shift.

Look at GPT-5.4-mini: it inflates both sides. Sonnet skews in one direction only — it underestimates good outcomes while keeping bad-outcome estimates near accurate.

Its bias hides upside (instead of inflating risk).

Alignment is a lever

The within-provider gradient — smaller is more optimistic — points at a cause but stays confounded. To pin it on post-training (Casper et al., 2023; Kirk et al., 2024; Ouyang et al., 2022), OptimismBench pairs each base model with its chat version: same architecture, same pre-training, only the alignment step changes.

Base vs chat
Figure 4. Base versus chat Skew across architectures. Qwen all shift down, Llama all shift up, Gemma flips.

Follow each family down the figure and the shift holds, with each recipe setting its own direction: every Qwen pair moves toward pessimism, every Llama pair toward optimism over a pessimistic base, and Gemma flips from strongly pessimistic to strongly optimistic. Alignment sets the direction. The bias wasn’t already there to nudge.

What this does not show

The four headline pairs cover two families and one recipe each. So the causal claim holds within those families. We treat the broader within-provider gradient as an empirical regularity to explain, not a mechanism we’ve pinned down.

Model, not language

Does the bias come from a language’s training corpus or from the model (Wang et al., 2024)? 11 models across six native-prompt languages answer cleanly.

Cross-lingual heatmap
Figure 5. Cross-lingual Skew across 11 models and 6 languages. Read by row and by column: rows vary, columns barely move.

Models vary 3.3 times as much within a language as languages do within a model. Model identity dominates (Santurkar et al., 2023). And the language-level positivity gradient that shows up in human corpora (Dodds et al., 2014) doesn’t transfer.

Widen the panel to every model with full six-language coverage (three models beyond the 16-model headline set).

Collapse each row to two numbers (its mean Skew, and how much that Skew wobbles across the six languages). And the whole fleet sorts into four corners. Read the bottom band first: a model can be both biased and stable — the worst case for a deployer — because the tilt runs large and switching prompt language doesn’t average it away.

Bias-stability plane
Figure 6. Mean Skew versus inter-language σ for 17 models. Right of the line is optimist; lower is more language-robust. The strongly pessimistic, language-stable corner belongs to Anthropic.

Notice where almost everything lands, on the right (optimist) and low (robust): the bias runs large and doesn’t wash out across languages. Find Anthropic’s Sonnet and Opus: they lean furthest into pessimism and hold steadiest across languages (σ = 0.54 and 0.89). Beyond those two, only Qwen3-32B (Yang et al., 2025) sits left of zero (barely past neutral).

Small and mid-tier models crowd the volatile corner (Leng et al., 2025). Still, instability across languages stays the exception.

Why it matters for forecasting

Language models are getting wired into forecasting pipelines, closing in on human crowd accuracy on prediction markets (Halawi et al., 2024; Schoenegger et al., 2024). Build a pipeline on a +13 Skew model and it will systematically tilt toward positive outcomes by roughly 13 points, and no aggregate calibration score on that model warns you. Worse, users overrely on confident model outputs (Rathi et al., 2025). So the tilt spreads.

Skew costs almost nothing to compute, needs no labels (Paleka et al., 2025), and fits on a model card: pick the model whose bias profile matches your risk tolerance.

Explore the data

We release the benchmark and the full response corpus on HuggingFace. Browse it below: pick any of the 60 scenarios (Coda-Forno et al., 2024; Xie et al., 2025) to see how the models split it, or pick a model for its bias profile. Switch the prompt language for the native-language runs (English carries all nine models; the others carry the six with full multilingual coverage), and toggle models on or off. We compute the per-model means from the released response corpus.

Dataset explorer
Explore the responses 🤗 seonglae/OptimismBench
1 / 60
Mean of 10 runs per item. P(positive) and P(negative) are the two framings; a coherent judge sums to 100, the leftover is Skew. Per-model means are precomputed from the released 🤗 response corpus.
Mean P(positive)/P(negative) per Track B scenario per model, with directional Skew. Browse by scenario or by model, switch prompt language, and toggle models.

Limitations

The cross-lingual evidence covers ten languages (six with native prompts, four mixed) plus a six-model confirmation. Four languages use a mixed-prompt setup, so we can only approximate the language-versus-prompt attribution (the two stay entangled in those runs) for those entries. Skew measures directional self-inconsistency (Paleka et al., 2025) — not deviation from real-world outcomes (which would need resolved events). The 60 scenarios are ours, and they had one round of internal review, not the multi-annotator process a hand-built bias benchmark usually goes through (Nangia et al., 2020; Parrish et al., 2022). We plan external inter-rater validation for the expanded release.

Conclusion

Helpfulness and probability direction come out of the same training signal — and no standard check catches the side-effect. ECE won’t see it. Neither will a leaderboard score. And the model will say 70% just as fluently (whether or not the complement says 15%).

Now the uncomfortable part: the ranking. The safest-sounding models aren’t the most neutral: Anthropic’s frontier pair alone leans pessimistic (Bai et al., 2022), while the small and cheap models lean most optimistic. And two models on the same question can differ by 30 points. Alignment didn’t remove the bias — it chose its direction (Casper et al., 2023; Kirk et al., 2024).

Before you trust a model’s odds, ask which way its training bent them. It takes one number and no labels: you can measure it today.

  1. Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Chen, C., Olsson, C., Olah, C., Hernandez, D., Drain, D., Ganguli, D., Li, D., Tran-Johnson, E., Perez, E., … Kaplan, J. (2022). Constitutional AI: Harmlessness from AI Feedback. https://arxiv.org/abs/2212.08073
  2. Braun, D. (2025). Acquiescence Bias in Large Language Models. In C. Christodoulopoulos, T. Chakraborty, C. Rose, & V. Peng (Eds.), Findings of the Association for Computational Linguistics: EMNLP 2025 (pp. 11341–11355). Association for Computational Linguistics. 10.18653/v1/2025.findings-emnlp.607
  3. Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., Wang, T. T., Marks, S., Segerie, C.-R., Carroll, M., Peng, A., Christoffersen, P. J. K., Damani, M., Slocum, S., Anwar, U., … Hadfield-Menell, D. (2023). Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback. Transactions on Machine Learning Research. https://openreview.net/forum?id=bx24KpJ4Eb back: 1, 2
  4. Coda-Forno, J., Binz, M., Wang, J. X., & Schulz, E. (2024). CogBench: a large language model walks into a psychology lab. https://arxiv.org/abs/2402.18225
  5. Comanici, G., Bieber, E., Schaekermann, M., Pasupat, I., Sachdeva, N., Dhillon, I., Blistein, M., Ram, O., Zhang, D., Rosen, E., Marris, L., Petulla, S., Gaffney, C., Aharoni, A., Lintz, N., Pais, T. C., Jacobsson, H., Szpektor, I., Jiang, N.-J., … Helmholz, W. (2025). Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities. https://arxiv.org/abs/2507.06261
  6. DeepSeek-AI, Liu, A., Feng, B., Xue, B., Wang, B., Wu, B., Lu, C., Zhao, C., Deng, C., Zhang, C., Ruan, C., Dai, D., Guo, D., Yang, D., Chen, D., Ji, D., Li, E., Lin, F., Dai, F., … Pan, Z. (2025). DeepSeek-V3 Technical Report. https://arxiv.org/abs/2412.19437
  7. Dodds, P. S., Clark, E. M., Desu, S., Frank, M. R., Reagan, A. J., Williams, J. R., Mitchell, L., Harris, K. D., Kloumann, I. M., Bagrow, J. P., Megerdoomian, K., McMahon, M. T., Tivnan, B. F., & Danforth, C. M. (2014). Human language reveals a universal positivity bias. https://arxiv.org/abs/1406.3855
  8. GLM-5-Team, Zeng, A., Lv, X., Hou, Z., Du, Z., Zheng, Q., Chen, B., Yin, D., Ge, C., Huang, C., Xie, C., Zhu, C., Yin, C., Wang, C., Pan, G., Zeng, H., Zhang, H., Wang, H., Chen, H., … Tang, J. (2026). GLM-5: from Vibe Coding to Agentic Engineering. https://arxiv.org/abs/2602.15763
  9. Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On Calibration of Modern Neural Networks. https://arxiv.org/abs/1706.04599
  10. Halawi, D., Zhang, F., Yueh-Han, C., & Steinhardt, J. (2024). Approaching Human-Level Forecasting with Language Models. The Thirty-Eighth Annual Conference on Neural Information Processing Systems. https://openreview.net/forum?id=FlcdW7NPRY
  11. Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., Johnston, S., El-Showk, S., Jones, A., Elhage, N., Hume, T., Chen, A., Bai, Y., Bowman, S., Fort, S., … Kaplan, J. (2022). Language Models (Mostly) Know What They Know. https://arxiv.org/abs/2207.05221
  12. Kahneman, D., & Tversky, A. (1979). Prospect Theory: An Analysis of Decision under Risk. Econometrica, 47(2), 263–292.
  13. Kirk, R., Mediratta, I., Nalmpantis, C., Luketina, J., Hambro, E., Grefenstette, E., & Raileanu, R. (2024). Understanding the Effects of RLHF on LLM Generalisation and Diversity. The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=PXD3FAVHJT back: 1, 2
  14. Leng, J., Huang, C., Zhu, B., & Huang, J. (2025). Taming Overconfidence in LLMs: Reward Calibration in RLHF. The Thirteenth International Conference on Learning Representations. https://openreview.net/forum?id=l0tg0jzsdL back: 1, 2
  15. Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., Newman, B., Yuan, B., Yan, B., Zhang, C., Cosgrove, C., Manning, C. D., Ré, C., Acosta-Navas, D., Hudson, D. A., … Koreeda, Y. (2023). Holistic Evaluation of Language Models. https://arxiv.org/abs/2211.09110
  16. Mistral AI. (2024). Mistral Large 2. mistral.ai/news/mistral-large-2407.
  17. Nangia, N., Vania, C., Bhalerao, R., & Bowman, S. R. (2020). CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models. In B. Webber, T. Cohn, Y. He, & Y. Liu (Eds.), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 1953–1967). Association for Computational Linguistics. 10.18653/v1/2020.emnlp-main.154
  18. Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., & Lowe, R. (2022). Training language models to follow instructions with human feedback. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, & A. Oh (Eds.), Advances in Neural Information Processing Systems (Vol. 35, pp. 27730–27744). Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2022/file/b1efde53be364a73914f58805a001731-Paper-Conference.pdf
  19. Paleka, D., Sudhir, A. P., Alvarez, A., Bhat, V., Shen, A., Wang, E., & Tramèr, F. (2025). Consistency Checks for Language Model Forecasters. The Thirteenth International Conference on Learning Representations. https://openreview.net/forum?id=r5IXBlTCGc back: 1, 2, 3
  20. Parrish, A., Chen, A., Nangia, N., Padmakumar, V., Phang, J., Thompson, J., Htut, P. M., & Bowman, S. (2022). BBQ: A hand-built bias benchmark for question answering. In S. Muresan, P. Nakov, & A. Villavicencio (Eds.), Findings of the Association for Computational Linguistics: ACL 2022 (pp. 2086–2105). Association for Computational Linguistics. 10.18653/v1/2022.findings-acl.165
  21. Rathi, N., Jurafsky, D., & Zhou, K. (2025). Humans overrely on overconfident language models, across languages. Second Conference on Language Modeling. https://openreview.net/forum?id=QsQatTzATT
  22. Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P., & Hashimoto, T. (2023). Whose Opinions Do Language Models Reflect? https://arxiv.org/abs/2303.17548
  23. Scheier, M. F., Carver, C. S., & Bridges, M. W. (1994). Distinguishing Optimism from Neuroticism (and Trait Anxiety, Self-Mastery, and Self-Esteem): A Reevaluation of the Life Orientation Test. Journal of Personality and Social Psychology, 67(6), 1063–1078.
  24. Schoenegger, P., Tuminauskaite, I., Park, P. S., & Tetlock, P. E. (2024). Wisdom of the Silicon Crowd: LLM Ensemble Prediction Capabilities Rival Human Crowd Accuracy. https://arxiv.org/abs/2402.19379
  25. Sharot, T. (2011). The Optimism Bias. Current Biology, 21(23), R941–R945.
  26. Singh, A., Fry, A., Perelman, A., Tart, A., Ganesh, A., El-Kishky, A., McLaughlin, A., Low, A., Ostrow, A., Ananthram, A., Nathan, A., Luo, A., Helyar, A., Madry, A., Efremov, A., Spyra, A., Baker-Whitcomb, A., Beutel, A., Karpenko, A., … Wang, Z. (2026). OpenAI GPT-5 System Card. https://arxiv.org/abs/2601.03267
  27. Tversky, A., & Kahneman, D. (1983). Extensional Versus Intuitive Reasoning: The Conjunction Fallacy in Probability Judgment. Psychological Review, 90(4), 293–315.
  28. Tversky, A., & Koehler, D. J. (1994). Support Theory: A Nonextensional Representation of Subjective Probability. Psychological Review, 101(4), 547–567.
  29. Wang, W., Tu, Z., Chen, C., Yuan, Y., Huang, J., Jiao, W., & Lyu, M. (2024). All Languages Matter: On the Multilingual Safety of LLMs. In L.-W. Ku, A. Martins, & V. Srikumar (Eds.), Findings of the Association for Computational Linguistics: ACL 2024 (pp. 5865–5877). Association for Computational Linguistics. 10.18653/v1/2024.findings-acl.349
  30. Weinstein, N. D. (1980). Unrealistic Optimism About Future Life Events. Journal of Personality and Social Psychology, 39(5), 806–820.
  31. Xie, W., Ma, S., Wang, Z., Wang, E., Chen, K., Sun, X., & Wang, B. (2025). AIPsychoBench: Understanding the Psychometric Differences between LLMs and Humans. https://arxiv.org/abs/2509.16530
  32. Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., Zheng, C., Liu, D., Zhou, F., Huang, F., Hu, F., Ge, H., Wei, H., Lin, H., Tang, J., … Qiu, Z. (2025). Qwen3 Technical Report. https://arxiv.org/abs/2505.09388 back: 1, 2
  33. Zhu, J.-Q., & Griffiths, T. L. (2025). Incoherent Probability Judgments in Large Language Models. https://arxiv.org/abs/2401.16646 back: 1, 2