meta analysis

BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices

Toriqul Rahat*University of ScholarsAnka Reuel*Stanford University

* Corresponding author

Received

8/27/2026

Revised

Accepted

8/27/2026

Published

8/27/2026

No DOI registered yet CC BY 4.0

Abstract

AI models are increasingly prevalent in high-stakes environments, necessitating thorough assessment of their capabilities and risks. Benchmarks are popular for measuring these attributes and for comparing model performance, tracking progress, and identifying weaknesses in foundation and non-foundation models. They can inform model selection for downstream tasks and influence policy initiatives. However, not all benchmarks are the same: their quality depends on their design and usability. In this paper, we develop an assessment framework considering 46 best practices across an AI benchmark’s lifecycle and evaluate 24 AI benchmarks against it. We find that there exist large quality differences and that commonly used benchmarks suffer from significant issues. We further find that most benchmarks do not report statistical significance of their results nor allow for their results to be easily replicated. To support benchmark developers in aligning with best practices, we provide a checklist for minimum quality assurance based on our assessment. We also develop a living repository of benchmark assessments to support benchmark comparability, accessible at betterbench.stanford.edu.

Figures

Figures will be displayed here when provided during production. Cloudinary-hosted figure assets are rendered per article.

Tables

Tables will be displayed here when provided during production.

References

References placeholder

References will be listed here. Structured reference metadata is stored per article and exported via JATS.

Cite this article

@article{rahat2026,
  title={BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices},
  author={Rahat, Toriqul and Reuel, Anka},
  journal={AI Review Letters},
  year={2026},
  volume={},
  number={},
  pages={},
  doi={},
  url={http://localhost:3000/articles/betterbench-assessing-ai-benchmarks-uncovering-issues-and-establishing-best-prac-2026-19ygfw}
}

Canonical: http://localhost:3000/articles/betterbench-assessing-ai-benchmarks-uncovering-issues-and-establishing-best-prac-2026-19ygfw

Metrics

Views

Downloads

Citations

Metrics placeholders — integrate Altmetric / Crossref / Dimensions when available.

Article info

Article number: 2026-19YGFW

Type: meta analysis

Journal: AI Review Letters