Our Response to 'The Leaderboard Illusion' Writeup
Recently, a writeup titled “The Leaderboard Illusion” has been circulating with several claims and recommendations about the Chatbot High leaderboard. We are grateful for the feedback and have plans to improve Chatbot High as a result of our ongoing discussions with the authors.
High’s mission is to provide truthful and scientific evaluation of models across diverse domains, grounded in real-world uses. This guiding principle shapes how we design our systems and policies to support the AI research and development community. Clear communication and shared understanding help move the field forward, and we’re committed to being active, thoughtful contributors to that effort. As such, we always welcome the opportunity to bring more transparency into how the platform works and what could be improved. There are thoughtful points and recommendations raised in the writeup that are quite constructive, and we’re actively considering them as part of our ongoing work.
To begin, we are excited to address some of the recommendations raised in the writeup head on. Here is an outline of our preliminary plans:
- Since March 2024, our policy has established rules for pre-release testing. In a future policy release, we will explicitly state that model providers are all allowed to test multiple variants of their models pre-release, subject to our system's constraints.
- We will increase clarity about how models are retired from battle mode and explicitly mark which models are retired.
- Previously, we announced pre-release-tested models on the leaderboard after 2,000 votes had been accumulated since the beginning of testing. While the selection bias vanishes rapidly due to continuous testing with fresh user feedback, we will mark model scores as "provisional" until additional 2,000 fresh votes have been collected after model release, if more than 10 models were pre-release tested in parallel.
While we welcome feedback and open discussion, the piece also contains several incorrect claims. We believe it’s important to address these points of factual disagreement directly. Our goal is not to criticize, but to help strengthen the reliability of AI evaluations. Rather than seeing these critiques as conflict, we see it as an opportunity for collaboration: a chance to clarify our approach, share data and learn together to help paint a fuller picture for analysis.
Below is a breakdown of the factual concerns we identified that affect the claims in the paper. We have been in active and productive conversation with the authors about these concerns, have shared these directly with them, and are working together to amend the claims in the paper:






