AI capability claims

Did Llama 4 Maverick really rank #2 on LMArena?

Meta's Llama 4 launch cited a #2 position on the LMArena leaderboard. The model on the board was an unreleased experimental variant; the public weights, re-tested, ranked 32nd. What does the evidence say about the launch claim?

by WeirdTrue · 2026-09-22

Claims

#1 Meta, Llama 4 launch announcement said on 2025-04-05 · source

Llama 4 Maverick achieves an ELO of 1417 on LMArena, ranking #2.
Partially contradictedVersion 1 · Model qwen-max · verdict_v1 · ai-claims@1 · 2026-09-22 15:56

The claim that Llama 4 Maverick ranked #2 on LMArena with an ELO of 1417 is partially contradicted. While Meta did make this claim in their launch announcement, evidence shows that the version tested and ranked on LMArena was an experimental variant, not the public release. The public version, when re-tested, ranked 32nd. This discrepancy indicates that the claim, as stated, is not fully supported by the evidence.

Citations

  • ai.meta.com: “launch post presents Maverick's LMArena ELO of 1417 and #2 rank; a footnote states the tested version was an experimental chat version optimised for conversationality.
  • www.theregister.com: “the unmodified public model, when re-tested, ranked 32nd.
  • arxiv.org: “Meta privately tested 27 Llama 4 variants on Chatbot Arena before launch and published only the best; the submitted model, Llama-4-Maverick-03-26-Experimental, differs from the released weights.

What would change this

  • Independent verification of the public Llama 4 Maverick's performance on LMArena.

Evidence (4)

Add evidence

Timeline

  1. 2026-09-22 15:56 · Verdict issued → Partially contradicted
  2. 2026-09-22 15:56 · Evidence added
  3. 2026-09-22 15:56 · Evidence added
  4. 2026-09-22 15:56 · Evidence added
  5. 2026-09-22 15:56 · Evidence added
  6. 2026-09-22 15:56 · Claim added
  7. 2026-09-22 15:56 · Case opened