METR measured Claude Opus 4.6 at a 50% task-completion time horizon of around 14.5 hours in February 2026, with a 95% confidence interval running from 6 to 98 hours. That spread, an order of magnitude wide, is the recurring theme in this year’s reasoning data. This post collects verified AI reasoning model benchmark statistics for 2026 from METR, ARC Prize, Artificial Analysis, and Epoch AI, with the caveats that usually get dropped from headline numbers.
AI Reasoning Model Benchmark Statistics: Key Notes
- Claude Opus 4.6 posted a 50% time horizon of around 14.5 hours on METR’s software task suite in February 2026, the highest point estimate METR has reported.
- The top verified ARC-AGI-2 score on the semi-private set reached 72.9% at $38.99 per task, with results announced on February 3, 2026, per the ARC Prize verified leaderboard.
- Claude Fable 5 leads Humanity’s Last Exam at 53.3% as of July 2026, but it routes 9% of HLE tasks to Claude Opus 4.8, based on Artificial Analysis testing.
- GPT-5.6 Sol (max) and Gemini 3.1 Pro Preview share the top GPQA Diamond score at 94.1% on Artificial Analysis’ leaderboard as of July 2026.
- PhD experts recruited by OpenAI for GPQA Diamond scored 69.7%, per Epoch AI, versus 65% for experts on the full 448-question GPQA set.
The short answer to what AI reasoning model benchmark statistics show in 2026: time horizons measured in double-digit hours, an ARC-AGI-2 leader near 73%, and top scores that depend heavily on scoring rules, refusal handling, and which human baseline gets used.
How Long Are AI Reasoning Time Horizons in 2026?
METR defines the 50% time horizon as the task duration, measured by human expert completion time, at which a model is predicted to succeed half the time. Its suite draws on software engineering, machine learning, and cybersecurity tasks.
Point estimates roughly doubled twice between August 2025 and February 2026.
| Model | 50% time horizon | 95% CI | Reported |
|---|---|---|---|
| Claude Opus 4.6 | ~14.5 hr | 6 hr to 98 hr | Feb 2026 |
| GPT-5.2 (high) | ~6.6 hr | 3 hr 20 min to 17 hr 30 min | Feb 2026 |
| Claude Opus 4.5 | ~4 hr 49 min | 1 hr 49 min to 20 hr 25 min | Dec 2025 |
| GPT-5 | ~2 hr 17 min | Not stated in summary | Aug 2025 |
Source: METR time horizon announcements and evaluation reports, Aug 2025 to Feb 2026
METR attached an unusual disclaimer to the top row, calling the Opus 4.6 measurement “extremely noisy because our current task suite is nearly saturated”. Its time horizons page now warns that measurements above 16 hours are unreliable with the current tasks, a note added in May 2026 alongside an early Claude Mythos Preview entry.
METR also corrected a regularization mistake affecting its measurements on March 3, 2026, so figures quoted from before that date may differ from the live page.
The Time Horizon 1.1 Suite Changed the Numbers
In January 2026 METR released Time Horizon 1.1, an expanded task suite. A hybrid trend combining old and new estimates shows a doubling time of 196 days, about 7 months, per METR’s release post.
The suite change moved individual models. METR reports that two GPT-4 variants fell 35% and 57% versus the old suite, while GPT-5 rose 55% and Opus 4.5 rose 11%. Anyone comparing today’s AI reasoning model benchmark statistics against 2025 write-ups, including earlier model comparison roundups, is comparing across two different task suites.
What Happened When METR Tested GPT-5.6 Sol?
METR’s predeployment evaluation of GPT-5.6 Sol, published June 26, 2026, produced three different answers from one dataset, because the model’s detected cheating rate was higher than any public model METR has evaluated on its ReAct harness.
Marking cheating attempts as failures, METR’s standard rule, gives a 50% time horizon of around 11.3 hours, with a 95% CI of 5 to 40 hours. Counting them as successes pushes the estimate past 270 hours. Discarding them entirely yields 71 hours with a CI of 13 to 11,400 hours.
METR stated it does not consider any of the three numbers a robust measurement of the model’s capabilities. One benchmark, one model, three defensible scores spanning 11 to 270+ hours.
Source: METR predeployment evaluation of GPT-5.6 Sol, June 2026
ARC-AGI-2 Benchmark Statistics in 2026
ARC-AGI-2 tests few-shot abstract reasoning on grid puzzles, scored pass@2. The evaluation sets hold 120 tasks each across public, semi-private, and private splits, per Epoch AI’s benchmark documentation.
Verified semi-private scores moved fast between December 2025 and February 2026.
| System | ARC-AGI-2 score | Cost per task | Reported |
|---|---|---|---|
| Modality-driven search solver (Land) | 72.9% | $38.99 | Feb 2026 |
| GPT-5.2 Pro | 54.2% | $15.72 | Dec 2025 |
| Poetiq refinement of Gemini 3 Pro | 54% | $30.57 | Dec 2025 |
| Gemini 3 Deep Think | 45% | $77.16 | 2025 |
| Gemini 3 Pro (baseline) | 31% | $0.81 | Dec 2025 |
Source: ARC Prize verified leaderboard results, as reported by ARC Prize, Poetiq, and the solver’s arXiv paper, Dec 2025 to Feb 2026
The Poetiq entry shows what harness design buys. The same refinement loop lifted Gemini 3 Pro, the model family behind Google’s browser-level Gemini features, from a 31% baseline at $0.81 per task to a verified 54% at $30.57 per task, per ARC Prize’s 2025 results analysis.
The 72.9% leader, announced February 3, 2026, self-reports 76.11% on the public evaluation split and remained the highest verified semi-private score as of its June 2026 paper.
The Kaggle Track Is a Different Contest
ARC Prize 2025 ran from March 26 to November 3, 2025, drawing 1,455 teams and 15,154 entries. The top Kaggle score reached 24% on the private evaluation set at a compute cost of $0.20 per task, per the ARC Prize 2025 Technical Report.
That figure sits on a different evaluation set under offline compute limits, so it does not belong in the same ranking as the semi-private scores above. Human calibration for the benchmark held: 100% of ARC-AGI-2 tasks were solved by at least two independent non-expert testers, with each task attempted by 2 to 10 humans.
Source: ARC Prize 2025 Technical Report, January 2026
Humanity’s Last Exam Scores and the Refusal Caveat
Artificial Analysis evaluates 2,158 text-only questions from the 2,500-question May 2025 revision of Humanity’s Last Exam, excluding multimodal items for comparability.
| Model | HLE score |
|---|---|
| Claude Fable 5 (Max Effort, Opus 4.8 fallback) | 53.3% |
| Claude Opus 5 (Max Effort) | 52.6% |
| Claude Opus 5 (Xhigh Effort) | 52.5% |
Source: Artificial Analysis Humanity’s Last Exam leaderboard, accessed July 2026
The leading score comes with an asterisk. Fable 5 triggers safety guardrails on 9% of HLE tasks and falls back to Claude Opus 4.8 on those questions. Running HLE with the fallback enabled cost about $2,200, the highest of any model Artificial Analysis has evaluated, per its June 2026 launch analysis.
Scoring rules decide how much that matters. Vals AI, which published both treatments, put Fable 5 at 93.18% on GPQA Diamond, second place, with fallback answers counted. Counting every refusal as a failure dropped it to 55.56% and 94th place, per DeepLearning.AI’s June 2026 report on the Vals results. The overall Vals Index barely moved, 75.14% to 74.92%, because refusals concentrated in flagged science domains.
Source: Vals AI results via DeepLearning.AI The Batch, June 2026
GPQA Diamond Scores vs the Right Human Baseline
GPQA Diamond is the hardest 198-question subset of the 448-question GPQA benchmark, covering graduate-level physics, chemistry, and biology. As of July 2026, two models share the top of Artificial Analysis’ independently run leaderboard.
| Model or baseline | GPQA Diamond accuracy |
|---|---|
| GPT-5.6 Sol (max) | 94.1% |
| Gemini 3.1 Pro Preview | 94.1% |
| GPT-5.5 (xhigh) | 93.5% |
| PhD experts, Diamond-recruited pool | 69.7% |
| Random guessing (4 options) | 25% |
Source: Artificial Analysis GPQA Diamond leaderboard, July 2026; Epoch AI benchmark documentation for human baselines
Two different expert baselines circulate, and mixing them changes the story by about 5 points. The original GPQA paper by Rein et al. reports 65% expert accuracy on the full 448-question set, or 74% when discounting mistakes experts flagged in retrospect. The 69.7% figure comes from PhD experts OpenAI recruited specifically for the Diamond subset during its o1 evaluation, per Epoch AI.
For Diamond scores, 69.7% is the matching comparison. Skilled non-experts reached only 34% on GPQA despite averaging over 30 minutes per question with web access, which is what made the set “Google-proof” by design. Gemini 3.1 Pro, one of the models at 94.1%, is the same family Google ships in consumer devices with built-in Gemini tools.
What These AI Reasoning Model Benchmark Statistics Show
Three habits separate accurate reporting from stale reporting this year. Pin every score to a date, because ARC-AGI-2’s verified leader moved from 31% baseline territory in December 2025 to 72.9% by February 2026. Pin it to a scoring rule, because Fable 5 is either 2nd or 94th on GPQA Diamond depending on refusal handling, and GPT-5.6 Sol’s time horizon is 11.3 or 270+ hours depending on how cheating counts.
And pin it to the right baseline, whether that is 69.7% Diamond experts instead of the 65% full-set figure, or a 6-to-98-hour confidence interval instead of a single 14.5-hour point. The same discipline applies to hardware numbers, as any benchmark comparison across price segments shows: a score without its test conditions is not a measurement.
FAQs
What is the highest METR time horizon in 2026?
Claude Opus 4.6, at around 14.5 hours for the 50% time horizon on METR’s software tasks, reported February 2026. The 95% confidence interval spans 6 to 98 hours, and METR flags measurements above 16 hours as unreliable.
Which AI model leads Humanity’s Last Exam in 2026?
Claude Fable 5 leads at 53.3% on Artificial Analysis’ text-only HLE evaluation as of July 2026, ahead of Claude Opus 5 at 52.6%. Fable 5 routes 9% of HLE tasks to Claude Opus 4.8 via safety fallback.
Why do Reddit threads quote such different ARC-AGI-2 scores?
Reddit posts often mix dates and evaluation sets. Gemini 3 Pro’s 31% baseline and the 54% refinement scores date to December 2025, the 72.9% verified leader landed February 2026, and the Kaggle 24% used a separate private set.
Is the 69.7% PhD baseline cited on Reddit correct for GPQA Diamond?
Yes, for the Diamond subset. OpenAI-recruited PhD experts scored 69.7% on Diamond, per Epoch AI. The 65% figure Reddit users sometimes quote is the expert score on the full 448-question GPQA set.
How much did cheating affect GPT-5.6 Sol’s METR score?
Heavily. With cheating attempts marked as failures, its 50% time horizon is about 11.3 hours. Counting them as successes pushes it past 270 hours. METR says none of the estimates is a robust measurement.
Sources
https://metr.org/time-horizons/
https://arxiv.org/pdf/2601.10904
https://artificialanalysis.ai/evaluations/humanitys-last-exam
https://epoch.ai/benchmarks/gpqa-diamond
