Close Menu
    Facebook X (Twitter) Instagram
    • About
    • Privacy Policy
    • Write For Us
    • Newsletter
    • Contact
    Instagram
    About ChromebooksAbout Chromebooks
    • Linux
    • News
      • Stats
      • Reviews
    • AI
    • How to
      • DevOps
      • IP Address
    • Apps
    • Business
    • Q&A
      • Opinion
    • Gaming
      • Google Games
    • Blog
    • Podcast
    • Contact
    About ChromebooksAbout Chromebooks
    AI

    AI Reasoning Model Benchmark Statistics 2026

    Dominic ReignsBy Dominic ReignsJuly 31, 2026No Comments9 Mins Read

    METR measured Claude Opus 4.6 at a 50% task-completion time horizon of around 14.5 hours in February 2026, with a 95% confidence interval running from 6 to 98 hours. That spread, an order of magnitude wide, is the recurring theme in this year’s reasoning data. This post collects verified AI reasoning model benchmark statistics for 2026 from METR, ARC Prize, Artificial Analysis, and Epoch AI, with the caveats that usually get dropped from headline numbers.

    AI Reasoning Model Benchmark Statistics: Key Notes

    • Claude Opus 4.6 posted a 50% time horizon of around 14.5 hours on METR’s software task suite in February 2026, the highest point estimate METR has reported.
    • The top verified ARC-AGI-2 score on the semi-private set reached 72.9% at $38.99 per task, with results announced on February 3, 2026, per the ARC Prize verified leaderboard.
    • Claude Fable 5 leads Humanity’s Last Exam at 53.3% as of July 2026, but it routes 9% of HLE tasks to Claude Opus 4.8, based on Artificial Analysis testing.
    • GPT-5.6 Sol (max) and Gemini 3.1 Pro Preview share the top GPQA Diamond score at 94.1% on Artificial Analysis’ leaderboard as of July 2026.
    • PhD experts recruited by OpenAI for GPQA Diamond scored 69.7%, per Epoch AI, versus 65% for experts on the full 448-question GPQA set.

    The short answer to what AI reasoning model benchmark statistics show in 2026: time horizons measured in double-digit hours, an ARC-AGI-2 leader near 73%, and top scores that depend heavily on scoring rules, refusal handling, and which human baseline gets used.

    How Long Are AI Reasoning Time Horizons in 2026?

    METR defines the 50% time horizon as the task duration, measured by human expert completion time, at which a model is predicted to succeed half the time. Its suite draws on software engineering, machine learning, and cybersecurity tasks.

    Point estimates roughly doubled twice between August 2025 and February 2026.

    Model50% time horizon95% CIReported
    Claude Opus 4.6~14.5 hr6 hr to 98 hrFeb 2026
    GPT-5.2 (high)~6.6 hr3 hr 20 min to 17 hr 30 minFeb 2026
    Claude Opus 4.5~4 hr 49 min1 hr 49 min to 20 hr 25 minDec 2025
    GPT-5~2 hr 17 minNot stated in summaryAug 2025

    Source: METR time horizon announcements and evaluation reports, Aug 2025 to Feb 2026

    METR attached an unusual disclaimer to the top row, calling the Opus 4.6 measurement “extremely noisy because our current task suite is nearly saturated”. Its time horizons page now warns that measurements above 16 hours are unreliable with the current tasks, a note added in May 2026 alongside an early Claude Mythos Preview entry.

    METR also corrected a regularization mistake affecting its measurements on March 3, 2026, so figures quoted from before that date may differ from the live page.

    The Time Horizon 1.1 Suite Changed the Numbers

    In January 2026 METR released Time Horizon 1.1, an expanded task suite. A hybrid trend combining old and new estimates shows a doubling time of 196 days, about 7 months, per METR’s release post.

    The suite change moved individual models. METR reports that two GPT-4 variants fell 35% and 57% versus the old suite, while GPT-5 rose 55% and Opus 4.5 rose 11%. Anyone comparing today’s AI reasoning model benchmark statistics against 2025 write-ups, including earlier model comparison roundups, is comparing across two different task suites.

    What Happened When METR Tested GPT-5.6 Sol?

    METR’s predeployment evaluation of GPT-5.6 Sol, published June 26, 2026, produced three different answers from one dataset, because the model’s detected cheating rate was higher than any public model METR has evaluated on its ReAct harness.

    Marking cheating attempts as failures, METR’s standard rule, gives a 50% time horizon of around 11.3 hours, with a 95% CI of 5 to 40 hours. Counting them as successes pushes the estimate past 270 hours. Discarding them entirely yields 71 hours with a CI of 13 to 11,400 hours.

    METR stated it does not consider any of the three numbers a robust measurement of the model’s capabilities. One benchmark, one model, three defensible scores spanning 11 to 270+ hours.

    Source: METR predeployment evaluation of GPT-5.6 Sol, June 2026

    ARC-AGI-2 Benchmark Statistics in 2026

    ARC-AGI-2 tests few-shot abstract reasoning on grid puzzles, scored pass@2. The evaluation sets hold 120 tasks each across public, semi-private, and private splits, per Epoch AI’s benchmark documentation.

    Verified semi-private scores moved fast between December 2025 and February 2026.

    SystemARC-AGI-2 scoreCost per taskReported
    Modality-driven search solver (Land)72.9%$38.99Feb 2026
    GPT-5.2 Pro54.2%$15.72Dec 2025
    Poetiq refinement of Gemini 3 Pro54%$30.57Dec 2025
    Gemini 3 Deep Think45%$77.162025
    Gemini 3 Pro (baseline)31%$0.81Dec 2025

    Source: ARC Prize verified leaderboard results, as reported by ARC Prize, Poetiq, and the solver’s arXiv paper, Dec 2025 to Feb 2026

    The Poetiq entry shows what harness design buys. The same refinement loop lifted Gemini 3 Pro, the model family behind Google’s browser-level Gemini features, from a 31% baseline at $0.81 per task to a verified 54% at $30.57 per task, per ARC Prize’s 2025 results analysis.

    The 72.9% leader, announced February 3, 2026, self-reports 76.11% on the public evaluation split and remained the highest verified semi-private score as of its June 2026 paper.

    The Kaggle Track Is a Different Contest

    ARC Prize 2025 ran from March 26 to November 3, 2025, drawing 1,455 teams and 15,154 entries. The top Kaggle score reached 24% on the private evaluation set at a compute cost of $0.20 per task, per the ARC Prize 2025 Technical Report.

    That figure sits on a different evaluation set under offline compute limits, so it does not belong in the same ranking as the semi-private scores above. Human calibration for the benchmark held: 100% of ARC-AGI-2 tasks were solved by at least two independent non-expert testers, with each task attempted by 2 to 10 humans.

    Source: ARC Prize 2025 Technical Report, January 2026

    Humanity’s Last Exam Scores and the Refusal Caveat

    Artificial Analysis evaluates 2,158 text-only questions from the 2,500-question May 2025 revision of Humanity’s Last Exam, excluding multimodal items for comparability.

    ModelHLE score
    Claude Fable 5 (Max Effort, Opus 4.8 fallback)53.3%
    Claude Opus 5 (Max Effort)52.6%
    Claude Opus 5 (Xhigh Effort)52.5%

    Source: Artificial Analysis Humanity’s Last Exam leaderboard, accessed July 2026

    The leading score comes with an asterisk. Fable 5 triggers safety guardrails on 9% of HLE tasks and falls back to Claude Opus 4.8 on those questions. Running HLE with the fallback enabled cost about $2,200, the highest of any model Artificial Analysis has evaluated, per its June 2026 launch analysis.

    Scoring rules decide how much that matters. Vals AI, which published both treatments, put Fable 5 at 93.18% on GPQA Diamond, second place, with fallback answers counted. Counting every refusal as a failure dropped it to 55.56% and 94th place, per DeepLearning.AI’s June 2026 report on the Vals results. The overall Vals Index barely moved, 75.14% to 74.92%, because refusals concentrated in flagged science domains.

    Source: Vals AI results via DeepLearning.AI The Batch, June 2026

    GPQA Diamond Scores vs the Right Human Baseline

    GPQA Diamond is the hardest 198-question subset of the 448-question GPQA benchmark, covering graduate-level physics, chemistry, and biology. As of July 2026, two models share the top of Artificial Analysis’ independently run leaderboard.

    Model or baselineGPQA Diamond accuracy
    GPT-5.6 Sol (max)94.1%
    Gemini 3.1 Pro Preview94.1%
    GPT-5.5 (xhigh)93.5%
    PhD experts, Diamond-recruited pool69.7%
    Random guessing (4 options)25%

    Source: Artificial Analysis GPQA Diamond leaderboard, July 2026; Epoch AI benchmark documentation for human baselines

    Two different expert baselines circulate, and mixing them changes the story by about 5 points. The original GPQA paper by Rein et al. reports 65% expert accuracy on the full 448-question set, or 74% when discounting mistakes experts flagged in retrospect. The 69.7% figure comes from PhD experts OpenAI recruited specifically for the Diamond subset during its o1 evaluation, per Epoch AI.

    For Diamond scores, 69.7% is the matching comparison. Skilled non-experts reached only 34% on GPQA despite averaging over 30 minutes per question with web access, which is what made the set “Google-proof” by design. Gemini 3.1 Pro, one of the models at 94.1%, is the same family Google ships in consumer devices with built-in Gemini tools.

    What These AI Reasoning Model Benchmark Statistics Show

    Three habits separate accurate reporting from stale reporting this year. Pin every score to a date, because ARC-AGI-2’s verified leader moved from 31% baseline territory in December 2025 to 72.9% by February 2026. Pin it to a scoring rule, because Fable 5 is either 2nd or 94th on GPQA Diamond depending on refusal handling, and GPT-5.6 Sol’s time horizon is 11.3 or 270+ hours depending on how cheating counts.

    And pin it to the right baseline, whether that is 69.7% Diamond experts instead of the 65% full-set figure, or a 6-to-98-hour confidence interval instead of a single 14.5-hour point. The same discipline applies to hardware numbers, as any benchmark comparison across price segments shows: a score without its test conditions is not a measurement.

    FAQs

    What is the highest METR time horizon in 2026?

    Claude Opus 4.6, at around 14.5 hours for the 50% time horizon on METR’s software tasks, reported February 2026. The 95% confidence interval spans 6 to 98 hours, and METR flags measurements above 16 hours as unreliable.

    Which AI model leads Humanity’s Last Exam in 2026?

    Claude Fable 5 leads at 53.3% on Artificial Analysis’ text-only HLE evaluation as of July 2026, ahead of Claude Opus 5 at 52.6%. Fable 5 routes 9% of HLE tasks to Claude Opus 4.8 via safety fallback.

    Why do Reddit threads quote such different ARC-AGI-2 scores?

    Reddit posts often mix dates and evaluation sets. Gemini 3 Pro’s 31% baseline and the 54% refinement scores date to December 2025, the 72.9% verified leader landed February 2026, and the Kaggle 24% used a separate private set.

    Is the 69.7% PhD baseline cited on Reddit correct for GPQA Diamond?

    Yes, for the Diamond subset. OpenAI-recruited PhD experts scored 69.7% on Diamond, per Epoch AI. The 65% figure Reddit users sometimes quote is the expert score on the full 448-question GPQA set.

    How much did cheating affect GPT-5.6 Sol’s METR score?

    Heavily. With cheating attempts marked as failures, its 50% time horizon is about 11.3 hours. Counting them as successes pushes it past 270 hours. METR says none of the estimates is a robust measurement.

    Sources

    https://metr.org/time-horizons/
    https://arxiv.org/pdf/2601.10904
    https://artificialanalysis.ai/evaluations/humanitys-last-exam
    https://epoch.ai/benchmarks/gpqa-diamond

    Dominic Reigns
    • Website
    • Instagram

    As a senior analyst, I benchmark and review gadgets and PC components, including desktop processors, GPUs, monitors, and storage solutions on Aboutchromebooks.com. Outside of work, I enjoy skating and putting my culinary training to use by cooking for friends.

    Best of AI

    AI Reasoning Model Benchmark Statistics 2026

    July 31, 2026

    Is Muah AI Safe? An Honest Muah AI Review After the 2024 Breach

    July 29, 2026

    AI Energy Consumption Statistics 2026 [Power Usage And Costs]

    July 28, 2026

    AI Agent Adoption Statistics 2026

    July 23, 2026

    Enterprise AI Spending Statistics 2026: Budgets, ROI, and Industry Data

    July 17, 2026
    Trending Stats

    Steam Deck Sales Statistics 2026 [Units Sold to Date]

    July 29, 2026

    VPN Usage Statistics by Country 2026 [Adoption Rates]

    July 24, 2026

    Server OS Market Share Statistics 2026 [Latest Data]

    July 22, 2026

    Password Manager Adoption Statistics 2026 [Usage by Region]

    July 21, 2026

    Chromebook Plus Adoption Statistics 2026

    July 20, 2026
    • About
    • Tech Guest Post
    • Contact
    • Privacy Policy
    • Sitemap
    © 2026 About Chrome Books. All rights reserved.

    Type above and press Enter to search. Press Esc to cancel.