Skip to content
BetaBenchAlert v0.1 is in Beta.Numbers are real, but pages and rules can still change.See what changed

Real tests of the AI coding tools you pay for.

Claude Code, Codex, and Grok Build — timed on the same paid plans you buy. How fast they write, and whether the answer is right. For people on a subscription, not an API key.

BenchAlert Speed

Live43m ago

modelefforttok/stotal
  1. Grok 4.6xhigh effort36.02 tok/s6.63s total
  2. Opus 5high effort35.21 tok/s6.84s total
  3. Fable 5rate limit warninghigh effort28.60 tok/s8.43s total
  4. GPT-5.6 Solhigh effort19.60 tok/s12.4s total
Measured
4:01 PM
Timezone
ET
Test kit
v0.2

Models measured
10
Every round
Rounds a day
24
One every hour
Answers served
98%
234 of 240 attempted checks
Rate limits
0
In the last 24 h

Quota

What 100% of a weekly limit is worth at the vendor's own API prices.

Full board →
QuotaAPI $ / week

Liveupdated 17h agonext run 03:10 UTC

  1. Claude Max 20x≥ lower bound≥ $1,571

    5-hour $2.20 · Fable 5 up to 50% of weekly

  2. SuperGrok Heavy≥ lower bound≥ $1,201

    Includes Cursor Ultra — $400/mo usage on top

  3. ChatGPT Pro ($100, 5×)≥ lower bound≥ $604

    Vendor's own figure $600 (15,000 credits)

Last reading
Aug 23
Method
v1
Interval
calibrated

Highlights

The last 24 hours, made simple

Live speed is in the hero. Here we step back and ask two steadier questions: what was typical, and how often did each scheduled check return an answer?

Typical writing speed

Visible answer tokens per second, last 24 h · higher is faster

  1. Grok 4.66.44s total wait37.09
  2. Opus 57.21s total wait33.45
  3. Fable 57.90s total wait30.52
  4. GPT-5.6 Sol11.2s total wait21.74
037.09 tok/s

4 of 10 models

Other models

Answered when scheduled

Completed answers out of all scheduled checks, last 24 h · higher is better

  1. Grok 4.624/24 answered100%
  2. Opus 524/24 answered100%
  3. Fable 524/24 answered100%
  4. GPT-5.6 Sol23/24 answered96%
0%100%

4 of 10 models

Lab status

Both boards cover the last 24 hours. Each board has its own scale · test kit v0.2. 4 of 10 pinned models shown · the rest live on Other models.

Speed

Typical speed and total wait

The 4 models we lead with, ranked by typical total wait over the last 24 hours. At least 12 completed checks are required.

Other models

Writing speed counts only the answer you can see. Total wait measures click-to-finish time. The table ranks the shorter total wait first.

The 4 models on the home boards, ranked by typical start-to-finish time over the last 24 hours, with visible writing speed beside it.
RankModelBar, on one shared scaleWriting speedTotal waitvs yesterdayAnsweredCoverage
1Grok 4.6xAI · effort xhigh37.096.44slikely 6.27s6.89s+3%100%24/24
2Opus 5Anthropic · effort high33.457.21slikely 6.78s7.69s-2%100%24/24
3Fable 5Anthropic · effort high · last round rate limit warning30.527.90slikely 7.34s8.24s-3%100%24/24
4GPT-5.6 SolOpenAI · effort high · last round failed21.7411.2slikely 10.6s12.0s-4%96%24/24

Window 24 hours · 4 of 4 models have 12 completed checks · 96 checks attempted · test kit v0.2. A dash under vs yesterday means that model did not have 12 completed checks the day before. 4 of 10 pinned models shown · the rest live on Other models.

24 hours

The last 24 hours, round by round

Each line joins one model's rounds, one an hour. Every dot is a real round, so the swings stay in plain sight. Point at the plot to read any round.

Open today's report

Visible answer tokens per second · higher is faster

Model writing speed over time

Writing speed (tokens / second)020406080
6pm12am6am12pmnow
Aug 22Aug 23

Lab time (ET)

Line: one model's visible writing speed, check to check. Dots: the checks themselves. A single missed hour is stepped over by a faint dotted link; anything longer breaks the line. Nothing is invented to fill a gap.

Full record — every round, every number
Visible answer tokens per second for every check in the last 24 hours. A dash means that check had no completed answer.
RoundGrok 4.6Opus 5Fable 5GPT-5.6 Sol
Aug 23 4:00pm36.0235.2128.60
Aug 23 3:00pm39.1533.2426.3319.60
Aug 23 2:00pm28.0133.3030.5322.23
Aug 23 1:00pm40.0834.5831.0715.38
Aug 23 12:00pm29.5733.4724.0520.31
Aug 23 11:00am39.1729.4532.8123.15
Aug 23 10:00am37.1735.5626.1822.92
Aug 23 9:00am14.8330.9630.0020.17
Aug 23 8:00am28.7331.0530.4919.51
Aug 23 7:00am37.8328.5531.6120.57
Aug 23 6:00am34.6938.9730.2123.01
Aug 23 5:00am37.0131.8029.2421.88
Aug 23 4:00am37.6532.7729.4921.18
Aug 23 3:00am36.0928.1530.5124.59
Aug 23 2:00am32.5331.3431.6522.60
Aug 23 1:00am37.4122.7828.5224.58
Aug 23 12:00am38.1333.4331.4921.93
Aug 22 11:00pm44.2034.1528.2620.04
Aug 22 10:00pm31.4844.2051.0623.35
Aug 22 9:00pm41.2051.4850.6820.14
Aug 22 8:00pm37.0247.1535.1218.01
Aug 22 7:00pm35.2233.9855.2021.74
Aug 22 6:00pm44.0847.0143.2020.72
Aug 22 5:00pm37.4153.7545.0723.90

Window 24 hours · one round every 60 minutes · 24 rounds on the clock · n = 95 completed readings drawn · test kit v0.2. 4 of 10 pinned models shown · the rest live on Other models.

We also ping each app every round to check it answers at all. Those checks, round by round, live on Lab status.

Method

How we get these numbers

We run the real paid apps on our own lab machine, not the hidden APIs. Every number traces back to a saved run.

Open How we test
Cadence
Every hour
Rounds a day
24
Test kit
v0.2
Window
24hours
Test kit v0.2
Old rows stay in Logs. They do not mix into these boards.
Changelog
Paid apps
We run Claude Code, Codex, and Grok Build ourselves. Not the hidden APIs.
Status
Graded answers
A round counts only when the answer is right. Not “it printed text.”
How we test
Saved proof
A public number traces back to a saved run in our archive.
Logs

Last round 34m ago · 312 results saved in the last 24 hours, 8 failed · Lab status.