Real tests of the AI coding tools you pay for.
Claude Code, Codex, and Grok Build — timed on the same paid plans you buy. How fast they write, and whether the answer is right. For people on a subscription, not an API key.
Live43m ago
- Grok 4.6xhigh effort36.02 tok/s6.63s total
- Opus 5high effort35.21 tok/s6.84s total
- Fable 5rate limit warninghigh effort28.60 tok/s8.43s total
- GPT-5.6 Solhigh effort19.60 tok/s12.4s total
- Measured
- 4:01 PM
- Timezone
- ET
- Test kit
- v0.2
—
- Models measured
- 10
- Every round
- Rounds a day
- 24
- One every hour
- Answers served
- 98%
- 234 of 240 attempted checks
- Rate limits
- 0
- In the last 24 h
Quota
What 100% of a weekly limit is worth at the vendor's own API prices.
Liveupdated 17h agonext run 03:10 UTC
- Claude Max 20x≥ lower bound≥ $1,571
5-hour $2.20 · Fable 5 up to 50% of weekly
- SuperGrok Heavy≥ lower bound≥ $1,201
Includes Cursor Ultra — $400/mo usage on top
- ChatGPT Pro ($100, 5×)≥ lower bound≥ $604
Vendor's own figure $600 (15,000 credits)
- Last reading
- Aug 23
- Method
- v1
- Interval
- calibrated
Highlights
The last 24 hours, made simple
Live speed is in the hero. Here we step back and ask two steadier questions: what was typical, and how often did each scheduled check return an answer?
Typical writing speed
Visible answer tokens per second, last 24 h · higher is faster
4 of 10 models
Other modelsAnswered when scheduled
Completed answers out of all scheduled checks, last 24 h · higher is better
4 of 10 models
Lab statusBoth boards cover the last 24 hours. Each board has its own scale · test kit v0.2. 4 of 10 pinned models shown · the rest live on Other models.
Speed
Typical speed and total wait
The 4 models we lead with, ranked by typical total wait over the last 24 hours. At least 12 completed checks are required.
Writing speed counts only the answer you can see. Total wait measures click-to-finish time. The table ranks the shorter total wait first.
| Rank | Model | Bar, on one shared scale | Writing speed | Total wait | vs yesterday | Answered | Coverage |
|---|---|---|---|---|---|---|---|
| 1 | Grok 4.6xAI · effort xhigh | 37.09 | 6.44slikely 6.27s–6.89s | +3% | 100% | 24/24 | |
| 2 | Opus 5Anthropic · effort high | 33.45 | 7.21slikely 6.78s–7.69s | -2% | 100% | 24/24 | |
| 3 | Fable 5Anthropic · effort high · last round rate limit warning | 30.52 | 7.90slikely 7.34s–8.24s | -3% | 100% | 24/24 | |
| 4 | GPT-5.6 SolOpenAI · effort high · last round failed | 21.74 | 11.2slikely 10.6s–12.0s | -4% | 96% | 24/24 |
Window 24 hours · 4 of 4 models have 12 completed checks · 96 checks attempted · test kit v0.2. A dash under vs yesterday means that model did not have 12 completed checks the day before. 4 of 10 pinned models shown · the rest live on Other models.
24 hours
The last 24 hours, round by round
Each line joins one model's rounds, one an hour. Every dot is a real round, so the swings stay in plain sight. Point at the plot to read any round.
Visible answer tokens per second · higher is faster
Model writing speed over time
Lab time (ET)
Line: one model's visible writing speed, check to check. Dots: the checks themselves. A single missed hour is stepped over by a faint dotted link; anything longer breaks the line. Nothing is invented to fill a gap.
Full record — every round, every number
| Round | Grok 4.6 | Opus 5 | Fable 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| Aug 23 4:00pm | 36.02 | 35.21 | 28.60 | — |
| Aug 23 3:00pm | 39.15 | 33.24 | 26.33 | 19.60 |
| Aug 23 2:00pm | 28.01 | 33.30 | 30.53 | 22.23 |
| Aug 23 1:00pm | 40.08 | 34.58 | 31.07 | 15.38 |
| Aug 23 12:00pm | 29.57 | 33.47 | 24.05 | 20.31 |
| Aug 23 11:00am | 39.17 | 29.45 | 32.81 | 23.15 |
| Aug 23 10:00am | 37.17 | 35.56 | 26.18 | 22.92 |
| Aug 23 9:00am | 14.83 | 30.96 | 30.00 | 20.17 |
| Aug 23 8:00am | 28.73 | 31.05 | 30.49 | 19.51 |
| Aug 23 7:00am | 37.83 | 28.55 | 31.61 | 20.57 |
| Aug 23 6:00am | 34.69 | 38.97 | 30.21 | 23.01 |
| Aug 23 5:00am | 37.01 | 31.80 | 29.24 | 21.88 |
| Aug 23 4:00am | 37.65 | 32.77 | 29.49 | 21.18 |
| Aug 23 3:00am | 36.09 | 28.15 | 30.51 | 24.59 |
| Aug 23 2:00am | 32.53 | 31.34 | 31.65 | 22.60 |
| Aug 23 1:00am | 37.41 | 22.78 | 28.52 | 24.58 |
| Aug 23 12:00am | 38.13 | 33.43 | 31.49 | 21.93 |
| Aug 22 11:00pm | 44.20 | 34.15 | 28.26 | 20.04 |
| Aug 22 10:00pm | 31.48 | 44.20 | 51.06 | 23.35 |
| Aug 22 9:00pm | 41.20 | 51.48 | 50.68 | 20.14 |
| Aug 22 8:00pm | 37.02 | 47.15 | 35.12 | 18.01 |
| Aug 22 7:00pm | 35.22 | 33.98 | 55.20 | 21.74 |
| Aug 22 6:00pm | 44.08 | 47.01 | 43.20 | 20.72 |
| Aug 22 5:00pm | 37.41 | 53.75 | 45.07 | 23.90 |
Window 24 hours · one round every 60 minutes · 24 rounds on the clock · n = 95 completed readings drawn · test kit v0.2. 4 of 10 pinned models shown · the rest live on Other models.
We also ping each app every round to check it answers at all. Those checks, round by round, live on Lab status.
Method
How we get these numbers
We run the real paid apps on our own lab machine, not the hidden APIs. Every number traces back to a saved run.
- Cadence
- Every hour
- Rounds a day
- 24
- Test kit
- v0.2
- Window
- 24hours
Last round 34m ago · 312 results saved in the last 24 hours, 8 failed · Lab status.