Skip to content
BetaBenchAlert v0.1 is in Beta.Numbers are real, but pages and rules can still change.See what changed

Side by side

Compare two models

Put two models side by side. Same tests, same units. Each row names its leader.

The short answer · test kit v0.2

Grok 4.5 completed the fixed check 1.8× faster than GPT-5.6 Sol.

Grok 4.5
38.2answer tok/s · 6.26s total
Opus 5
33.5answer tok/s · 7.21s total
GPT-5.6 Sol
21.9answer tok/s · 11.1s total

Speed is one question. Check the answered-when-scheduled row too — a quick model that keeps failing is not the quicker tool in practice.

Pick the models

Your picks live in the link. Copy it and the matchup travels with it.

Swap the two columns

Row by row

The marked cell leads its row. An empty typical result means fewer than 12 completed checks in this window — not zero.

RowGPT-5.6 SolOpus 5Grok 4.5
Write speedvisible answer tokens per second · hidden thinking does not raise it · higher is better21.9answer tok/s33.5answer tok/s38.2leadsanswer tok/s
Time to first wordhow long the screen stays blank · Codex does not stream, so its cell stays empty · lower is better2.72sleads3.10s
Typical waitstart to finish for the whole command · lower is better11.1s7.21s6.26sleads
Answered when scheduledcompleted correct answers out of all scheduled checks in 24 hours · higher is better21/24 (88%)24/24 (100%)leads24/24 (100%)

Every row covers the last 24 hours. Test kit v0.2 rows only.

How to read this

Add a third model to the web address with ?c=sonnet. Every number follows the same rules as the home board: current or raw-recovered compatible speed rows, right answers only, one try per round. Units are spelled out on How we test.

Matchups worth opening

One click loads the whole card.

Read next

GPT-5.6 Sol pageOpus 5 pageThe model boardHow we test