Skip to content
BetaBenchAlert v0.1 is in Beta.Numbers are real, but pages and rules can still change.See what changed

Codex

GPT-5.6 Terra

Inside Codex we ask for gpt-5.6-terra every hour. Everything below comes from those runs on our lab machine.

Codexgpt-5.6-terraThink setting highLast round 3h agoOther models

Write speed, last 24 hours

23.5answer tok/s

10.4s typical total wait

Likely middle: 9.40s11.2s

Visible answer tokens per second across the whole command. Hidden thinking does not make this number larger. Higher is faster. These are the middle values of 19 completed checks out of 24 scheduled in the last 24 hours.

Time to first word
This app does not stream text, so we cannot time the first word.
Typical total wait
10.4s
How long the whole command took, start to finish, including the app booting. Lower is faster.
Last ping
The newest tiny “reply ok” check. It measures wake-up time, not writing speed.

What stands out

Four facts pulled straight from the saved runs.

Board position
#7 of 10
Just ahead: Haiku 4.5 at 8.67s total wait. Just behind: GPT-5.5 at 10.4s total wait.
Answered when scheduled
19/24
Completed answers out of all scheduled checks in the last 24 hours.
Served as
gpt-5.6-terra
The model name the app reported back. If it differs from the pin, we say so.
Round-to-round swing
1.6×
How much the fastest round beat the slowest one. A big swing means a noisy day.

Recent rounds

Newest first. Fails stay in the list with their reason.

WhenResultWriting speedTotal wait
1m agofailedno tool items
1h agofailedno tool items
2h agofailedno tool items
3h agookno tool items23.3 tok/s10.5s
4h agookno tool items25.1 tok/s9.68s
5h agowrong answerno tool items
6h agookno tool items21.5 tok/s11.3s
6h agookno tool items22.0 tok/s11.0s

All rounds are in Logs.

Read next

Compare with Sonnet 5

Its closest rival in Claude Code — same window, same units, with the leader marked on every row.

Compare with GPT-5.6 Sol

The pin next to it in Codex. See what the tier jump buys inside one app.

More models

How to read this

Every number here comes from our lab machine running the paid app, one try per hourly round. The live check uses test kit v0.2; recovered history appears only after its raw file passes the same visible-speed rules. Our measurements, not the vendor’s.