# BenchAlert BenchAlert is an independent benchmarking lab. We time Claude Code, Codex, and Grok Build — the paid command-line apps people type into. Not the hidden HTTP APIs. We run every test ourselves, on our own dedicated lab machine, every hour, on subscriptions we pay for. This site publishes the results. Readers do not install, run, or host anything. Test kit v0.2. Boards use current rows plus older speed rows only when their saved raw files replayed cleanly and the visible-speed value could be rebuilt exactly. One lab machine. These exact flags. Every published number traces back to a saved CLI file in our archive. These are our own measurements. They are not official numbers from Anthropic, OpenAI, or xAI. ## For people - https://benchalert.com/ — live pulse board - https://benchalert.com/start — how to read the boards - https://benchalert.com/today — the last 24 hours in one card - https://benchalert.com/compare — two models side by side - https://benchalert.com/models — one page per paid model - https://benchalert.com/design — one landing-page brief, seven models, every page they wrote - https://benchalert.com/method — exact prompts, flags, and units - https://benchalert.com/quota — what a full week of each paid plan is worth at the vendor's own API list price - https://benchalert.com/logs — every sample and what our lab spent - https://benchalert.com/status — is our lab healthy - https://benchalert.com/changelog — every change to how we test - https://benchalert.com/about — who runs the lab - https://benchalert.com/teams — using the data before you buy seats - https://benchalert.com/developers — the JSON feeds - https://benchalert.com/privacy - https://benchalert.com/terms ## For machines - https://benchalert.com/api/today - https://benchalert.com/api/speed?hours=24 - https://benchalert.com/api/probes?hours=24 - https://benchalert.com/api/runs?limit=100 - https://benchalert.com/api/status - https://benchalert.com/api/quota?days=90 - https://benchalert.com/badge - https://benchalert.com/feed.xml ## Units - Live speed in the hero is the latest completed check only. It is never a median or average. - Answer tok/s counts only visible answer tokens over the full start-to-finish time. Hidden thinking cannot inflate it. - Total wait is click-to-finish latency. Typical boards rank the shorter median wait first and require 12 completed checks. - Ping is a tiny "ok" check. It is not a speed test. - A served round is one where the app completed the correct answer. Rate-limit-warning rounds stay in speed summaries because the reader really experienced that wait. - Quota (Q) is the dollar value of 100% of a paid plan's usage window, priced at the vendor's published API list price. It is measured on our own accounts, never anyone else's, and it is a floor: the meter counts devices we cannot see. - User reports (when enabled) are visitor votes about how a tool feels. They are opinion, kept apart from lab numbers. ## If you quote us Say the number came from BenchAlert, name the test kit version, and link the page you took it from.