17 accounts, 162 games: what actually makes an EloCoach review slow
While building the Experiments page I kept making claims about speed without numbers behind them. “A review takes about twenty seconds.” “Running things in parallel is faster.” So I stopped guessing and ran a baseline: real games from real, public players, through the real app on a real phone, and the coach on top. This is what the numbers say, including the parts that surprised me and the parts I cannot honestly generalise.
The setup
Players: 9 Chess.com accounts and 8 Lichess accounts, mostly titled players and streamers (GothamChess, Chessbrah, GM Benjamin Bok, GM Benjamin Finegold, Eric Rosen, Magnus Carlsen, Alireza Firouzja, Andrew Tang, Nihal Sarin, Sergei Zhigalko and others) plus one beginner-level account rated 600 to 1000. They are public figures playing in public, and I only read games they had already published.
Games: the newest 15 rated rapid or blitz games per player, played before 6 October 2026. That produced 162 games.
Engine side: a Pixel 9 running the release build. Stockfish 19 at depth 12 plus Maia 3, the model that predicts what a human at a given rating would play. All of it on the phone.
Coach side: one language model call per game (Gemini through OpenRouter), made from my PC against a local copy of the backend.
Surprise one: elite Lichess players play bullet
The same rule gave 117 Chess.com games and only 45 from Lichess. Our sync reads a player’s most recent hundred or so Lichess games, and for Tang, Zhigalko, Sarin and Firouzja, all 108 of them were bullet. Carlsen’s were 106 bullet and 2 blitz. The strongest players on Lichess are playing a format where a game lasts a minute or two, which is wonderful to watch and close to useless for a review coach. It was a useful correction to my assumption that “grandmaster data” would be easy to get. The sync side of this is in the fetching post.
The engine pass tracks game length, and almost nothing else
The median on-device analysis took 18.7 seconds, with a 90th percentile of 29.9. About 227 milliseconds per move. One game was left out because the app went to the background while it ran and it took 228 seconds, which is a different story about phones, not engines.
One dot per game, 161 games. Hover a dot for its exact numbers. Lichess games are shorter, so they sit lower left.
The correlation with game length is 0.85. Key moments found (0.35) and player rating (0.37) matter much less, and that is reassuring: the engine is not secretly spending longer on strong players. 67 of the games involve a rating above 2600, which is Maia’s ceiling, so those are read as a 2600 player would see them. That is a limit of the model, not the phone, and I would rather say it than hide it.
The coach is slow because it writes, not because it reads
This was the clearest result. A review sends the model about 25,000 tokens (the game, the engine findings, your memory) and gets back about 4,200. Prompt size grows with game length almost perfectly (0.99). You would expect a longer game to mean a longer wait. It does not.
One dot per review, 162 reviews. The wait is almost entirely how much the model writes.
Total time correlates 0.97 with output tokens and 0.04 with prompt size. The median review took 21.6 seconds, and the first visible text arrived after 12.1. If I want a faster coach, the lever is a tighter answer, not a smaller game. It also means the engine and the coach are independent: engine time and coach time correlate at 0.10, so a slow phone does not make for a slow coach and vice versa.
Engine time follows game length. Coach time follows the length of its own answer. Amber is the engine, blue is the coach.
What changes when many reviews run at once
I ran the coach one review at a time, then four workers, then eight. Four workers finished 97 reviews in roughly 13 minutes, about four times faster overall, with the median per review barely moving (20.4 to 21.8 seconds). The cost showed up in the tail: the 90th percentile went from 25.3 to 36.5 seconds, and first visible text from 16.6 to 26.9.
Grouped by how many other calls overlapped each review at any point in its life. Sample sizes are small in the middle buckets, so read the shape, not each bar.
Up to seven overlapping calls the median stays between 19 and 24 seconds. At eight or more it jumps to 36. That is the number I would plan around.
The eight-worker run stopped for a reason I did not expect, and it was not speed. It stopped on credit. OpenRouter reserves the full cost of each in-flight request up front. With eight at once, those reservations added up to more than the account balance, and requests started returning HTTP 402. Three of those in a row and the fallback logic marks the model as degraded and fails the calls after it. In production the real ceiling on concurrency is not how fast the model answers, it is how much headroom the balance has.
Does the coach get the board right?
I ran an automatic check over every “piece on square” claim in the 162 reviews, replaying each one against every position in the game. 1,081 claims, 6 flagged. Four of the six were false alarms: general opening theory, a suggested move that was never played, and a correct promotion line inside a variation that I verified by hand. Two were real, both wrong-square slips: a fork said to hit “the king on h1” when it stood on g2 (the fork itself was real), and “keeping your pawn on e6” when it was never there. That is about 0.2 percent. It is the same kind of mistake I wrote about in the post on invented moves, now measured at scale instead of found by accident.
I also ran seven reviews through Apollo, the second coach. They wrote more (a median of about 7,200 tokens against 4,200) and took longer (a median of 31.8 seconds), and the 20 board claims in them were all correct. Seven reviews is a sample that proves nothing, and the run stopped there, so treat that as a hint, not a result.
What this does not tell you
One phone, one PC. A Pixel 9 is a recent phone. An older Android will be slower on the engine pass, and I have not measured one.
A local backend. The coach calls went through a local copy running on an embedded database that handles one query at a time, not production Postgres. The load numbers are real for that setup and probably pessimistic for production.
Almost all blitz. No rapid or classical games came out of the newest-games rule, so I have nothing on how long, slower games behave.
Strong players, mostly. The player pool skews toward titled players. It is useful for stress-testing, less so for knowing what a 1200 player’s game looks like.