← All posts

Experiments9 min read

By Benjamin from EloCoach

17 accounts, 162 games: what actually makes an EloCoach review slow

While building the Experiments page I kept making claims about speed without numbers behind them. “A review takes about twenty seconds.” “Running things in parallel is faster.” So I stopped guessing and ran a baseline: real games from real, public players, through the real app on a real phone, and the coach on top. This is what the numbers say, including the parts that surprised me and the parts I cannot honestly generalise.

The setup

  • Players: 9 Chess.com accounts and 8 Lichess accounts, mostly titled players and streamers (GothamChess, Chessbrah, GM Benjamin Bok, GM Benjamin Finegold, Eric Rosen, Magnus Carlsen, Alireza Firouzja, Andrew Tang, Nihal Sarin, Sergei Zhigalko and others) plus one beginner-level account rated 600 to 1000. They are public figures playing in public, and I only read games they had already published.
  • Games: the newest 15 rated rapid or blitz games per player, played before 6 October 2026. That produced 162 games.
  • Engine side: a Pixel 9 running the release build. Stockfish 19 at depth 12 plus Maia 3, the model that predicts what a human at a given rating would play. All of it on the phone.
  • Coach side: one language model call per game (Gemini through OpenRouter), made from my PC against a local copy of the backend.

Surprise one: elite Lichess players play bullet

The same rule gave 117 Chess.com games and only 45 from Lichess. Our sync reads a player’s most recent hundred or so Lichess games, and for Tang, Zhigalko, Sarin and Firouzja, all 108 of them were bullet. Carlsen’s were 106 bullet and 2 blitz. The strongest players on Lichess are playing a format where a game lasts a minute or two, which is wonderful to watch and close to useless for a review coach. It was a useful correction to my assumption that “grandmaster data” would be easy to get. The sync side of this is in the fetching post.

The engine pass tracks game length, and almost nothing else

The median on-device analysis took 18.7 seconds, with a 90th percentile of 29.9. About 227 milliseconds per move. One game was left out because the app went to the background while it ran and it took 228 seconds, which is a different story about phones, not engines.

01020304004080120160game length in plies (half-moves)on-device analysis, seconds59 plies, 22.2 s73 plies, 15.1 s21 plies, 5.2 s150 plies, 40.3 s74 plies, 20.3 s65 plies, 16 s75 plies, 23.8 s68 plies, 15.2 s61 plies, 18.2 s83 plies, 26.3 s86 plies, 17.7 s87 plies, 16.6 s114 plies, 22.1 s106 plies, 24.5 s108 plies, 26.3 s54 plies, 11.1 s101 plies, 20.3 s71 plies, 17.3 s94 plies, 18.4 s52 plies, 13.3 s61 plies, 17.3 s102 plies, 18.4 s63 plies, 15.8 s69 plies, 13.1 s98 plies, 16.9 s37 plies, 7.9 s110 plies, 20.6 s81 plies, 15 s111 plies, 21.3 s56 plies, 11.6 s58 plies, 12.9 s74 plies, 16.3 s97 plies, 19.7 s10 plies, 4.3 s135 plies, 28.4 s54 plies, 12.9 s55 plies, 15.3 s73 plies, 22 s29 plies, 9.4 s122 plies, 26.2 s65 plies, 20.6 s7 plies, 2.7 s51 plies, 12.1 s126 plies, 24 s65 plies, 24.2 s142 plies, 29 s105 plies, 35.5 s180 plies, 31.9 s99 plies, 32.9 s139 plies, 31.5 s23 plies, 9.9 s30 plies, 12.5 s156 plies, 36.6 s84 plies, 18.7 s80 plies, 24 s86 plies, 18.4 s57 plies, 15.5 s59 plies, 21.3 s146 plies, 34.1 s40 plies, 11.1 s70 plies, 24 s93 plies, 24.3 s109 plies, 19.4 s116 plies, 29.3 s113 plies, 21.2 s157 plies, 36.2 s115 plies, 19.9 s63 plies, 16.9 s66 plies, 20.9 s83 plies, 30.6 s62 plies, 17.9 s141 plies, 28.3 s55 plies, 14.2 s87 plies, 26.7 s90 plies, 20.8 s119 plies, 24.9 s67 plies, 18 s53 plies, 14.4 s113 plies, 23.6 s133 plies, 32.5 s104 plies, 39.7 s98 plies, 22.3 s70 plies, 18.9 s84 plies, 29.2 s71 plies, 16.9 s41 plies, 15.5 s165 plies, 30.1 s70 plies, 29.9 s48 plies, 11.6 s45 plies, 10.1 s114 plies, 27.2 s124 plies, 25.1 s12 plies, 4.5 s35 plies, 10.4 s67 plies, 20.4 s178 plies, 39.8 s79 plies, 18.6 s109 plies, 19.6 s133 plies, 24.4 s84 plies, 19.1 s162 plies, 43 s83 plies, 16.6 s123 plies, 29.2 s88 plies, 26.5 s89 plies, 19.3 s82 plies, 17.5 s30 plies, 19.4 s92 plies, 25.9 s121 plies, 22.4 s105 plies, 19.9 s183 plies, 32.9 s68 plies, 12.7 s68 plies, 16.1 s105 plies, 20 s57 plies, 11.6 s99 plies, 19.1 s99 plies, 18.9 s39 plies, 13 s54 plies, 10.6 s89 plies, 20.5 s42 plies, 9.2 s64 plies, 13 s57 plies, 12.4 s113 plies, 19.7 s49 plies, 10.2 s69 plies, 13.1 s66 plies, 12.8 s110 plies, 20 s86 plies, 16.5 s53 plies, 10.9 s36 plies, 8.3 s20 plies, 4.7 s24 plies, 5.7 s18 plies, 6.4 s18 plies, 4.3 s126 plies, 23.1 s63 plies, 12.1 s42 plies, 8.1 s99 plies, 18.2 s16 plies, 4.3 s73 plies, 14 s43 plies, 9.2 s122 plies, 20.5 s14 plies, 3.7 s89 plies, 18.1 s136 plies, 23.1 s120 plies, 33.5 s41 plies, 8.9 s90 plies, 23.2 s164 plies, 29.4 s103 plies, 27.9 s79 plies, 16.6 s46 plies, 13.4 s75 plies, 13.8 s55 plies, 11.3 s123 plies, 24.4 s49 plies, 17.4 s82 plies, 17.9 s116 plies, 21 s55 plies, 12.3 s124 plies, 21.7 sChess.comLichess
One dot per game, 161 games. Hover a dot for its exact numbers. Lichess games are shorter, so they sit lower left.

The correlation with game length is 0.85. Key moments found (0.35) and player rating (0.37) matter much less, and that is reassuring: the engine is not secretly spending longer on strong players. 67 of the games involve a rating above 2600, which is Maia’s ceiling, so those are read as a 2600 player would see them. That is a limit of the model, not the phone, and I would rather say it than hide it.

The coach is slow because it writes, not because it reads

This was the clearest result. A review sends the model about 25,000 tokens (the game, the engine findings, your memory) and gets back about 4,200. Prompt size grows with game length almost perfectly (0.99). You would expect a longer game to mean a longer wait. It does not.

01020304050030006000900012000tokens the coach wrotewhole review, seconds7831 tokens written, 34.2 s4092 tokens written, 20.4 s10659 tokens written, 40.7 s3351 tokens written, 20.1 s3230 tokens written, 15.6 s4769 tokens written, 23.1 s3245 tokens written, 19.1 s3887 tokens written, 21.3 s8799 tokens written, 40.9 s8226 tokens written, 36 s5127 tokens written, 24.3 s8368 tokens written, 36.5 s5159 tokens written, 25.1 s3550 tokens written, 19.1 s3141 tokens written, 19.1 s4513 tokens written, 21.6 s2226 tokens written, 15.6 s7493 tokens written, 33 s3792 tokens written, 18.9 s10104 tokens written, 40.6 s5333 tokens written, 25.9 s6311 tokens written, 29.1 s3596 tokens written, 18.4 s4260 tokens written, 20.1 s3612 tokens written, 18.4 s4172 tokens written, 20.8 s2953 tokens written, 19.7 s4083 tokens written, 19.5 s4472 tokens written, 20.4 s3241 tokens written, 17.6 s4144 tokens written, 24.6 s3790 tokens written, 20.2 s3058 tokens written, 17.9 s6084 tokens written, 30 s11122 tokens written, 48.9 s6906 tokens written, 34.6 s5225 tokens written, 28.2 s4420 tokens written, 21.8 s6830 tokens written, 30.3 s3520 tokens written, 18.4 s7221 tokens written, 31 s2634 tokens written, 14.9 s5507 tokens written, 25.7 s3693 tokens written, 20 s6536 tokens written, 30 s7101 tokens written, 31.3 s4509 tokens written, 22.4 s4062 tokens written, 19.6 s4325 tokens written, 21.5 s11357 tokens written, 50.5 s8752 tokens written, 41.6 s4585 tokens written, 24.9 s5339 tokens written, 25.9 s3030 tokens written, 21.7 s3783 tokens written, 20.6 s4046 tokens written, 18.9 s9450 tokens written, 40.1 s7610 tokens written, 36.1 s4606 tokens written, 23.2 s3934 tokens written, 20.9 s3356 tokens written, 17.4 s4783 tokens written, 22.2 s3147 tokens written, 15.6 s3251 tokens written, 19.6 s4481 tokens written, 24.3 s11555 tokens written, 50.9 s5685 tokens written, 26.1 s3776 tokens written, 18.6 s3700 tokens written, 20.1 s10374 tokens written, 43.3 s3425 tokens written, 18.9 s6846 tokens written, 29.5 s7197 tokens written, 33.8 s4820 tokens written, 22.2 s4196 tokens written, 19.7 s4321 tokens written, 22.9 s5191 tokens written, 23.9 s4260 tokens written, 21.8 s3128 tokens written, 16.4 s3495 tokens written, 18 s5869 tokens written, 27 s4957 tokens written, 24.7 s3664 tokens written, 19.6 s3858 tokens written, 19.4 s3325 tokens written, 16.9 s3119 tokens written, 17.2 s3707 tokens written, 19.2 s4993 tokens written, 26.5 s2610 tokens written, 15.9 s4647 tokens written, 29.8 s3350 tokens written, 17.8 s2805 tokens written, 15.8 s4714 tokens written, 21 s3229 tokens written, 16.3 s4196 tokens written, 20.5 s5119 tokens written, 27.2 s2966 tokens written, 16.1 s3750 tokens written, 24.6 s3818 tokens written, 18.6 s3255 tokens written, 17.6 s3884 tokens written, 19.9 s8433 tokens written, 39.7 s4150 tokens written, 20.8 s3026 tokens written, 16.4 s4352 tokens written, 22.1 s4103 tokens written, 23.5 s5337 tokens written, 27.2 s2965 tokens written, 16.1 s4286 tokens written, 24.2 s3449 tokens written, 18 s4168 tokens written, 25.3 s10309 tokens written, 43.3 s4615 tokens written, 21.6 s4155 tokens written, 21.1 s5357 tokens written, 24.7 s3463 tokens written, 18.6 s4110 tokens written, 21.7 s3443 tokens written, 19.2 s3115 tokens written, 22.1 s11778 tokens written, 50.7 s9642 tokens written, 45 s7774 tokens written, 40.1 s3498 tokens written, 17.5 s3399 tokens written, 20.4 s4508 tokens written, 23.9 s3040 tokens written, 16.7 s3110 tokens written, 18.1 s8584 tokens written, 41 s3312 tokens written, 19.6 s2920 tokens written, 21.3 s3898 tokens written, 20.1 s2990 tokens written, 17.2 s3031 tokens written, 27.5 s4428 tokens written, 22.4 s4230 tokens written, 24.2 s3929 tokens written, 27.6 s6541 tokens written, 32.3 s4197 tokens written, 28.4 s4343 tokens written, 22.9 s4996 tokens written, 25.1 s4100 tokens written, 22.4 s7069 tokens written, 32.8 s2902 tokens written, 16.8 s4759 tokens written, 25.1 s4105 tokens written, 25.1 s4258 tokens written, 21.8 s3363 tokens written, 24.9 s3320 tokens written, 17.4 s4801 tokens written, 23.1 s3147 tokens written, 18.2 s5191 tokens written, 26.4 s11192 tokens written, 48.3 s2745 tokens written, 15.3 s8464 tokens written, 35.9 s2820 tokens written, 16.2 s3611 tokens written, 20 s3803 tokens written, 22.5 s5501 tokens written, 26.1 s4389 tokens written, 24.1 s3469 tokens written, 18.9 s3716 tokens written, 22.5 s5988 tokens written, 31 sChess.comLichess
One dot per review, 162 reviews. The wait is almost entirely how much the model writes.

Total time correlates 0.97 with output tokens and 0.04 with prompt size. The median review took 21.6 seconds, and the first visible text arrived after 12.1. If I want a faster coach, the lever is a tighter answer, not a smaller game. It also means the engine and the coach are independent: engine time and coach time correlate at 0.10, so a slow phone does not make for a slow coach and vice versa.

00.250.50.751correlation (0 = unrelated, 1 = moves in lockstep)Engine time vs game lengthEngine time vs game length: 0.850.85Engine time vs player ratingEngine time vs player rating: 0.370.37Engine time vs key momentsEngine time vs key moments: 0.350.35Coach time vs answer lengthCoach time vs answer length: 0.970.97Coach time vs calls in flightCoach time vs calls in flight: 0.260.26Coach time vs engine timeCoach time vs engine time: 0.10.10Coach time vs game lengthCoach time vs game length: 0.050.05Coach time vs prompt sizeCoach time vs prompt size: 0.040.04
Engine time follows game length. Coach time follows the length of its own answer. Amber is the engine, blue is the coach.

What changes when many reviews run at once

I ran the coach one review at a time, then four workers, then eight. Four workers finished 97 reviews in roughly 13 minutes, about four times faster overall, with the median per review barely moving (20.4 to 21.8 seconds). The cost showed up in the tail: the 90th percentile went from 25.3 to 36.5 seconds, and first visible text from 16.6 to 26.9.

01020304050seconds, whole review1 in flight: median 21.6 s, p90 30 s, 42 reviews1n=422 in flight: median 21.3 s, p90 22.9 s, 7 reviews2n=73 in flight: median 23.9 s, p90 32.3 s, 22 reviews3n=224 in flight: median 20 s, p90 41 s, 15 reviews4n=155 in flight: median 19.2 s, p90 22.1 s, 15 reviews5n=156 in flight: median 21.3 s, p90 30.3 s, 19 reviews6n=197 in flight: median 24.3 s, p90 36.6 s, 34 reviews7n=348+ in flight: median 36 s, p90 43.3 s, 15 reviews8+n=15other calls in flight at any point during the reviewbar = median · tick = 90th percentile
Grouped by how many other calls overlapped each review at any point in its life. Sample sizes are small in the middle buckets, so read the shape, not each bar.

Up to seven overlapping calls the median stays between 19 and 24 seconds. At eight or more it jumps to 36. That is the number I would plan around.

The eight-worker run stopped for a reason I did not expect, and it was not speed. It stopped on credit. OpenRouter reserves the full cost of each in-flight request up front. With eight at once, those reservations added up to more than the account balance, and requests started returning HTTP 402. Three of those in a row and the fallback logic marks the model as degraded and fails the calls after it. In production the real ceiling on concurrency is not how fast the model answers, it is how much headroom the balance has.

Does the coach get the board right?

I ran an automatic check over every “piece on square” claim in the 162 reviews, replaying each one against every position in the game. 1,081 claims, 6 flagged. Four of the six were false alarms: general opening theory, a suggested move that was never played, and a correct promotion line inside a variation that I verified by hand. Two were real, both wrong-square slips: a fork said to hit “the king on h1” when it stood on g2 (the fork itself was real), and “keeping your pawn on e6” when it was never there. That is about 0.2 percent. It is the same kind of mistake I wrote about in the post on invented moves, now measured at scale instead of found by accident.

I also ran seven reviews through Apollo, the second coach. They wrote more (a median of about 7,200 tokens against 4,200) and took longer (a median of 31.8 seconds), and the 20 board claims in them were all correct. Seven reviews is a sample that proves nothing, and the run stopped there, so treat that as a hint, not a result.

What this does not tell you

  • One phone, one PC. A Pixel 9 is a recent phone. An older Android will be slower on the engine pass, and I have not measured one.
  • A local backend. The coach calls went through a local copy running on an embedded database that handles one query at a time, not production Postgres. The load numbers are real for that setup and probably pessimistic for production.
  • Almost all blitz. No rapid or classical games came out of the newest-games rule, so I have nothing on how long, slower games behave.
  • Strong players, mostly. The player pool skews toward titled players. It is useful for stress-testing, less so for knowing what a 1200 player’s game looks like.