← All posts

Engineering8 min read

My coach kept inventing chess moves

The first version of the EloCoach review worked like this. You tap a game. The backend sends the PGN to the language model. The model starts writing, and while it writes it drops little board diagrams into the answer: “here is the line you missed”, followed by a position and a few moves.

The prose was good. The boards were a coin flip. Sometimes the line was fine. Sometimes it moved a bishop through a pawn, or captured something that had left the square three moves earlier, or started from the wrong position entirely. The model was not lying. It was doing what language models do, which is produce the most plausible next token, and a plausible chess move is not the same thing as a legal one.

Order of operations was the real bug

In that first version the coach started writing the moment you tapped, before Stockfish had looked at a single position. It improvised variations from memory, cited moments it had no data on, and embedded chips that were statistically likely to contain illegal moves. We were asking a text predictor to do the job of an engine and then acting surprised.

The fix that mattered was not a validator. It was a rule: the coach writes last. Stockfish and Maia run first, on your phone. Their output is packaged into one object with the annotated game, a win probability and centipawn score on every ply, how often players at your rating pick each move, and a handful of selected moments. The model receives that and explains it. It can say “74% of players around your rating find Nd5 here” because the 74 came from a model that was run, not recalled.

That cost us roughly 8 to 15 seconds of loading before the first word appears. I would make the trade again every time.

Then trust nothing anyway

Even with the numbers in front of it, the model still writes the chips. So every chip that carries moves gets replayed against a real rules engine, chess.js, before it reaches a screen. The check is almost embarrassingly simple:

const chess = new Chess(startFen)
const tokens = chip.moves.replace(/\d+\.{1,3}/g, '').trim().split(/\s+/)
for (const san of tokens) {
  try { if (!chess.move(san)) return false } catch { return false }
}
return true

If the whole line is legal, it ships. If it is not, we do not throw the chip away. We keep the longest legal prefix and drop the rest, so a line that went wrong on its fifth move still shows the four moves that were right. If not even the first move is legal, the chip is removed. Wrong boards are worse than no board.

The bug that got through

Validation catches illegal lines. It does not catch a legal line shown from the wrong starting point, and that is how one slipped through. A “sequence” chip with no starting position renders from the initial board. The coach had written a mid-game continuation, something like 16...Bxd1 17.Rxh3 Bxa4, with no position attached. From move one of a fresh game that is nonsense, and the app showed an empty board with “No move data available”. Live and when you reopened the session, because both paths parse the same stored chip.

The obvious fix is to make the model supply a starting position as well. We did not do that. A hand-typed FEN from a language model is just a second hallucination that is harder to check than the first. Instead the prompt now tells the coach to use a position chip anchored at the real game moment, where the data already exists in the game’s own PGN, and the backend rejects any sequence chip whose moves do not parse from the position it claims. There is a test for it, added in the same change.

Do not fix one unverifiable input by asking for another.

Smaller lessons from the same stretch

  • Our own server was also eating the output. A Fastify response schema was silently stripping the metadata that tells the app which game a chip refers to. The model was fine. The pipe was not. Always check the response on the wire, not just the code that built it.
  • Chips refer to games by a six character alias, never the database ID, and the alias must already exist in the context the model was given. If it was never shown, it cannot be emitted.
  • The prompt carries a short internal checklist. Before claiming a piece is loose, the model must name the enemy piece that can capture it and confirm the capture is legal. It is not perfect. It is cheap and it helps.