← All posts

Engineering7 min read

What counts as a “brilliant” chess move? My five attempts at an answer.

When the review screen puts “!!” next to a move, the coach praises it, the move gets picked as a key moment, and you feel good about your game. So that label had better mean something. The current definition has a handful of conditions and a history of me getting it wrong, each time in a way I only saw after it reached a real game.

Version one: is the piece in danger?

In mid June I split the good moves in two. A “great” move (!) is one where the engine’s top choice is much better than the second best, so there is basically one move to find. A “brilliant” move (!!) is a great move that is also a sacrifice. To decide whether it was a sacrifice I asked chess.js whether the piece was now on an attacked square.

That is the kind of rule that sounds right and is not. An attacked square does not mean you gave anything up. It cannot tell a sacrifice from an exchange, because in an exchange your piece is attacked too and you simply get the other piece back.

Version two: wait until the position is quiet

A week later I rewrote it. Instead of looking at the board right after the move, the detector follows the engine’s expected continuation until nothing is being captured or checked any more, then counts material. If you are down at least a pawn once things settle, you gave something up on purpose. It follows at most eight plies.

That rewrite also added the other conditions that still exist:

  • The move has to be the engine’s top choice, with a win-probability gap of at least 10 points over the second best.
  • The position must still be live: you cannot be more than 90 percent winning before the move, because a flashy move in a won position is not brilliant, it is just a flourish.
  • A comprehension guard: if the move gets punished within your next two moves, it does not count. A sacrifice that only works because the opponent misses the refutation is luck.

I also added a “rare versus natural” flavour using Maia, so the coach could say whether most players at your rating would have found it. I come back to what happened to that below.

First wrong label: the forked knight

On 9 July I found a move being called brilliant that was nothing of the sort. A knight was forked, the player retreated it, and that cost a pawn somewhere else. Material before and after said “sacrifice”. But the knight was already lost one way or another, and retreating it was the least bad option, not an elective sacrifice. The fix subtracts the value of anything of yours that was already attacked and undefended before the move, so only material you chose to give up counts.

Second wrong thing: the prompt promised a field that did not exist

On 22 July I did a tidy-up pass and found three problems at once. The rare-versus-natural field was being computed and sent, but the code that builds the prompt never read it, even though the prompt told the coach it would be there. I deleted the field. A “best” class was a real value at runtime that the prompt’s list of classes did not mention. And “best” had a flat score floor that could outrank a real blunder when picking which moments to show, which is the opposite of what you want in a review. All three were the same kind of bug: two parts of the system disagreeing about what a label means, and nothing checking. The fix was a pair of tests that compare what the prompt documents with what the renderer and runtime actually produce.

Third wrong thing: the coach’s wording

On 26 August the problem was not the label but what the coach said about it. “Brilliant” guarantees a big gap over the second-best move. It does not guarantee your position got better. A brave defensive try in an already lost position is still brilliant, but the template wording around it said things like “wins games” and “real reward”. I kept the judgement about the act (brave, deliberate) and dropped every claim about the result, and added a regression test because this exact contradiction had come back once already.

What I learned about the definition

Every condition in there was added because a wrong label got through.

That is a fine way to build a heuristic and a poor way to be sure about it. Each patch fixed the case I had just seen. None of them started from the question “what should be true of a move before we print two exclamation marks”, and that is the question I should have asked on day one.

The conditions are now written down with their reasons, and a test pins what the thresholds mean in centipawns, so a change to the win-probability curve cannot silently shift them. How that curve changed, and why, is the subject of a later post.