Engineering9 min read
Maia 3, your rating, and the day its favourite opening move was h2h3
Stockfish tells you the best move. That is useful and, for most of us, a little cruel. The best move at move 23 is often a quiet queen retreat that no club player will ever find, and a coach that keeps pointing at it is not teaching, it is flexing.
Maia answers a different question: what does a human at this rating actually play here? It was trained on millions of real games, so its output is a probability for every legal move, shaped by how people of a given strength really choose. Put the two side by side and you can say something honest: “the engine likes Nd5, 74% of players around your rating find it, you played Bxf6.” That sentence is the reason I wanted Maia in the product at all.
Why Maia 3 beat picking a band
The older way to use Maia is to choose a rating band first. You load the model for 1500, or the one for 1700, and ask it for moves. That is workable for a research paper and awkward for a phone app. You ship several models or you ship one and swap, you round every player to the nearest band, and the ends of the range are thin: the first generation stopped at 1900.
Maia 3 takes the rating as an input. One model file, and for every position we hand it two numbers, the rating of the player to move and the rating of their opponent. It accepts anything from 600 to 2600. That changed three things for us:
- One download. The model is fetched once during setup and never swapped. No band logic, no per-band files to ship or get out of sync.
- Everyone fits. A 750 and a 2150 both get a model that was conditioned for them, not rounded to something nearby.
- Two cohorts for free. Because rating is just an input, we can ask the same model twice for each position: once as you, once as a player 300 points stronger. “Peers play this” and “stronger players play this” come from the same weights, which makes the gap between them mean something.
Maia thinks in blitz, so we convert
Maia 3 was trained on Lichess blitz games. A 1500 in blitz and a 1500 in rapid are not the same player, and the rapid pools skew weaker for the same number. If you feed a rapid rating straight into the model, it describes someone a bit worse than you.
So before inference we nudge the rating for the time control: no change for bullet and blitz, plus 150 for rapid and classical, plus 250 for daily. The offset touches only the number handed to Maia. The rating we use for everything else, blunder thresholds, which moments to pick, anything you see, is your real one.
The ends of the range are enforced too. Anything outside 600 to 2600 is clamped, because outside it the model is guessing. That mattered more than I expected. When I first measured Maia on the opponents of a grandmaster, every player came back rated 2829 to 3242, all clamped to the ceiling, and my experiment compared strong players to a 2600 baseline instead of to their peers. Same code, different sampling, and the answer flipped. Any measurement now checks its sample sits inside the model’s range first.
The day Maia’s favourite opening move was h2h3
Here is the part I would rather not write and think you should read.
For a while, Maia’s “human move” suggestions were arbitrary. Not wrong in a subtle way. Arbitrary. From the starting position, for a 1500 player, the model’s top pick was h2h3 at 15%. Nobody opens like that. The output was still a perfectly valid probability distribution over legal moves, so everything looked fine.
It was two bugs stacked on each other:
- The output was decoded with the wrong move vocabulary. The model’s move head has 4352 entries, indexed from-square times 64 plus to-square, with extra slots for promotions. I had used the 1858-entry move table from a different engine family. Every probability was attached to the wrong move.
- The input board was wrong too. Rows were read upside down, so White’s back rank landed on the wrong squares. And for Black, I mirrored the board by rotating it 180 degrees, which also swaps files. The model saw Black’s short castle as a long castle.
Either bug alone would have made the other fix look like it did nothing, which is how they survived. I only believed it when I measured against real play. On 600 positions from real Chess.com blitz games, rated 800 to 2200, with Maia conditioned on each player’s own rating:
| Measure | Before | After |
|---|---|---|
| Top pick matches the move played | 6.7% | 51.3% |
| Played move in the top three | 16.8% | 76.7% |
| Start position at 1500 | h2h3 15% | e4 64%, d4 24% |
The rating input started doing real work too. The Sicilian reply c5 climbs from 3% for an 800 to 33% for a 2600, and a beginner’s favourite wild knight move fades off the board. Treat anything Maia told you before that fix as noise. That includes percentages we showed.
Why the tests did not catch it
The old test checked two things: the top move was legal, and the probabilities summed to one. Both are true of a scrambled distribution. They are properties of the softmax, not of whether the model is speaking chess.
A test that cannot fail on the bug is a decoration.
The suite now asserts plausibility: from the start position e4 or d4 must lead, e4 alone must hold at least 40%, h2h3 must not be in the top five, and the distribution must be peaked rather than flat. It also checks the input tensor, with the king on the right square and pawns on the right rank, and that Black’s kingside castle is not the same move as White’s queenside one. Reverting the fix fails 18 of the 29 checks. I verified that, because a test you have never watched fail is a guess.