Engineering7 min read
By EloCoach.ai
I swapped in Stockfish’s own win model. Two days later I took it back out.
Every grade in EloCoach rests on one number: how likely you are to win from a position. The engine gives you a score in centipawns, and something has to turn that into a percentage before you can say “that move cost you 12 points” or “this is a blunder”. On 21 September I replaced the curve that does this with Stockfish’s own. On 23 September I put the old one back. This is what happened in between.
The curve I started with
The first curve is a simple sigmoid with a constant of 0.00368208. It is the one Lichess uses, and Lichess did not pick it by hand: it came from a regression over tens of thousands of their own games. At +100 centipawns it says you win 59 percent of the time. At +200, about 68 percent. That feels like real chess between real people: a pawn up is nice, not over.
Why Stockfish’s own model looked better
I was upgrading the engine to Stockfish 19 anyway, and Stockfish ships its own win model. It is the engine’s own, it knows how much material is on the board, and someone else keeps it up to date instead of me carrying a constant forever. That sounded like exactly the right thing, so I ported it exactly. There was one detail to get right: the raw win rate sits near 0.7 percent at equality because draws dominate at long time controls, so I used the expected score (wins plus half the draws) to keep a level position at 50 percent.
I also noted in the commit that grading would get stricter, and left the grading thresholds alone on purpose, because rescaling them to reproduce the old labels would make adopting the engine’s model pointless. That was the plan. It lasted two days.
What it did to real reviews
One review after the change graded three of six key moments as blunders. A single positional mistake for a player rated around 850 came out as 100 to 0. Here is why.
| Position | Lichess curve | Stockfish model |
|---|---|---|
| +100 centipawns | 59% | 75% |
| +200 centipawns | 68% | 99.7% |
Stockfish’s model is fitted to long games between engines, where one pawn really is close to decisive. People are not engines. Measured in centipawns the two curves differ in scale by about a factor of thirteen, roughly 20 against 272. A model that is right about engines is wrong about someone at 850 who will miss a tactic on the next move.
Three separate causes
Putting the old curve back fixed the headline problem, and the exercise turned up two others that I had been blaming on the curve.
- The curve itself. One function body to restore.
- Position descriptions that depended on the curve. Things like “this position is decided” and “this move is natural for a human” were defined in win probability. When the curve changed, they changed silently: “decided” went from about +471 centipawns to +117, and the slack allowed around a one-pawn edge shrank from about 54 centipawns to about 8. That second one is why Maia’s suggestions stopped looking human. These now live in centipawns, so a change to the curve cannot move them. This is the part I would call the lasting improvement.
- Grading bands that were harsher than Lichess’s. Lichess marks an inaccuracy at 10 points, a mistake at 20 and a blunder at 30. We were one band harsher at every level, so a two-pawn slip read as a mistake where Lichess says inaccuracy. I matched them, so a move gets the label you already know from analysing the same game there.
Did it work?
On device that same day, one game’s grades went from blunder, blunder, great, great, blunder, excellent to great, great, great, great, inaccuracy. That reads fairly through the game. It also looks a little generous, and I noted that “great” is still defined by a win-probability gap I deliberately did not touch. I will find out whether that matters in real use.
That evidence was a search away throughout; reasoning from first principles lost twice to it.
That is how I put it in the notes afterwards, and it is the honest summary. The answer was already published. The better question was never “which model is more advanced” but “whose games was this fitted on”.