Why we built it6 min read
I pasted my games into Gemini for a month. Then I built what I actually wanted.
For a few weeks this spring my chess study routine looked like this. Finish a game on Chess.com. Open the archive. Copy the PGN. Paste it into Gemini with a short prompt asking what went wrong. Read the answer. Repeat tomorrow.
It worked better than it had any right to. The explanations were readable, they had a point of view, and every so often one of them made me say “oh, that’s why” out loud, which no engine evaluation bar has ever done for me. It was also fun, which I did not expect from post-game analysis. Losing a game and then having someone talk about it is a different experience from losing a game and staring at a red number.
It was also not good. Not perfect by any means, and not convenient for sure.
What was broken about the copy-paste loop
- It was a chore. Every review was six small steps before I got to the useful part. Anything that takes six steps is something you stop doing in week three.
- It had no eyes. The model read a wall of notation and reasoned about it in prose. It could be wrong, and the only way I could tell was to set the position up on a board myself and check. That defeats the point.
- It had no memory. Every chat started from zero. Which brings me to the thing that finally made me build something.
“Check for blunders before you move”
You have heard this advice. I have heard it from coaches, from books, from videos, and from every chat window I pasted a game into. It is true and it is useless, because it is the advice you give a stranger. Nobody who has watched me play for a month would say it to me. They would say: you do fine until the clock drops under a minute, and then you stop looking at what your opponent’s last move attacked.
That sentence needs two things a one-off chat cannot have. It needs to have seen many of my games, and it needs to remember them. Getting the same generic advice every time is not really a flaw in the model. It is a flaw in the setup. The model never got to know me.
What I wanted instead
Let me review my games, get feedback, have fun, and never hear the same generic thing twice.
Concretely, that turned into a list I have been refining ever since:
- The chess facts should come from engines, not from a language model’s memory. Stockfish for what is objectively best. Maia for what a human at your rating actually plays, because “just play the engine move” is not advice anyone can follow.
- The language model should explain those facts, in a voice you like, and it should not be allowed to make up a board.
- There should be a memory. Not a transcript of old chats, but a small, honest picture of how you play that gets updated when the evidence changes, including when you stop making a mistake.
- It should pull your games itself. No PGN, no clipboard. Type a handle, tap a game.
None of that is a clever idea. It is just the list of things I kept doing by hand, written down. The interesting part is that each item turned out to hide a real engineering problem, and the next posts are about those: the model that kept inventing moves, the day Maia’s favourite opening move was h2h3, what it costs to run any of this, and what happens when a language model is the only thing standing between a user and a coach.
If you have a copy-paste routine of your own, I would like to hear what it looks like. That is the app I am trying to replace.