LLMs8 min read
We tried other models. Gemini is still the one, for now.
Every few weeks someone asks which AI model EloCoach uses and whether it is the best one. The answer is Gemini, and the better question is how we know. This is what I tried, what broke, and why I have stopped switching for now.
Spring: tuning one model by hand
The first month was the usual cycle of poking a model until it behaved. I lowered the temperature because coaching should not improvise. I relaxed safety settings that were stricter than a chess conversation needs. And one afternoon the model returned an empty response for no reason I could see, which taught me to log the full compiled prompt and the raw response before touching anything. Most “the model is bad” problems are “I cannot see what the model received” problems.
My favourite early mistake cost a rate limit. A first-time review of fifty games hit a 429 from the provider after nearly a minute. I had already generated a short summary for every game, paid for it, stored it, and then sent the raw PGN of all fifty games anyway, clock annotations and all. Somewhere between one hundred and two hundred thousand tokens for nothing. The same bug report showed three different model names hard-coded in three places. Both were my fault, and both led to the same rule: one source of truth for which model runs what.
Summer: model choice becomes data
In August the coach moved behind OpenRouter. The point was not to use a hundred models. It was that which model serves which job is now a row in a table keyed by coach, mode and tier. Changing it is a database update. At boot the server checks that every model in the table still exists, and if the main one is gone it refuses to start rather than quietly serving worse answers. If a model fails mid-request it falls through to the next in line.
That setup made the September experiments cheap to run, which was the idea.
September: the tests
I tried other models for the main coaching call, including a GPT-5 model and a Kimi model, and split a smaller model out for the short jobs like the home greeting and the banter you see when a review opens. I will not publish a leaderboard, because it would not be one. This is not a benchmark, it is how those models behaved on our prompts. Three things decided it.
- Following a long, fussy instruction set. Our prompt tells the coach how to read engine data, when to cite your memory, how to format board chips, what voice to use, and what it must never claim. Gemini follows that list more reliably than anything else I tested.
- Speed. A review has to feel like a conversation. Gemini starts talking sooner and finishes sooner.
- Price. It is also the cheapest of the options, and in this product cost lands on the person paying for the next diagnosis.
There were non-chess surprises too. One model I used for the short jobs was a reasoning model. Its answer budget was around four hundred tokens, and it spent nearly all of it thinking: 384 of 396. The response came back looking fine, with a polite “finished because of length”, and the half-written JSON failed to parse, so the greeting silently never appeared. No error anywhere in our logs. I only saw it in the provider’s own dashboard, where the reasoning-token count was sitting right next to the completion count. We turned reasoning off for that job and doubled the budget.
If a call returns “length” as its finish reason, nothing else matters yet.
A cheaper lesson came in the same week: one of the model names I put in the table was never a real model name, and another had been withdrawn from the provider. That is exactly the failure the boot check and the fallback chain were written for.
The prompt is bigger than I want
I should be straight about this. The prompt is long. It carries the engine summary, your memory, the rules for board chips, the voice and the safety rails, and it keeps growing as I find new ways for the model to misbehave. That costs money and adds seconds. Shrinking it is on the list, and I expect to get real savings from it without the coach getting worse. A shorter prompt with the same behaviour is the cheapest improvement left in the project.
Where this goes next
Gemini is the default because it earned it, not because I am attached to it. The plumbing already allows a different model per coach. So I can imagine a grandmaster-style coach running on Claude or ChatGPT, if that coach’s owner decides the output sounds more like them. The deciding question would not be which lab wins a benchmark. It would be whether that coach’s students recognise their teacher in the answer.