Nearly every gripe about these apps, including the repetition, the drift, and the odd gap between two runs of the same scene, traces back to a single mechanism. Ten minutes on it is well spent, because it swaps "the app is broken" for "the app is doing what it does, and here's the knob."
One piece at a time
The model doesn't draft a reply and then hand it over. It generates one token, roughly a word or a word fragment, then reads everything so far, including that token, and generates the next.
That's the entire loop. No plan, no outline, no draft to polish. A reply that ends elegantly had no idea it would when it began.
Two effects follow right away, and you've probably seen both.
No taking it back. Once the model writes "I've never been to Paris," everything afterward builds on that line. It will keep things consistent with a sentence it produced by accident before it will contradict itself.
Mistakes snowball. A slightly off word in sentence one pushes sentence two off, which pushes sentence three off. That's why long replies wander and short ones seldom do.
Weighted dice before every word
At each step the model holds a ranked list of candidates with odds attached: say "good" at 30%, "fine" at 12%, "terrible" at 3%, then a long tail.
If it always grabbed the top pick, the character would be predictable and painfully dull: same greeting each time, same three jokes. So the app draws at random, with the odds as weights.
A setting usually called temperature controls how those weights are shaped. Turn it down and the odds pile onto the obvious picks, giving you steady, predictable, eventually boring text. Turn it up and the odds spread out, giving you surprise and creativity, plus a higher chance of a reply that makes no sense.
You rarely get a slider. You get whatever spot the app picked, and it's a big reason people say one app "feels smarter" than another. Often it isn't smarter. It's just running warmer or cooler.
That also answers "why did regenerating fix it?" Nothing got fixed. You drew again and landed a better roll.
What the model can see
Before your message goes in, the app builds a block of text you never view: the character description, facts stored about you, a summary of your history, and the latest conversation. All of that forms the context window, and the reply gets written from it as one continuous document.
Here's a tip most people never use: the model answers the shape of what it sees. Feed it three short, flat messages and you'll get short, flat replies, since it's extending that pattern. Send something vivid and the tone moves. You aren't persuading a person. You're setting a pattern, and patterns spread both ways.
That also accounts for the most common self-inflicted problem here. Users drift into "hey," "how was your day," "what are you doing," then decide the app got worse. It's extending the document the two of you are writing.
Same engine, different feel
Plenty of apps here run on similar underlying models, yet they feel distinct. The gap comes from what's wrapped around the model:
- The system prompt. How the character is described, how long that description runs, and what it says about tone and pacing.
- Sampling settings. The temperature choice covered above.
- Retrieval. Which stored facts and which slice of history get put in front of the model on a given turn.
- The filter. What it refuses, and whether a refusal is graceful or a brick wall.
Nomi feels steady across weeks because of the third item, not a bigger model. Candy AI bundles chat, images and voice into one subscription, and that's a product call, not a modeling one. No benchmark would catch either difference, but both surface within two weeks of use, which is how we score them.
Four ways to use this
Steer at the start. The first couple of sentences set up the rest. When a scene is heading somewhere bad, stop and restate instead of arguing with paragraph four.
Regenerate with intent. If a reply is 80% right, regenerating discards that 80%. If the opening sentence is wrong, regenerate right away.
Write in the tone you want back. Short and flat in, short and flat out. It's the cheapest fix for anyone who thinks their companion turned boring.
Don't fight over invented facts. A model that wrote something false will stand by it, because staying consistent with its own output is how it works. Correct it in a fresh message. Demanding a confession won't work. This has its own article: why AI companions make things up.
What to remember
Nobody's home between your messages. The character exists only for one generation and gets rebuilt from text for the next. Every sense of continuity comes from the plumbing around the model: stored facts, summaries, retrieval. That plumbing is what really separates a $10 app from a $20 one.
It's also why our ranking gives memory and consistency so much weight. A landing page can't show you either.

