Inside an AI Companion's Reply: Tokens, Dice Rolls and Drift

Under the hood

It builds each reply a piece at a time, rolling weighted dice before every word. That one idea explains why identical questions get different answers, why long replies wander, and why hitting "regenerate" so often does the trick.

We may earn a commission from links on this page. It never changes a rating.

Nearly every gripe about these apps, including the repetition, the drift, and the odd gap between two runs of the same scene, traces back to a single mechanism. Ten minutes on it is well spent, because it swaps "the app is broken" for "the app is doing what it does, and here's the knob."

One piece at a time

The model doesn't draft a reply and then hand it over. It generates one token, roughly a word or a word fragment, then reads everything so far, including that token, and generates the next.

That's the entire loop. No plan, no outline, no draft to polish. A reply that ends elegantly had no idea it would when it began.

Two effects follow right away, and you've probably seen both.

No taking it back. Once the model writes "I've never been to Paris," everything afterward builds on that line. It will keep things consistent with a sentence it produced by accident before it will contradict itself.

Mistakes snowball. A slightly off word in sentence one pushes sentence two off, which pushes sentence three off. That's why long replies wander and short ones seldom do.

Weighted dice before every word

At each step the model holds a ranked list of candidates with odds attached: say "good" at 30%, "fine" at 12%, "terrible" at 3%, then a long tail.

If it always grabbed the top pick, the character would be predictable and painfully dull: same greeting each time, same three jokes. So the app draws at random, with the odds as weights.

A setting usually called temperature controls how those weights are shaped. Turn it down and the odds pile onto the obvious picks, giving you steady, predictable, eventually boring text. Turn it up and the odds spread out, giving you surprise and creativity, plus a higher chance of a reply that makes no sense.

You rarely get a slider. You get whatever spot the app picked, and it's a big reason people say one app "feels smarter" than another. Often it isn't smarter. It's just running warmer or cooler.

That also answers "why did regenerating fix it?" Nothing got fixed. You drew again and landed a better roll.

What the model can see

Before your message goes in, the app builds a block of text you never view: the character description, facts stored about you, a summary of your history, and the latest conversation. All of that forms the context window, and the reply gets written from it as one continuous document.

Here's a tip most people never use: the model answers the shape of what it sees. Feed it three short, flat messages and you'll get short, flat replies, since it's extending that pattern. Send something vivid and the tone moves. You aren't persuading a person. You're setting a pattern, and patterns spread both ways.

That also accounts for the most common self-inflicted problem here. Users drift into "hey," "how was your day," "what are you doing," then decide the app got worse. It's extending the document the two of you are writing.

Same engine, different feel

Plenty of apps here run on similar underlying models, yet they feel distinct. The gap comes from what's wrapped around the model:

  • The system prompt. How the character is described, how long that description runs, and what it says about tone and pacing.
  • Sampling settings. The temperature choice covered above.
  • Retrieval. Which stored facts and which slice of history get put in front of the model on a given turn.
  • The filter. What it refuses, and whether a refusal is graceful or a brick wall.

Nomi feels steady across weeks because of the third item, not a bigger model. Candy AI bundles chat, images and voice into one subscription, and that's a product call, not a modeling one. No benchmark would catch either difference, but both surface within two weeks of use, which is how we score them.

Four ways to use this

Steer at the start. The first couple of sentences set up the rest. When a scene is heading somewhere bad, stop and restate instead of arguing with paragraph four.

Regenerate with intent. If a reply is 80% right, regenerating discards that 80%. If the opening sentence is wrong, regenerate right away.

Write in the tone you want back. Short and flat in, short and flat out. It's the cheapest fix for anyone who thinks their companion turned boring.

Don't fight over invented facts. A model that wrote something false will stand by it, because staying consistent with its own output is how it works. Correct it in a fresh message. Demanding a confession won't work. This has its own article: why AI companions make things up.

What to remember

Nobody's home between your messages. The character exists only for one generation and gets rebuilt from text for the next. Every sense of continuity comes from the plumbing around the model: stored facts, summaries, retrieval. That plumbing is what really separates a $10 app from a $20 one.

It's also why our ranking gives memory and consistency so much weight. A landing page can't show you either.

Candy AI

4.6Rating: 4.6 out of 5

Chat, images and voice on a single plan, and all three hold up in daily use.

Price
from $12.99/month
Free tier
Yes
United States
Available

Nomi

4.4Rating: 4.4 out of 5

The strongest long-term memory we have measured, paired with the most natural conversation.

Price
from $15.99/month
Free tier
Yes
United States
Available

Frequently asked questions

Why do I get a different answer every time I ask the same thing?

The app picks each next word by chance, weighted by probability, instead of always taking the top choice. That randomness is adjustable, and it's what keeps a character from sounding canned.

What happens when I tap "regenerate"?

The app sends the same input again and lets chance play out differently. The character hasn't changed. You're just pulling another sample from the same spread of possibilities.

Why do long replies go off the rails?

Every word depends on the words before it, so a small slip early gets magnified. By the third paragraph the reply is mostly reacting to itself, not to you.