Pictures, Voice and Video: What Happens Behind the Media Features

Under the hood

Your companion's words and its pictures come from two unrelated systems that hardly communicate. That's the reason its face keeps shifting, and the reason every app puts a meter on media first.

We may earn a commission from links on this page. It never changes a rating.

It's easy to picture an AI companion as one mind that talks, draws and speaks. It isn't. It's three or four separate systems stitched together by the app, and most media complaints trace back to the stitching.

Three engines behind one screen

The language model does the chatting. Text out, nothing more.

The image model is another beast entirely, usually a diffusion model. Give it a written description and it draws a picture. It has never read a word of your chat.

The speech system converts text into sound. It's also its own thing, and it knows only the single sentence it's handed.

Here's what happens when you ask for a selfie. The language model drafts a brief description of the shot. The app tacks on the character's saved appearance tags. That combined prompt goes to the image model, and a picture comes back. Then the chat responds to a picture it can't actually see, going off the description it wrote itself.

That handoff sits behind almost every media gripe in this category.

Why her face changes

A diffusion model doesn't look up your character. It invents someone who fits a description. "Long dark hair, green eyes, mid-twenties" fits millions of faces, so a different one turns up each time.

Apps handle it with varying success:

  • Saved appearance tags. The same text string every time. Cheap, and only loosely consistent.
  • A reference image guiding each render. Far better, and the usual trick when a face stays recognizable.
  • A model tuned to one character. The most consistent option and by far the priciest, which is why it mostly lives on upper tiers.

This is worth testing on a free tier before you pay. Make four pictures of one character in four different settings. If four different women come back, no subscription will fix that. Consistency across images is a big reason Candy AI tops our ranking. It's also why Secret Desires scores well: it builds looks, voice and personality as one package and adds short video.

Why media always gets a meter

Text is cheap. A chat reply costs the operator a fraction of a cent.

An image costs noticeably more each time. Voice is charged per second of audio. Video costs far more than either.

That gap explains the pricing pattern you'll see everywhere: roomy message allowances, tight image caps, voice minutes sold on the side, video locked to top tiers. Nobody is inventing scarcity to push an upsell. The limits mirror the invoice the operator actually receives.

So when an app says "unlimited," it nearly always means unlimited text. Read it that way and the pricing pages finally make sense. The math is in what image generation really costs you.

Voice: what you gain and what you pay

Hearing a reply feels different from reading it, and the difference is bigger than the spec sheet suggests.

Check two things before you pay for it:

Lag. Wait three seconds before every answer and the spell is broken. Try it on a free tier or trial. A demo video proves nothing.

Note or call? A voice note, where the character reads its reply aloud, is common and cheap. A live call, where you talk and it talks back, is a separate product and usually a separate tier. Nomi lists voice calls among its features. Plenty of apps say "voice" and mean notes.

Where images fall short

Two limits to know about up front:

Hands, lettering and small details stay shaky in every generator I've tried in this category. Nobody has cracked it, and an app claiming otherwise is showing you its best shot, not its typical one.

The picture doesn't know your scene. The chat model writes a one-line prompt, so anything outside that line vanishes: the room you described, the outfit from three messages back, the time of day. Spelling out what you want in the request fixes more than any setting will.

Judging the media side in one evening

Burn through a free tier's whole allowance in one sitting instead of a picture a day. You're checking four things:

  1. Consistency. Four pictures, one character, different scenes.
  2. Prompt adherence. Ask for something specific and see how much survives.
  3. Lag. For images and voice both.
  4. What eats the cap. Does a failed or refused generation still use up a slot?

That last one is the least documented and the most irritating to find out after you've paid. Along with the rest of what the free tiers actually include, it only surfaces when you use the app properly.

Candy AI

4.6Rating: 4.6 out of 5

Chat, images and voice on a single plan, and all three hold up in daily use.

Price
from $12.99/month
Free tier
Yes
United States
Available

Secret Desires

4.3Rating: 4.3 out of 5

Unfiltered roleplay that keeps its characters consistent long after other apps lose the thread.

Free tier
Yes
United States
Available

Frequently asked questions

Why doesn't my companion look the same twice?

The image model builds the character fresh from a written description every time. Apps with a steady face lean on a saved reference image or a model tuned to that character. Without one, every request is a new roll of the dice.

What makes voice so tightly limited?

Speech is billed by the length of the audio, so each second costs the operator real money. Text costs a tiny fraction of that. So message caps are loose and voice minutes are stingy.

Can the image generator read our chat?

No. It sees only a short prompt the app writes on its behalf and never your history. That's why a picture can miss details that seemed obvious in the conversation.