"Moderation" sounds like a single thing. Inside a companion app it's at least four, running at different speeds and involving different people.
Four layers, top to bottom

Every step down sees less, but looks closer.
1. Live filters. Automated classifiers can scan each message you send and each reply the model drafts before anything shows up. They target prohibited categories: anything involving minors, certain violence, signs of self-harm. When they fire, the message is blocked, the reply is toned down, or a warning appears. That's the origin of refusals and crisis cards. Our article on content filters explains why they sometimes fire in the middle of a scene.
2. Flagging across a conversation. Separate systems can score whole chats or accounts for patterns, such as repeated attempts to dodge a filter, harassment, or hints that a minor is using an adult app. A flag isn't a penalty. It's a note that a person might take a look.
3. People. Trust and safety staff, frequently working for outside contractors, go through a portion of flagged chats, user reports and appeals. Some companies also pull ordinary conversations at random to check quality or train models. Privacy policies describe this layer with phrases like "to improve our services" or "to enforce our terms."
4. Legal demands. Courts, police and regulators can compel data with legal orders. What a company does next depends on its own policy and the law where it operates.
What usually brings a human in
| Trigger | Reason |
|---|---|
| Anything involving minors | A legal duty in most countries, and often reported onward |
| Signs of self-harm or suicide | Crisis protocols are now legally required in several US states |
| Threats against real people | Risk of real-world harm |
| A report or support ticket from you | You asked someone to look |
| Repeated tries to get around filters | Enforcing the terms of service |
| Random quality sampling | Product improvement, where the policy permits |
What recent laws change
US chatbot laws for companions, beginning with New York in 2025 and California in 2026, make apps detect talk of suicide and self-harm and point users to crisis help. California further requires companies to report on their protocols. The practical effect is more automated monitoring of the most sensitive conversations, not less. AI companion laws in the US has the details.
How to check a policy
Search the privacy policy and terms for these words:
- "review," "moderate," "human": do people read chats, and why?
- "contractors" or "service providers": are the reviewers employed by someone else?
- "improve," "train": are ordinary chats sampled?
- "law enforcement," "legal process," "court order": when is data handed over, and does that need a court order?
A careful policy lists the triggers and says reviewers open conversations only when necessary. A loose one that gives staff access to "all content for any business purpose" also tells you something. Our four privacy checks walk through the rest of the document.
Takeaways
- Type as though a reviewer might see it, because occasionally one will.
- Leave other people's identifying details out, especially in explicit scenes. A reviewer reading a flagged chat sees them too.
- If a filter misfires during fiction, a quick out-of-character note usually fixes it. Arguing in character can trigger more flags.
- If who can read your chats matters most to you, the only arrangement where nobody else can is a companion running on your own computer. See running an AI companion locally.
Moderation isn't the same as surveillance. For nearly every conversation, nothing and no one looks past the automated filter. But the policy, not the app's friendly tone, sets the exceptions, and you want to know them before you need them.