How to Evaluate an AI Companion With Memory Before You Commit
October 5, 2026 6 min read
Memory in an AI companion is not one thing. It ranges from holding a single conversation together, to storing facts you can read and edit, to selectively recalling details days later. The gap between what a product page calls memory and what it actually does is where most disappointment happens. The practical way to close that gap is to run a small set of observable tests on any candidate before you share anything you would not want retained, and to read the privacy policy before the first meaningful conversation, not after. The framework below is built from the documentation currently published by the two providers whose claims were checked for this article — Character.AI and SentX, the company behind this site — and it applies to any other platform in the same way. Because SentX publishes this article and Victoria is its foundation model, the comparison section states what each side documents rather than which one we prefer.
What does the product actually store?
Start with the vocabulary. Character.AI's official memory announcement describes four controls: Story Memory, pinned messages, Facts, and a Memory Usage view, plus tools for carrying selected information into a new chat. These are provider-described controls in rollout, not independent proof that recall is complete or that every account sees the same set. SentX's homepage describes its own memory differently: selective, changing over time, with neither complete recall nor confidentiality following from a remembered conversation. The distinction matters because a product that lets you open and edit what it remembers is doing something structurally different from one that simply carries context forward. When you evaluate any option, ask for the equivalent of an inspection screen: can you see what is stored, correct it, and delete it? If the answer is unclear, record it as unknown rather than assuming the best or the worst.
How do you test recall without spending weeks?
A recall test needs nothing but patience, and it runs across several later sessions rather than one sitting.
- In a first session, tell the companion a small, distinctive, non-sensitive detail — a made-up favorite book, a fictional project name, a preference you have never stated anywhere else.
- End the session. Return the next day, or a few days later, and start a fresh conversation.
- Ask a related question without reminding it of the detail. Note whether it surfaces unprompted, appears only when prompted, or is gone.
- Give it a second detail, then deliberately correct it in a later session and see whether the correction sticks or the original version resurfaces.
- Delete what you can through the product's controls and note whether the item still appears on the next visit.
Each step records an observation about a specific reply, and that is all it is. A detail that fails to surface may simply have been left out of what the model assembled for that session — its absence proves nothing about what is stored, or whether it could resurface later. A detail that does come back may have been picked up from the current exchange rather than recalled from storage at all. None of this is a benchmark — the sample is one person, one device, a few days — but tracked across sessions, the observations separate advertised memory from working memory in a way a feature list cannot.
Does remembered mean understood?
There is a real difference between a person remembering what you told them and a system retrieving a stored entry. A friend carries your words into how they treat you, weighs them against everything else they know, and you can hold them to it. A companion app, as its documentation describes it, carries them forward under whatever controls and data policy its provider publishes — behavior defined by software, not by judgment. Recall that works is evidence about a mechanism, not about understanding; successful recall alone does not make the exchange a reciprocal human relationship. Judge the memory the way you would judge any tool — reliable, inspectable, deletable — and keep the rest of the interaction honest about what it is.
What should the privacy policy say before you trust it?
Read three things. First, whether submissions are treated as confidential at all. Second, whether retained interactions can enter shared systems or train models that serve other users. Third, what account deletion actually removes. On the second and third points, the published policies differ in ways that matter. SentX's privacy policy states plainly that submissions are not confidential, that retained interaction representations can enter shared memory and influence other users' interactions, and that deleting an account does not promise removal of material already incorporated into shared memory or model weights. That is a hard stop for anyone who intends to share work secrets, health information, relationship details, or anything else they would not want retained and potentially reused. Character.AI's memory announcement describes the controls but does not, in that document, settle the same three questions; the policy page is where those answers live, and it deserves the same close reading. A policy that says "we may use your data to improve our services" without scoping what improve includes is not an answer.
Who is writing this comparison, and why should you care?
SentX AI publishes this site, develops Victoria, and has a commercial interest in how this comparison lands. No head-to-head test was run, no user outcomes were collected, and neither product's accuracy was measured. Every feature statement above comes from the provider's own current documentation, checked on the same day, and every limitation is stated the way the provider states it. Treat that symmetry as the whole of the fairness claim: same criteria, same date, same standard of evidence — nothing more. If you want a verdict from an independent lab or a long-term user study, none of that exists in the sources checked for this article, and this article will not pretend otherwise.
Which option fits which reader?
Based only on the documented differences: if your use is confidential in any sense — professional, medical, financial, deeply personal — the published SentX policy rules it out before testing begins, and the same caution applies to any platform whose policy you have not fully read. If your use is low-stakes continuity — keeping track of an ongoing project, a creative thread, a habit you are building — both providers document mechanisms for that, and the recall test above is the honest way to find out which one holds up for you. Cost, signup requirements, and regional availability are secondary filters; check them at the point of purchase rather than letting them shape the memory assessment.
When should you just walk away?
Three conditions end the evaluation early. The product offers no way to view or delete what it has stored. The privacy policy routes your submissions into shared training without a meaningful opt-out, and your use case includes anything sensitive. Or the recall test comes up empty — in the test described above, none of the tested details appeared in the sampled later replies — while the product still markets itself primarily on memory. In the last case, the gap between the marketing and the observed behavior is the finding.
Sources
- SentX and Victoria — SentX
- Smarter Memory for Smarter Chats — Character.AI
- SentX Privacy Policy — SentX