Two different kinds of “memory”
A language model has knowledge built into it during training: grammar, facts about the world, how conversations usually go. It does not have your chats inside it, and a personal clone should not need to train a new model on them.
Personal memory comes from outside the model. The clone keeps a store of your messages and notes, looks things up when needed and passes what it finds to the model for that one reply. Researchers call this combination of built-in knowledge and an external, searchable store retrieval-augmented generation, a term introduced in a 2020 paper by Patrick Lewis and colleagues.
Style is handled separately. How you write (length, rhythm, emoji, nicknames) is learned as a profile and as examples. What you remember (people, places, events) is retrieved from the index.
What a context window is
A language model writes each reply by reading a block of text called its context: instructions, the recent conversation and any extra material. The maximum size of that block is the context window. Anything outside it does not exist for the model at that moment.
Context windows have grown a lot, but years of personal chats can still run to millions of words across many contacts. Even where a huge window is available, sending everything with every message would be slow and expensive. It would also send far more personal data than one reply needs.
Why a persona prompt cannot hold years of chats
A common shortcut is to write a summary of a person (“Anna is warm, teases her brother, loves hiking”) and give it to the model as a persona. A summary like this is useful for tone. It is not memory.
- It is lossy. A few paragraphs cannot contain the date of a trip, the name of a neighbour’s dog or what was said after an argument.
- It invites invention. When asked about a detail that is not in the summary, a model tends to fill the gap with something plausible.
- Long inputs are not read evenly. Liu and colleagues (2023) found that language models used information at the start and end of a long context better than information in the middle. Pasting everything in does not mean everything is used.
This is why Memory Clone’s Persona Blueprint describes style and character, while facts come from searching the complete index.
How retrieval works, step by step
- Index the messagesEach imported message is stored with its sender, date and conversation, so it can be found again later.
- Understand the new messageWhen you write to the clone, the system works out what you are asking about: a person, a place, a time or an event.
- Search the indexIt searches for passages that match, including nearby messages so the meaning is not lost.
- Filter by relationshipMemories from the current relationship come first, and private facts the current contact would not know are kept out.
- Build the contextThe best passages, style examples and the recent conversation are combined into one bounded context.
- Write the replyThe language model writes a reply in the person’s style, using the retrieved details.
Dated recall: remembering when, not just what
Chat exports carry a date and time on every message. A good memory system keeps those dates, so it can answer questions such as “what did we do last New Year?” or “when did you start that job?”. It can also tell an old plan from a recent one.
Dates also help the clone avoid mixing periods of life. Something true in 2016 may not be true now, and the date lets the system say so.
Relationship scoping: keeping memories in their place
Real people know things from one friend that they would never tell another. A clone that searched everything for everyone would leak those secrets.
Relationship scoping means that, when the clone talks to one contact, its memories come first from the conversations with that contact. Background from elsewhere can help it understand a message, but names, secrets and facts the current person would not know are kept out of the reply. In Memory Clone, the “Myself” mode is the exception: talking to yourself, the clone can use your full history.
Where AI memory goes wrong
| Failure | What it looks like | What helps |
|---|---|---|
| Missing data | The clone cannot recall something that really happened | Import the chat where it was discussed, or add a note |
| Wrong retrieval | It brings up a similar but different event | Ask more specifically, with a name, place or year |
| Hallucination | It states a detail that never happened | Check important facts in Memory Search; treat replies as simulation |
| Outdated facts | It talks about an old job or address as current | Dated recall, plus notes about what has changed |
| Wrong person | Two contacts with the same name are mixed up | Link and separate contacts during import |
No retrieval system is perfect. A clone can sound confident and still be wrong, so never rely on it for facts that matter.
How Memory Clone does it
Memory Clone builds the memory index locally on your Windows PC from your WhatsApp, Messenger, Instagram and Discord exports and your notes. The original files and the index stay on your computer. Our guide to local-first AI explains what leaves your PC and when.

- Memory Search lets you find exact, dated memories yourself, with the messages around them.
- Brain Map shows the people and memories the clone knows about.
- Relationship memory has priority, and private facts from one relationship are kept out of other conversations.
- Full-corpus recall: replies can draw on the complete indexed history, not only the Persona Blueprint.
- Excerpts only with consent: cloud processing is off until you switch it on for a clone. Then only the selected excerpts needed for a reply are sent through the Memory Clone server to the AI provider. Whole archives are never uploaded, and your content is not used for training.
Memory Clone is in alpha, so expect occasional misses. You can download the app and try Memory Search on your own exports. To understand which data helps most, read how many messages an AI needs and how to write notes that help a clone.
Häufig gestellte Fragen
Does an AI clone remember everything I imported?
It can search everything in its index, but each reply only uses the passages it finds most relevant. If something is never retrieved, it will not appear in the reply.
Is my data used to retrain the AI model?
Not in Memory Clone. The model is used as it is; your messages are searched locally and only selected excerpts are sent to write a reply, once you enable cloud processing.
What is RAG?
Retrieval-augmented generation: searching a document store for relevant passages and giving them to a language model before it answers.
Why did my clone make up a memory?
Language models fill gaps with plausible text when the right passage is missing or not found. Adding the relevant chat or a note, and asking more specifically, usually helps.
Can the clone tell one friend what another friend said?
It is designed not to. Relationship scoping keeps private facts from one relationship out of conversations with other people.
Quellen
Geprüft am . Produkte und Hilfeseiten von Drittanbietern ändern sich häufig; folge den Links für aktuelle Details.
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al., 2020) (arXiv / NeurIPS 2020)
- Lost in the Middle: How Language Models Use Long Contexts (Liu et al., 2023) (arXiv / Transactions of the ACL)
App menus change over time, so labels on your device may differ slightly. Memory Clone creates an AI simulation from the material you provide; it is not the person.