Suppose your team planned to launch on October 15. Two weeks later, you moved the date to November 1 because a supplier needed more time.
Both meetings are transcribed. A model with a large context window may be able to read them together and explain the change. But next month, in a new conversation, how will it find those records? Which date will it treat as current? Where can you check the reason?
The context window provides room to work with information. An AI product also needs ways to store that information, retrieve it later and handle changes. Increasing the first does not automatically solve the others.
What the window holds
The context window is the token budget available for one model request. It covers input and output and, for some models, reasoning tokens. OpenAI's conversation-state documentation explains those limits and how applications maintain information across requests.
A larger window can be valuable. It may let you compare several full transcripts, keep a long brief available while drafting, or analyze a document without splitting it into small pieces. When the relevant material fits, the model has a chance to work with it together.
Persistent memory concerns what remains available beyond that request. The surrounding application may store messages, save selected details or retrieve documents from an archive. It then supplies relevant material to the model when needed. A large window gives that application more room; it does not choose or maintain the archive for it.
Reading more does not settle which fact is current
In the launch example, searching for "launch date" might return both October 15 and November 1. Each transcript accurately records what people said at the time.
Answering "When are we launching?" requires the later decision. Answering "Why did the launch move?" requires the sequence and the supplier's explanation. A summary that collapses both discussions into an undated date field may lose the information needed for the second question.
Keeping dates and source references helps. So does preserving the difference between "we might need to delay" and "we agreed to November 1." These are decisions about how a product stores and retrieves information, even when the model is capable of reading every transcript.
There is also a separate question of how well the model uses the input. In the 2023 paper Lost in the Middle, researchers varied where relevant information appeared in long inputs. On the models and tasks they tested, performance was often better when that information appeared near the beginning or end than in the middle.
That finding is not a measurement of every current model. It does explain why a context-capacity number alone cannot establish whether a system will reliably recover your particular detail.
The roles of history and retrieval
A saved chat gives an application a record it may use again. It does not guarantee that every old message will appear in every future request. Products choose how to select, summarize or retrieve earlier material.
ChatGPT, for example, distinguishes saved memories from chat history in its Memory documentation. It also states that Memory does not retain every detail from every conversation. The behavior depends on the available features and settings.
Retrieval-augmented generation, usually called RAG, is one way to bring external information into an answer. A system searches a source, adds relevant material to the model's input and asks the model to respond using it. The original RAG paper studied combining a language model with an external information index.
For our launch question, retrieval could find the dated discussions before the model writes an answer. That still leaves practical responsibilities: searching the right account, finding the revision, returning a useful source reference and respecting a later deletion request.
Research such as MemGPT explores managing information across different memory tiers while working within a model's context limit. This is another way to approach the problem: organize what the system can bring back, alongside deciding how much it can process at once.
What to ask of a conversation memory app
Try returning to a decision that changed. Can you find the original discussion and the revision? Can you see the date and the relevant words? Can you correct a mistaken attribution or find out how to remove a saved record?
These questions tell you more about the product's usefulness than the size of the model's window alone. They also expose different failures. A missing transcript is a capture or storage problem. Retrieving the older decision without the revision is a retrieval problem. Misreading both available discussions is an answer-quality problem.
Draki's app captures conversations on your phone without requiring a bracelet, then uses OpenAI cloud transcription after your permission. The optional bracelet makes it convenient to keep capture within reach during the day, with recording you can turn on or off whenever you want.
You can set up Draki's optional custom MCP connection in an AI assistant that supports remote MCP, subject to its account requirements and settings. If you authorize read-only access, the assistant can search shared text and context and retrieve transcript passages. The assistant works with that shared material in its own service; local Draki processing does not move the assistant's model onto your phone. Our connection guide shows how to ask for those sources. Data controls and removal explains the separate question of managing what remains stored.
A useful answer to "When are we launching?" should return November 1 and show you where that decision was made. An answer to "Why?" should recover the change. The window makes room for those records; the product has to make them available when you ask.
FAQ
- Is a larger context window the same as memory?
- No. The context window limits how much a model can process in a request. Persistent memory depends on information being stored and made available again across requests or conversations.
- What is a context window?
- It is the token budget available to a model for one request, including input and output and, for some models, reasoning tokens. The exact limits depend on the model.
- Can a model miss information that fits inside its context window?
- Yes. Capacity alone does not establish how reliably a model will use every detail. The 2023 Lost in the Middle study found position-dependent performance in the models and retrieval tasks it tested.
- How does RAG relate to AI memory?
- Retrieval-augmented generation brings information from an external source into a model request. It can support persistent memory, but the surrounding system still has to manage storage, permissions, updates and removal.
- What should I check in an AI memory product?
- Check whether it can find a past detail, identify its source and date, distinguish an old decision from a revision, and explain how you can correct or remove saved information.
Written by Darijan Ducic
Draki for iPhone



