How do I add memory to an AI agent?
Short-Term vs. Long-Term Memory
Short-term memory is the conversation history you include in the prompt, limited by the model's context window (e.g., 8k to 128k tokens). You can manage it by keeping recent messages and summarizing older ones. Long-term memory involves storing information outside the context window, typically in a vector database, and retrieving it when relevant.
To implement long-term memory, embed text (e.g., past conversations, documents) into vectors and store them. When the agent needs information, it queries the vector DB for similar vectors and includes the results in the prompt. This is the core of retrieval-augmented generation (RAG).
- Short-term: maintain a list of recent messages
- Summarization: condense older history to save tokens
- Long-term: use a vector DB (Pinecone, Chroma, etc.)
- Embedding: convert text to vectors with an embedding model
- Retrieval: fetch relevant memories based on similarity
Best Practices
Decide what to store: not everything is worth remembering. Store facts, user preferences, and key decisions. Use metadata (timestamps, user IDs) to filter retrievals. Regularly prune or summarize old memories to keep the system efficient.
For production, consider a hybrid approach: use a fast in-memory store for recent context and a vector DB for long-term. Also, implement memory write policies to avoid storing redundant or sensitive information.
Common mistakes
- Assuming the context window is unlimited—it's not, and you must manage it.
- Storing everything without filtering, which can lead to irrelevant retrievals and higher costs.
- Neglecting privacy and security when storing user data in memory.
