Memory
Memory is an agent’s ability to reuse information from earlier in a
conversation, and across conversations. Without it, every prompt starts from
scratch: a user who asks “what’s the status of my order?” and then “cancel it”
leaves the model with no idea what “it” refers to. Rig covers both ends: attach
a conversation-memory backend and history is loaded and saved for you, or keep
the Vec<Message> yourself and decide exactly what the model sees.
There are two layers:
- Conversation history (short-term): the messages of the current conversation, sent back to the model on every turn and bounded so they fit the context window.
- Long-term memory: facts, observations, and user profiles that outlive a session, usually kept in a database or vector store.
Automatic conversation memory
Section titled “Automatic conversation memory”Give the agent a memory backend with .memory(...) and pick a conversation id
per request with .conversation(...). Before the run, Rig loads that
conversation’s history; after it succeeds, Rig appends the new turn, including
any tool calls and their results.
use rig::memory::InMemoryConversationMemory;use rig::providers::openai::{self, OpenAI};
let agent = AgentBuilder::new(OpenAI::from_env()?.completion(openai::GPT_5_5)) .preamble("You are a helpful assistant.") .memory(InMemoryConversationMemory::new()) .build();
// Each conversation id keeps its own history.agent.prompt("My name is Ada.").conversation("user-42").await?;let reply = agent.prompt("What's my name?").conversation("user-42").await?;println!("{}", reply.output());Your name is Ada.Memory works the same way for streaming: agent.prompt(p).conversation(id).stream()
saves the turn once the stream finishes.
How memory interacts with the rest of the prompt API:
- Set the id per request with
.conversation("..."), or give the agent a default withAgentBuilder::conversation("..."). No id means no memory. - Passing explicit history with
.history(...), or callingagent.chat(...), bypasses memory: nothing is loaded and nothing is saved. .without_memory()turns memory off for one request.- If loading fails, the run fails with
PromptError::Memorybefore any model call. If saving fails, you still get the answer;response.memory_append()tells you whether the backend acknowledged the write.
InMemoryConversationMemory lives in process memory, so it suits tests and
short-lived agents but forgets everything on restart. For durable sessions,
implement the ConversationMemory
trait over your own store:
pub trait ConversationMemory: WasmCompatSend + WasmCompatSync { /// Load the full history for a conversation (empty Vec if none). fn load<'a>(&'a self, conversation_id: &'a ConversationId) -> WasmBoxedFuture<'a, Result<Vec<Message>, MemoryError>>;
/// Append the messages of a successful turn. fn append<'a>(&'a self, conversation_id: &'a ConversationId, messages: Vec<Message>) -> WasmBoxedFuture<'a, Result<(), MemoryError>>;
/// Remove all stored messages for a conversation. fn clear<'a>(&'a self, conversation_id: &'a ConversationId) -> WasmBoxedFuture<'a, Result<(), MemoryError>>;}Message implements Serialize and Deserialize, so persisting a conversation
is ordinary serde work; a JSON column per conversation id is a fine first
backend. Return backend failures with MemoryError::backend(err). Keep
append cheap: it runs before the agent returns its response.
Bounding history with policies
Section titled “Bounding history with policies”Raw history grows without limit, and long histories are worse than just
expensive: models get distracted by stale content, and eventually the
conversation no longer fits the context window. The fix is to shape what load
returns so the model sees a bounded, relevant window.
Enable the memory feature to get Rig’s reusable policies in rig::memory:
rig = { version = "0.44.0", features = ["memory"] }The simplest policy is a sliding window over the most recent messages:
use rig::memory::{InMemoryConversationMemory, IntoFilter, SlidingWindowMemory};
let memory = InMemoryConversationMemory::new() .with_filter(SlidingWindowMemory::last_messages(20).into_filter());TokenWindowMemory bounds by estimated token count instead of message count,
which tracks what you pay for:
use rig::memory::{HeuristicTokenCounter, InMemoryConversationMemory, IntoFilter, TokenWindowMemory};
let memory = InMemoryConversationMemory::new().with_filter( TokenWindowMemory::new(4_000, HeuristicTokenCounter::default()).into_filter(),);Both policies also drop a leading tool result whose matching tool call was cut off, because most providers reject a tool result without its call.
with_filter is specific to the in-memory backend. To put a policy in front of
any backend, wrap it with PolicyMemory::new(backend, policy).
Truncation throws the old turns away. Two adapters keep something from them:
DemotingPolicyMemoryhands evicted messages to aDemotionHook, so you can archive them (a vector store for semantic recall, cold storage for audit) instead of losing them.CompactingMemoryreplaces evicted messages with a summary placed at the start of the history, so the model keeps a compressed view of the whole conversation:
use rig::memory::{CompactingMemory, InMemoryConversationMemory, SlidingWindowMemory, TemplateCompactor};use rig::providers::openai::{self, OpenAI};
let memory = CompactingMemory::new( InMemoryConversationMemory::new(), SlidingWindowMemory::last_messages(20), TemplateCompactor::new(), // plain-text rollup, no model call);
let agent = AgentBuilder::new(OpenAI::from_env()?.completion(openai::GPT_5_5)) .preamble("You are a helpful assistant.") .memory(memory) .build();For better summaries, implement the
Compactor trait
with a model call. Its carry_over argument is the previous summary, so each
compaction can fold in what came before. Compactors and demotion hooks run while
history loads, so a slow one delays the agent’s next turn.
Managing history yourself
Section titled “Managing history yourself”When you want full control, keep the history yourself as a Vec<Message>.
agent.chat(prompt, &mut history) sends the history with the prompt and then
appends the new turn (user prompt, any tool calls and results, and the answer)
to your vector:
use rig::providers::openai::{self, OpenAI};
let agent = AgentBuilder::new(OpenAI::from_env()?.completion(openai::GPT_5_5)) .preamble("You are a helpful assistant. Be concise.") .build();
let mut history: Vec<Message> = Vec::new();agent.chat("My name is Ada.", &mut history).await?;let reply = agent.chat("What's my name?", &mut history).await?;println!("{}", reply.output());If you want to decide what gets recorded, use .history(...) instead. It sends
the history but leaves your vector alone; the messages the run produced are in
response.messages:
let mut history = vec![ Message::user("What is the Rust programming language?"), Message::assistant("Rust is a systems programming language..."),];
let response = agent.prompt("Who created it?").history(history.clone()).await?;history.extend(response.messages);To bound a hand-managed history, apply a policy directly with
MemoryPolicy::apply, for example SlidingWindowMemory::last_messages(20).apply(history)?.
For a complete interactive loop, see Build a CLI chatbot.
Long-term memory
Section titled “Long-term memory”Bounded history keeps one conversation healthy, but many applications need memory that survives across sessions. Common kinds:
- Conversation observations: what was decided, open questions, topics of strong interest.
- User profile: stable facts about the user, such as stated preferences or location. Keep these separate from conversation history, update them incrementally, and re-check them before relying on them.
- Grounded facts: verifiable data gathered during a session (retrieved documents, computed results, API responses), stored with source and time.
The mechanics are the same for all three. After a significant exchange, use an
extractor or a plain prompt to distill what matters,
then store it. When a new conversation starts, fetch the most relevant items
and add them to the preamble or the opening messages. For semantic lookup, embed
memories into a vector store and retrieve them with
dynamic_context. A DemotionHook is a natural place to
feed evicted turns into such a store.
See also
Section titled “See also”- Agents: where memory,
.history(...), andchatplug in - Vector Stores & RAG: retrieve long-term memories semantically
- Build a CLI chatbot: a working conversational loop
