Skip to content

Memory

Memory is an agent’s ability to reuse information from earlier in a conversation, and across conversations. Without it, every prompt starts from scratch: a user who asks “what’s the status of my order?” and then “cancel it” leaves the model with no idea what “it” refers to. Rig covers both ends: attach a conversation-memory backend and history is loaded and saved for you, or keep the Vec<Message> yourself and decide exactly what the model sees.

There are two layers:

  • Conversation history (short-term): the messages of the current conversation, sent back to the model on every turn and bounded so they fit the context window.
  • Long-term memory: facts, observations, and user profiles that outlive a session, usually kept in a database or vector store.

Give the agent a memory backend with .memory(...) and pick a conversation id per request with .conversation(...). Before the run, Rig loads that conversation’s history; after it succeeds, Rig appends the new turn, including any tool calls and their results.

use rig::memory::InMemoryConversationMemory;
use rig::providers::openai::{self, OpenAI};
let agent = AgentBuilder::new(OpenAI::from_env()?.completion(openai::GPT_5_5))
.preamble("You are a helpful assistant.")
.memory(InMemoryConversationMemory::new())
.build();
// Each conversation id keeps its own history.
agent.prompt("My name is Ada.").conversation("user-42").await?;
let reply = agent.prompt("What's my name?").conversation("user-42").await?;
println!("{}", reply.output());
Your name is Ada.

Memory works the same way for streaming: agent.prompt(p).conversation(id).stream() saves the turn once the stream finishes.

How memory interacts with the rest of the prompt API:

  • Set the id per request with .conversation("..."), or give the agent a default with AgentBuilder::conversation("..."). No id means no memory.
  • Passing explicit history with .history(...), or calling agent.chat(...), bypasses memory: nothing is loaded and nothing is saved.
  • .without_memory() turns memory off for one request.
  • If loading fails, the run fails with PromptError::Memory before any model call. If saving fails, you still get the answer; response.memory_append() tells you whether the backend acknowledged the write.

InMemoryConversationMemory lives in process memory, so it suits tests and short-lived agents but forgets everything on restart. For durable sessions, implement the ConversationMemory trait over your own store:

pub trait ConversationMemory: WasmCompatSend + WasmCompatSync {
/// Load the full history for a conversation (empty Vec if none).
fn load<'a>(&'a self, conversation_id: &'a ConversationId)
-> WasmBoxedFuture<'a, Result<Vec<Message>, MemoryError>>;
/// Append the messages of a successful turn.
fn append<'a>(&'a self, conversation_id: &'a ConversationId, messages: Vec<Message>)
-> WasmBoxedFuture<'a, Result<(), MemoryError>>;
/// Remove all stored messages for a conversation.
fn clear<'a>(&'a self, conversation_id: &'a ConversationId)
-> WasmBoxedFuture<'a, Result<(), MemoryError>>;
}

Message implements Serialize and Deserialize, so persisting a conversation is ordinary serde work; a JSON column per conversation id is a fine first backend. Return backend failures with MemoryError::backend(err). Keep append cheap: it runs before the agent returns its response.

Raw history grows without limit, and long histories are worse than just expensive: models get distracted by stale content, and eventually the conversation no longer fits the context window. The fix is to shape what load returns so the model sees a bounded, relevant window.

Enable the memory feature to get Rig’s reusable policies in rig::memory:

rig = { version = "0.44.0", features = ["memory"] }

The simplest policy is a sliding window over the most recent messages:

use rig::memory::{InMemoryConversationMemory, IntoFilter, SlidingWindowMemory};
let memory = InMemoryConversationMemory::new()
.with_filter(SlidingWindowMemory::last_messages(20).into_filter());

TokenWindowMemory bounds by estimated token count instead of message count, which tracks what you pay for:

use rig::memory::{HeuristicTokenCounter, InMemoryConversationMemory, IntoFilter, TokenWindowMemory};
let memory = InMemoryConversationMemory::new().with_filter(
TokenWindowMemory::new(4_000, HeuristicTokenCounter::default()).into_filter(),
);

Both policies also drop a leading tool result whose matching tool call was cut off, because most providers reject a tool result without its call.

with_filter is specific to the in-memory backend. To put a policy in front of any backend, wrap it with PolicyMemory::new(backend, policy).

Truncation throws the old turns away. Two adapters keep something from them:

  • DemotingPolicyMemory hands evicted messages to a DemotionHook, so you can archive them (a vector store for semantic recall, cold storage for audit) instead of losing them.
  • CompactingMemory replaces evicted messages with a summary placed at the start of the history, so the model keeps a compressed view of the whole conversation:
use rig::memory::{CompactingMemory, InMemoryConversationMemory, SlidingWindowMemory, TemplateCompactor};
use rig::providers::openai::{self, OpenAI};
let memory = CompactingMemory::new(
InMemoryConversationMemory::new(),
SlidingWindowMemory::last_messages(20),
TemplateCompactor::new(), // plain-text rollup, no model call
);
let agent = AgentBuilder::new(OpenAI::from_env()?.completion(openai::GPT_5_5))
.preamble("You are a helpful assistant.")
.memory(memory)
.build();

For better summaries, implement the Compactor trait with a model call. Its carry_over argument is the previous summary, so each compaction can fold in what came before. Compactors and demotion hooks run while history loads, so a slow one delays the agent’s next turn.

When you want full control, keep the history yourself as a Vec<Message>. agent.chat(prompt, &mut history) sends the history with the prompt and then appends the new turn (user prompt, any tool calls and results, and the answer) to your vector:

use rig::providers::openai::{self, OpenAI};
let agent = AgentBuilder::new(OpenAI::from_env()?.completion(openai::GPT_5_5))
.preamble("You are a helpful assistant. Be concise.")
.build();
let mut history: Vec<Message> = Vec::new();
agent.chat("My name is Ada.", &mut history).await?;
let reply = agent.chat("What's my name?", &mut history).await?;
println!("{}", reply.output());

If you want to decide what gets recorded, use .history(...) instead. It sends the history but leaves your vector alone; the messages the run produced are in response.messages:

let mut history = vec![
Message::user("What is the Rust programming language?"),
Message::assistant("Rust is a systems programming language..."),
];
let response = agent.prompt("Who created it?").history(history.clone()).await?;
history.extend(response.messages);

To bound a hand-managed history, apply a policy directly with MemoryPolicy::apply, for example SlidingWindowMemory::last_messages(20).apply(history)?.

For a complete interactive loop, see Build a CLI chatbot.

Bounded history keeps one conversation healthy, but many applications need memory that survives across sessions. Common kinds:

  • Conversation observations: what was decided, open questions, topics of strong interest.
  • User profile: stable facts about the user, such as stated preferences or location. Keep these separate from conversation history, update them incrementally, and re-check them before relying on them.
  • Grounded facts: verifiable data gathered during a session (retrieved documents, computed results, API responses), stored with source and time.

The mechanics are the same for all three. After a significant exchange, use an extractor or a plain prompt to distill what matters, then store it. When a new conversation starts, fetch the most relevant items and add them to the preamble or the opening messages. For semantic lookup, embed memories into a vector store and retrieve them with dynamic_context. A DemotionHook is a natural place to feed evicted turns into such a store.