Vector Stores & RAG
Retrieval-Augmented Generation (RAG) looks up documents relevant to a query and includes them in the prompt, grounding the model’s answer in your data. It reduces hallucinations and lets a model use information that isn’t in its training set. Rig provides the building blocks: embeddings, vector stores, and agents that retrieve context on every prompt.
How RAG works
Section titled “How RAG works”A RAG pipeline has two phases:
- Ingestion. Split documents into chunks, embed each chunk, and store the embeddings with the chunk and its metadata in a vector store.
- Retrieval. When a question arrives, embed it with the same model, find the nearest stored vectors (usually by cosine similarity), and put the matching chunks into the prompt.
If you are building a support bot or assistant grounded in documents you own, you want RAG. For simple classification with well-known categories, you probably don’t.
RAG in Rig
Section titled “RAG in Rig”Two traits define a vector store:
VectorStoreIndex: search for the documents closest to a query (top_n,top_n_ids).InsertDocuments: add embedded documents to the store.
InMemoryVectorStore ships with Rig and needs no external service, which makes it a good fit for development and small datasets. For production, Rig integrates with LanceDB, MongoDB, Neo4j, PostgreSQL, Qdrant, SurrealDB, and more; see Vector Stores.
Minimal RAG agent
Section titled “Minimal RAG agent”Attach an index to an agent with dynamic_context(n, index). A hook runs before every model call in the run: it searches the index with the prompt’s text (or, on a turn that only carries tool results, the latest user message’s text) and adds the top n documents to that request as context.
use rig::prelude::*;use rig::providers::openai::{self, OpenAI};
#[derive(rig::Embed, serde::Serialize, Clone)]struct WordDefinition { id: String, word: String, #[embed] definitions: Vec<String>,}
#[tokio::main]async fn main() -> anyhow::Result<()> { let client = OpenAI::from_env()?; let embedding_model = client.embedding(openai::TEXT_EMBEDDING_3_SMALL, None);
// Ingestion: embed the documents and load them into an in-memory store. let embeddings = EmbeddingsBuilder::new(embedding_model.clone()) .documents(vec![ WordDefinition { id: "doc0".into(), word: "flurbo".into(), definitions: vec!["A green alien that lives on cold planets.".into()], }, WordDefinition { id: "doc1".into(), word: "glarb-glarb".into(), definitions: vec!["An ancient farming tool from planet Jiro.".into()], }, ])? .build() .await?; let index = InMemoryVectorStore::from_documents(embeddings).index(embedding_model);
// Retrieval: the agent fetches the best match for every prompt. let agent = AgentBuilder::new(client.completion(openai::GPT_5_5)) .preamble("You are a dictionary assistant. Use the definitions provided as context.") .dynamic_context(1, index) .build();
let answer = agent.prompt("What does \"glarb-glarb\" mean?").await?.output(); println!("{answer}"); Ok(())}"Glarb-glarb" is an ancient farming tool that comes from the planet Jiro.You can call dynamic_context more than once to retrieve from several indexes, and combine it with static documents added through .context(...).
Searching an index directly
Section titled “Searching an index directly”To get the documents yourself, build a VectorSearchRequest and call top_n. Each result carries the similarity score, the document id, and the stored document deserialized into the type you ask for:
let req = VectorSearchRequest::builder() .query("What type of money can I use in a fictional universe?") .samples(2) .threshold(0.3) // optional: drop weak matches .build();
for result in index.top_n::<WordDefinition>(req).await? { println!("score={:.2} id={} word={}", result.score, result.id, result.document.word);}score=0.81 id=doc0 word=flurboUse top_n_ids when you only need ids and scores, for example to fetch full records from your own database.
VectorSearchRequest also takes a metadata filter. Each store translates the generic filter into its own query language:
use rig::vector_store::request::{Filter, SearchFilter};
let req = VectorSearchRequest::builder() .query("cold planets") .samples(5) .filter(Filter::eq("word", json!("flurbo"))) .build();let results = index.top_n::<serde_json::Value>(req).await?;To pick your own document ids instead of the generated doc0, doc1, … build the store with InMemoryVectorStore::from_documents_with_id_f(embeddings, |doc| doc.id.clone()).
Retrieved tools
Section titled “Retrieved tools”An agent with many tools wastes context on tool definitions it won’t use, and large tool lists make models choose worse. Rig can store tool definitions in a vector index and send the model only the ones most relevant to each prompt, using AgentBuilder::retrieved_tools(n, index, toolset). The Tools page covers the ToolEmbedding trait and setup.
Reranking
Section titled “Reranking”Vector search is fast but coarse. A reranking model reads the query and each candidate together and reorders them by relevance, which often improves the final context noticeably. Clients that offer reranking (Voyage AI, and OpenAI-shaped clients for endpoints that serve it) build a rerank model with client.rerank(id). You call it with a rig::operation::RerankRequest and get back a rig::rerank::RerankResponse whose results hold each document’s original index and its relevance_score; DynModel<rig::operation::Rerank> is the provider-independent type:
use rig::operation::RerankRequest;use rig::providers::voyageai::{self, VoyageAi};
// Requires VOYAGE_API_KEY.let reranker = VoyageAi::from_env()?.rerank(voyageai::RERANK_2_5);
let response = reranker .call(RerankRequest { query: "Which animal lives on cold planets?".into(), documents: vec![ "A glarb-glarb is an ancient farming tool.".into(), "A flurbo is a green alien that lives on cold planets.".into(), ], }) .await?;
for result in response.results { println!("#{} scored {:.3}", result.index, result.relevance_score);}A common pattern is to over-fetch with top_n (say 20 results), rerank them, and keep the best few for the prompt.
Design tips
Section titled “Design tips”- Chunking. Relevant information can span chunk boundaries. Overlap chunks by 10–20%, or retrieve small chunks and pass their larger parent section to the model.
- Hybrid search. Semantic search can miss exact terms such as product codes. Combine a full-text search with vector search and merge the lists, for example with Reciprocal Rank Fusion.
- Conflicting or stale data. Filter by metadata, weight by recency or source authority, and re-embed documents when the source changes.
- Memory. The same machinery can store long-term facts and conversation summaries for agent memory.
See also
Section titled “See also”- Embeddings: the vectors RAG is built on
- Build a RAG system: an end-to-end tutorial
- Vector Stores: supported stores and setup
- Tools: retrieved tools with
ToolEmbedding
