Skip to content

Vector Stores & RAG

Retrieval-Augmented Generation (RAG) looks up documents relevant to a query and includes them in the prompt, grounding the model’s answer in your data. It reduces hallucinations and lets a model use information that isn’t in its training set. Rig provides the building blocks: embeddings, vector stores, and agents that retrieve context on every prompt.

A RAG pipeline has two phases:

  1. Ingestion. Split documents into chunks, embed each chunk, and store the embeddings with the chunk and its metadata in a vector store.
  2. Retrieval. When a question arrives, embed it with the same model, find the nearest stored vectors (usually by cosine similarity), and put the matching chunks into the prompt.

If you are building a support bot or assistant grounded in documents you own, you want RAG. For simple classification with well-known categories, you probably don’t.

Two traits define a vector store:

  • VectorStoreIndex: search for the documents closest to a query (top_n, top_n_ids).
  • InsertDocuments: add embedded documents to the store.

InMemoryVectorStore ships with Rig and needs no external service, which makes it a good fit for development and small datasets. For production, Rig integrates with LanceDB, MongoDB, Neo4j, PostgreSQL, Qdrant, SurrealDB, and more; see Vector Stores.

Attach an index to an agent with dynamic_context(n, index). A hook runs before every model call in the run: it searches the index with the prompt’s text (or, on a turn that only carries tool results, the latest user message’s text) and adds the top n documents to that request as context.

use rig::prelude::*;
use rig::providers::openai::{self, OpenAI};
#[derive(rig::Embed, serde::Serialize, Clone)]
struct WordDefinition {
id: String,
word: String,
#[embed]
definitions: Vec<String>,
}
#[tokio::main]
async fn main() -> anyhow::Result<()> {
let client = OpenAI::from_env()?;
let embedding_model = client.embedding(openai::TEXT_EMBEDDING_3_SMALL, None);
// Ingestion: embed the documents and load them into an in-memory store.
let embeddings = EmbeddingsBuilder::new(embedding_model.clone())
.documents(vec![
WordDefinition {
id: "doc0".into(),
word: "flurbo".into(),
definitions: vec!["A green alien that lives on cold planets.".into()],
},
WordDefinition {
id: "doc1".into(),
word: "glarb-glarb".into(),
definitions: vec!["An ancient farming tool from planet Jiro.".into()],
},
])?
.build()
.await?;
let index = InMemoryVectorStore::from_documents(embeddings).index(embedding_model);
// Retrieval: the agent fetches the best match for every prompt.
let agent = AgentBuilder::new(client.completion(openai::GPT_5_5))
.preamble("You are a dictionary assistant. Use the definitions provided as context.")
.dynamic_context(1, index)
.build();
let answer = agent.prompt("What does \"glarb-glarb\" mean?").await?.output();
println!("{answer}");
Ok(())
}
"Glarb-glarb" is an ancient farming tool that comes from the planet Jiro.

You can call dynamic_context more than once to retrieve from several indexes, and combine it with static documents added through .context(...).

To get the documents yourself, build a VectorSearchRequest and call top_n. Each result carries the similarity score, the document id, and the stored document deserialized into the type you ask for:

let req = VectorSearchRequest::builder()
.query("What type of money can I use in a fictional universe?")
.samples(2)
.threshold(0.3) // optional: drop weak matches
.build();
for result in index.top_n::<WordDefinition>(req).await? {
println!("score={:.2} id={} word={}", result.score, result.id, result.document.word);
}
score=0.81 id=doc0 word=flurbo

Use top_n_ids when you only need ids and scores, for example to fetch full records from your own database.

VectorSearchRequest also takes a metadata filter. Each store translates the generic filter into its own query language:

use rig::vector_store::request::{Filter, SearchFilter};
let req = VectorSearchRequest::builder()
.query("cold planets")
.samples(5)
.filter(Filter::eq("word", json!("flurbo")))
.build();
let results = index.top_n::<serde_json::Value>(req).await?;

To pick your own document ids instead of the generated doc0, doc1, … build the store with InMemoryVectorStore::from_documents_with_id_f(embeddings, |doc| doc.id.clone()).

An agent with many tools wastes context on tool definitions it won’t use, and large tool lists make models choose worse. Rig can store tool definitions in a vector index and send the model only the ones most relevant to each prompt, using AgentBuilder::retrieved_tools(n, index, toolset). The Tools page covers the ToolEmbedding trait and setup.

Vector search is fast but coarse. A reranking model reads the query and each candidate together and reorders them by relevance, which often improves the final context noticeably. Clients that offer reranking (Voyage AI, and OpenAI-shaped clients for endpoints that serve it) build a rerank model with client.rerank(id). You call it with a rig::operation::RerankRequest and get back a rig::rerank::RerankResponse whose results hold each document’s original index and its relevance_score; DynModel<rig::operation::Rerank> is the provider-independent type:

use rig::operation::RerankRequest;
use rig::providers::voyageai::{self, VoyageAi};
// Requires VOYAGE_API_KEY.
let reranker = VoyageAi::from_env()?.rerank(voyageai::RERANK_2_5);
let response = reranker
.call(RerankRequest {
query: "Which animal lives on cold planets?".into(),
documents: vec![
"A glarb-glarb is an ancient farming tool.".into(),
"A flurbo is a green alien that lives on cold planets.".into(),
],
})
.await?;
for result in response.results {
println!("#{} scored {:.3}", result.index, result.relevance_score);
}

A common pattern is to over-fetch with top_n (say 20 results), rerank them, and keep the best few for the prompt.

  • Chunking. Relevant information can span chunk boundaries. Overlap chunks by 10–20%, or retrieve small chunks and pass their larger parent section to the model.
  • Hybrid search. Semantic search can miss exact terms such as product codes. Combine a full-text search with vector search and merge the lists, for example with Reciprocal Rank Fusion.
  • Conflicting or stale data. Filter by metadata, weight by recency or source authority, and re-embed documents when the source changes.
  • Memory. The same machinery can store long-term facts and conversation summaries for agent memory.