Skip to content

Build a RAG system

Retrieval-augmented generation (RAG) grounds a language model in your own data: when a query arrives, you first retrieve the most relevant passages from a knowledge base, then hand them to the model alongside the question. The model answers from that context instead of relying solely on what it memorized during training, which reduces hallucination and lets it use up-to-date, private, or domain-specific information.

This guide builds a complete RAG system that extracts text from PDFs, embeds it page by page, stores the vectors in memory, and wires the store into an agent as dynamic context, all in under 100 lines. For the concepts behind RAG (embeddings, similarity search, re-ranking, tool-RAG), see Vector Stores & RAG.

Create a new project and add the dependencies:

Terminal window
cargo new rag_system
cd rag_system
[dependencies]
rig = { version = "0.44.0", features = ["pdf"] }
tokio = { version = "1", features = ["full"] }
anyhow = "1"
serde = { version = "1", features = ["derive"] }
  • rig: the Rig library; the pdf feature enables PdfFileLoader.
  • tokio: async runtime.
  • anyhow: ergonomic error handling.
  • serde: the stored pages are serialized into the vector store.

Set your OpenAI API key:

Terminal window
export OPENAI_API_KEY=your_api_key_here

Rig works with plain text, so first turn each PDF into text. Embedding models have an input limit (about 8,000 tokens for OpenAI’s), so a whole book won’t fit in one embedding; splitting by page keeps every chunk small and makes retrieval more precise. PdfFileLoader reads every file matching a glob and yields one string per page:

use rig::Embed;
use rig::loaders::PdfFileLoader;
use serde::Serialize;
/// One page of a PDF. `#[embed]` marks the field that gets embedded.
#[derive(Embed, Serialize, Clone)]
struct Page {
id: String,
#[embed]
text: String,
}
fn load_pages(glob: &str) -> anyhow::Result<Vec<Page>> {
let pages = PdfFileLoader::with_glob(glob)?
.load()
.ignore_errors() // skip files that fail to open
.by_page()
.into_iter()
.collect::<Result<Vec<String>, _>>()?;
Ok(pages
.into_iter()
.enumerate()
.filter(|(_, text)| !text.trim().is_empty())
.map(|(i, text)| Page { id: format!("page-{i}"), text })
.collect())
}

Create an embedding model, embed every page with EmbeddingsBuilder, and load the results into an InMemoryVectorStore. The in-memory store is ideal for small to medium collections; for larger corpora, swap in a persistent store such as LanceDB, MongoDB, Neo4j, or Qdrant.

use rig::embeddings::EmbeddingsBuilder;
use rig::providers::openai::{self, OpenAI};
use rig::vector_store::in_memory_store::InMemoryVectorStore;
let openai_client = OpenAI::from_env()?;
let embedding_model = openai_client
.embedding(openai::TEXT_EMBEDDING_3_SMALL, None)
.erase();
let pages = load_pages("documents/*.pdf")?;
let embeddings = EmbeddingsBuilder::new(embedding_model.clone())
.documents(pages)?
.build()
.await?;
let vector_store = InMemoryVectorStore::from_documents_with_id_f(embeddings, |page| page.id.clone());

EmbeddingsBuilder::new(model).documents(pages)? queues the pages for embedding; .build().await? calls the embedding API in batches. InMemoryVectorStore::from_documents_with_id_f builds the store directly from those embeddings, using each page’s id as its id in the store. The same model embeds the documents now and the queries later; models are cheap to clone. .erase() turns the typed model into a provider-independent DynModel, so the rest of the code doesn’t name OpenAI’s model type.

Turn the store into a searchable index, then attach it to an agent as dynamic context. On every prompt, the agent runs a vector search and injects the top matches into the model’s context automatically:

let index = vector_store.index(embedding_model);
let rag_agent = AgentBuilder::new(openai_client.completion(openai::GPT_5_5))
.preamble("You are a helpful assistant that answers questions using the provided PDF context.")
.dynamic_context(4, index)
.build();

dynamic_context(4, index) tells the agent to retrieve the four most relevant pages for each query. If you need to run searches yourself rather than through an agent, build a request and call top_n on the index:

use rig::vector_store::{VectorSearchRequest, VectorStoreIndex};
let req = VectorSearchRequest::builder()
.query("what did Sam Altman write?")
.samples(2)
.build();
for hit in index.top_n::<Page>(req).await? {
println!("{:.3} {}: {}", hit.score, hit.id, hit.document.text);
}

Rig ships a cli_chatbot helper that wraps any agent in an interactive prompt loop with chat history:

use rig::integrations::cli_chatbot::ChatBotBuilder;
let chatbot = ChatBotBuilder::new().agent(rag_agent).build();
chatbot.run().await?;
use rig::Embed;
use rig::embeddings::EmbeddingsBuilder;
use rig::integrations::cli_chatbot::ChatBotBuilder;
use rig::loaders::PdfFileLoader;
use rig::prelude::*;
use rig::providers::openai::{self, OpenAI};
use rig::vector_store::in_memory_store::InMemoryVectorStore;
use serde::Serialize;
#[derive(Embed, Serialize, Clone)]
struct Page {
id: String,
#[embed]
text: String,
}
fn load_pages(glob: &str) -> anyhow::Result<Vec<Page>> {
let pages = PdfFileLoader::with_glob(glob)?
.load()
.ignore_errors()
.by_page()
.into_iter()
.collect::<Result<Vec<String>, _>>()?;
Ok(pages
.into_iter()
.enumerate()
.filter(|(_, text)| !text.trim().is_empty())
.map(|(i, text)| Page { id: format!("page-{i}"), text })
.collect())
}
#[tokio::main]
async fn main() -> anyhow::Result<()> {
let openai_client = OpenAI::from_env()?;
let embedding_model = openai_client
.embedding(openai::TEXT_EMBEDDING_3_SMALL, None)
.erase();
let embeddings = EmbeddingsBuilder::new(embedding_model.clone())
.documents(load_pages("documents/*.pdf")?)?
.build()
.await?;
let index = InMemoryVectorStore::from_documents_with_id_f(embeddings, |page| page.id.clone())
.index(embedding_model);
let rag_agent = AgentBuilder::new(openai_client.completion(openai::GPT_5_5))
.preamble("You are a helpful assistant that answers questions using the provided PDF context.")
.dynamic_context(4, index)
.build();
let chatbot = ChatBotBuilder::new().agent(rag_agent).build();
chatbot.run().await?;
Ok(())
}

Place a few PDFs (for example Moores_Law_for_Everything.pdf and The_Last_Question.pdf) in a documents/ folder, then run:

Terminal window
cargo run

You now have a chatbot that answers questions grounded in your PDFs: summarizing a document, analyzing themes, or drawing connections across several, because each response is built from the pages the vector search surfaces for that query.

  • Persistent storage: replace InMemoryVectorStore with a dedicated vector store (LanceDB, MongoDB, Neo4j, Qdrant) for large collections, so you embed once instead of on every start.
  • Chunking: pages are a simple chunk boundary; for dense documents, split further into overlapping passages. See Loaders.
  • Model selection: use a cheaper completion model where answer quality allows.
  • Observability: Rig integrates with OpenTelemetry and Langfuse; see Observability.