Embeddings
An embedding is a vector representation of data, usually text, where semantically similar items map to nearby points. Rig turns your data into these vectors so you can power semantic search, similarity comparison, and Retrieval-Augmented Generation.
In Rig, an Embedding holds the original text (document) alongside its vector (vec: Vec<f64>).
Minimal example
Section titled “Minimal example”Create an embedding model from a provider client with client.embedding(model, ndims) and embed a few strings with EmbeddingsBuilder. Each .document(...) call adds one item to embed. Pass None for ndims to use the model’s default width.
use rig::prelude::*;use rig::providers::openai::{self, OpenAI};
#[tokio::main]async fn main() -> anyhow::Result<()> { // Requires the OPENAI_API_KEY environment variable. let model = OpenAI::from_env()?.embedding(openai::TEXT_EMBEDDING_3_SMALL, None);
let embeddings = EmbeddingsBuilder::new(model) .document("Some text")? .document("More text")? .build() .await?;
println!("Generated {} embeddings", embeddings.len()); Ok(())}Generated 2 embeddingsbuild() returns a Vec<(T, Vec<Embedding>)>: each document paired with its embeddings. The builder splits the work into batches that fit the provider’s limit and runs them concurrently, keeping the original order.
Embedding a single query
Section titled “Embedding a single query”At query time you often need one vector, not a batch. embed_text embeds a single text:
let model = OpenAI::from_env()?.embedding(openai::TEXT_EMBEDDING_3_SMALL, None);
let vector: Vec<f64> = model.embed_text("What is Rig?").await?.vec;println!("{} dimensions", vector.len());For several texts in one request, call the model directly: model.call(texts).await? returns an EmbeddingResponse with one Embedding per input, in order.
Vector stores do this for you when you search them (see Vector Stores & RAG), so you rarely need it by hand.
The Embed trait
Section titled “The Embed trait”To embed your own types, implement Embed, which tells Rig which text to turn into vectors. The easiest way is the derive macro: mark the field(s) to embed with #[embed].
use rig::Embed;
#[derive(Embed)]struct WordDefinition { id: String, word: String, #[embed] definitions: Vec<String>,}A Vec<String> field produces one embedding per element, so a document can have several vectors. The in-memory vector store scores such a document by its best-matching vector.
For full control, implement Embed by hand and push each piece of text into the TextEmbedder:
use rig::embeddings::{Embed, EmbedError, TextEmbedder};
struct WordDefinition { id: String, word: String, definition: String,}
impl Embed for WordDefinition { fn embed(&self, embedder: &mut TextEmbedder) -> Result<(), EmbedError> { // Embed the word together with its definition. embedder.embed(format!("{}: {}", self.word, self.definition)); Ok(()) }}String, &str, numbers, bool, serde_json::Value, and Vec<T: Embed> implement Embed already.
Embedding many documents
Section titled “Embedding many documents”Pass a collection of Embed values to .documents(...):
#[derive(rig::Embed, serde::Serialize, Clone)]struct Doc { id: String, #[embed] text: String,}
let documents = vec![ Doc { id: "doc0".into(), text: "Rig is a Rust library for LLM apps.".into() }, Doc { id: "doc1".into(), text: "Embeddings map text to vectors.".into() },];
let model = OpenAI::from_env()?.embedding(openai::TEXT_EMBEDDING_3_SMALL, None);
let embeddings = EmbeddingsBuilder::new(model) .documents(documents)? .build() .await?;Storing embeddings
Section titled “Storing embeddings”Embeddings become useful once they are searchable. Load them into the built-in InMemoryVectorStore, or insert them into any store that implements InsertDocuments:
// In memory, keyed by generated ids ("doc0", "doc1", ...).let store = InMemoryVectorStore::from_documents(embeddings);
// Turn the store into a searchable index with the same embedding model.let index = store.index(model);use rig::vector_store::InsertDocuments;
store.insert_documents(embeddings).await?;See Vector Stores & RAG for searching an index, and Vector Stores for each supported store.
Best practices
Section titled “Best practices”- Prepare documents. Clean text before embedding, and split large documents into focused chunks.
- Match models. Embed queries with the same model as the stored documents. Vectors from different models are not comparable.
- Batch. Prefer
.documents(...)for many items so Rig can batch requests.
See also
Section titled “See also”- Vector Stores & RAG: use embeddings for retrieval
- Loaders: read files into text ready to embed
- Vector Stores: supported store integrations
