Skip to content

MongoDB

rig::mongodb (feature mongodb) runs Rig’s vector search on MongoDB Atlas Vector Search. The search runs inside MongoDB as an aggregation pipeline, so you get persistence and server-side filtering next to the rest of your data.

[dependencies]
rig = { version = "0.44.0", features = ["mongodb"] }
mongodb = "3"
tokio = { version = "1", features = ["full"] }

You need a MongoDB Atlas cluster (Vector Search is an Atlas feature) and its connection string.

The collection needs a vector search index before you can query it. Create one in the Atlas UI or through the API, with numDimensions matching your embedding model (1536 for text-embedding-3-small) and path naming the field that holds the vector:

{
"fields": [
{
"type": "vector",
"path": "embedding",
"numDimensions": 1536,
"similarity": "cosine"
}
]
}

To filter on a field, add it to the index as a "type": "filter" field as well.

Embed documents, write them to the collection, then search through MongoDbVectorIndex:

use mongodb::{Client as MongoClient, Collection, bson::{self, doc}};
use rig::mongodb::{MongoDbVectorIndex, SearchParams};
use rig::prelude::*;
use serde::{Deserialize, Serialize};
use rig::providers::openai::{self, OpenAI};
#[derive(Embed, Clone, Serialize, Deserialize, Debug)]
struct Word {
id: String,
#[embed]
definition: String,
}
#[tokio::main]
async fn main() -> Result<(), anyhow::Error> {
let model = OpenAI::from_env()?.embedding(openai::TEXT_EMBEDDING_3_SMALL, None);
let mongo = MongoClient::with_uri_str(std::env::var("MONGODB_CONNECTION_STRING")?).await?;
let collection: Collection<bson::Document> =
mongo.database("knowledgebase").collection("context");
// Embed the documents and store each with its vector in `embedding`.
let words = vec![
Word { id: "doc0".into(), definition: "A flurbo is a green alien.".into() },
Word { id: "doc1".into(), definition: "A glarb-glarb is an ancient farming tool.".into() },
];
let embeddings = EmbeddingsBuilder::new(model.clone())
.documents(words)?
.build()
.await?;
let records = embeddings
.iter()
.map(|(word, embeddings)| {
doc! {
"id": word.id.clone(),
"definition": word.definition.clone(),
"embedding": embeddings.first().map(|e| e.vec.clone()),
}
})
.collect::<Vec<_>>();
collection.insert_many(records).await?;
// "vector_index" is the Atlas vector search index on this collection.
let index = MongoDbVectorIndex::new(collection, model, "vector_index", SearchParams::new()).await?;
let req = VectorSearchRequest::builder()
.query("What does glarb-glarb mean?")
.samples(1)
.build();
for result in index.top_n::<Word>(req).await? {
println!("{:.3} {}", result.score, result.document.definition);
}
Ok(())
}
  • MongoDbVectorIndex::new(collection, model, index_name, params) checks that the named Atlas index exists and is queryable, and reads the vector field from its definition. It errors if the index is missing or not ready yet.
  • top_n::<T>(req) embeds the query, runs a $vectorSearch stage, adds the search score, and projects the vector field out. T is deserialized from the rest of the stored document, so it must not require the embedding field. result.id is the document’s _id.
  • top_n_ids(req) returns only _id and score.
  • insert_documents (the InsertDocuments trait) writes one record per embedding with the shape { document, embedding, embedded_text }. Point the Atlas index at embedding, and read results back into a type with a document field. Writing records yourself, as above, lets you pick the layout.

SearchParams::new() uses approximate search with numCandidates set to ten times the requested samples. .exact(true) switches to exact search, and .num_candidates(n) overrides the candidate count. A request threshold becomes a minimum score.

MongoDbSearchFilter builds the filter document of the $vectorSearch stage. Its values are BSON:

use mongodb::bson::Bson;
use rig::mongodb::MongoDbSearchFilter;
use rig::vector_store::request::SearchFilter;
let req = VectorSearchRequest::<MongoDbSearchFilter>::builder()
.query("What does glarb-glarb mean?")
.samples(3)
.filter(MongoDbSearchFilter::eq("id", Bson::from("doc1")))
.build();

Besides eq, gt, lt, and and or, it has gte, lte, not, is_type, size, all and any. Filtered fields must be declared as filter fields in the Atlas index.