Skip to content

LanceDB

rig::lancedb (feature lancedb) searches LanceDB tables, a serverless vector database built on Apache Arrow. It runs embedded, on local disk or on object storage (S3, GCS, Azure), with no server to operate.

[dependencies]
rig = { version = "0.44.0", features = ["lancedb"] }
lancedb = "0.30"
arrow-array = "58"
tokio = { version = "1", features = ["full"] }

You create and fill tables with LanceDB’s own API (and Arrow arrays); rig::lancedb::LanceDbVectorIndex is the read side, implementing VectorStoreIndex over an existing table.

lancedb::connect opens (or creates) a database at a URI: a path for local storage, or an s3://, gs:// or az:// URI for object storage.

// Local, on-disk store.
let db = lancedb::connect("data/lancedb-store").execute().await?;
// Object storage on S3 (see LanceDB's storage guide for credentials).
let db = lancedb::connect("s3://my-lancedb-bucket").execute().await?;

A table needs an id column, your document columns, and a fixed-size list column for the embedding whose length matches the model’s width:

use std::sync::Arc;
use lancedb::arrow::arrow_schema::{DataType, Field, Fields, Schema};
fn schema(dims: usize) -> Schema {
Schema::new(Fields::from(vec![
Field::new("id", DataType::Utf8, false),
Field::new("definition", DataType::Utf8, false),
Field::new(
"embedding",
DataType::FixedSizeList(
Arc::new(Field::new("item", DataType::Float64, true)),
dims as i32,
),
false,
),
]))
}

Embed your documents with EmbeddingsBuilder, turn them into an Arrow RecordBatch, and create the table from it:

use std::sync::Arc;
use arrow_array::{ArrayRef, FixedSizeListArray, RecordBatch, StringArray, types::Float64Type};
use rig::embeddings::Embedding;
#[derive(rig::Embed, Clone, serde::Deserialize, Debug)]
struct Word {
id: String,
#[embed]
definition: String,
}
fn as_record_batch(records: Vec<(Word, Vec<Embedding>)>, dims: usize) -> anyhow::Result<RecordBatch> {
let id = StringArray::from_iter_values(records.iter().map(|(w, _)| &w.id));
let definition = StringArray::from_iter_values(records.iter().map(|(w, _)| &w.definition));
let embedding = FixedSizeListArray::from_iter_primitive::<Float64Type, _, _>(
records.into_iter().map(|(_, embeddings)| {
embeddings
.into_iter()
.next()
.map(|e| e.vec.into_iter().map(Some).collect::<Vec<_>>())
}),
dims as i32,
);
Ok(RecordBatch::try_from_iter(vec![
("id", Arc::new(id) as ArrayRef),
("definition", Arc::new(definition) as ArrayRef),
("embedding", Arc::new(embedding) as ArrayRef),
])?)
}
// ...
let embeddings = EmbeddingsBuilder::new(model.clone()).documents(words)?.build().await?;
let dims = model.capabilities().ndims;
let table = db
.create_table("definitions", vec![as_record_batch(embeddings, dims)?])
.execute()
.await?;

LanceDbVectorIndex::new wraps a table with the embedding model (used to embed queries), the name of the id column, and search parameters:

use rig::lancedb::{LanceDbVectorIndex, SearchParams};
use rig::providers::openai::{self, OpenAI};
let model = OpenAI::from_env()?.embedding(openai::TEXT_EMBEDDING_3_SMALL, None);
let db = lancedb::connect("data/lancedb-store").execute().await?;
let table = db.open_table("definitions").execute().await?;
let index = LanceDbVectorIndex::new(table, model, "id", SearchParams::default()).await?;

SearchParams is a builder:

  • distance_type(lancedb::DistanceType::Cosine): the metric, which must match the one the table’s vector index was built with. LanceDB defaults to L2.
  • search_type(SearchType::Flat | SearchType::Approximate): force exact or approximate search. Unset, LanceDB searches approximately when the table has a vector index and exhaustively otherwise.
  • nprobes(n) and refine_factor(n): approximate-search tuning, used only with SearchType::Approximate.
  • post_filter(true): apply filters after the vector search instead of before.
  • column(name): the embedding column, needed only when the table has more than one.

Without a vector index, LanceDB scans every row (exact). For large tables, build an IVF-PQ index for approximate search; LanceDB needs at least 256 rows to train one:

use lancedb::index::{Index, vector::IvfPqIndexBuilder};
if table.index_stats("embedding").await?.is_none() {
table
.create_index(&["embedding"], Index::IvfPq(IvfPqIndexBuilder::default()))
.execute()
.await?;
}
use rig::lancedb::{LanceDbVectorIndex, SearchParams};
use rig::providers::openai::{self, OpenAI};
#[derive(Deserialize, Debug)]
struct Word {
id: String,
definition: String,
}
let req = VectorSearchRequest::builder()
.query("My boss says I zindle too much, what does that mean?")
.samples(3)
.build();
for result in index.top_n::<Word>(req).await? {
println!("{:.3} {} {}", result.score, result.id, result.document.definition);
}

Each row comes back with its embedding columns removed, deserialized into your type. Unlike most stores, score is LanceDB’s distance, so lower is closer, and a request threshold is a maximum distance.

LanceDB filters are SQL predicates. Use LanceDBFilter as the request’s filter type:

use rig::lancedb::LanceDBFilter;
use rig::vector_store::request::SearchFilter;
let req = VectorSearchRequest::<LanceDBFilter>::builder()
.query("search query")
.samples(5)
.filter(LanceDBFilter::eq("id", json!("doc1")).or(LanceDBFilter::like("definition", "%moon%")))
.build();

Besides eq, gt, lt, and and or, it has not, in_values, like, ilike, is_null, is_not_null, between, and array conditions (array_has_any, array_has_all, array_length). Column names are spliced into the SQL as written, so never build them from untrusted input; string values are escaped.