Skip to content

Loaders

Loaders read files from disk (or bytes) and turn them into text you can give an agent as context or embed into a vector store. They handle glob matching, directory traversal, and per-file errors, so one bad file doesn’t stop a batch. Rig ships loaders for:

  • Any text file: FileLoader
  • PDFs: PdfFileLoader (the pdf feature)
  • ePub books: EpubFileLoader (the epub feature)

If the model can read the file format itself (PDFs and images on most large providers), you can also skip text extraction and send the file in the message.

FileLoader handles text files. Point it at a glob, a directory, or raw bytes, then read the contents. ignore_errors() skips files that fail to load instead of returning their errors:

use rig::loaders::FileLoader;
// Glob: every Rust file in a directory, with each file's path.
let examples = FileLoader::with_glob("examples/*.rs")?
.read_with_path() // yields Result<(PathBuf, String), _>
.ignore_errors()
.into_iter();
// Directory: every file directly inside a folder.
let dir_files = FileLoader::with_dir("data/")?
.read() // yields Result<String, _>
.ignore_errors();
// Bytes: content from any source, e.g. a download.
let from_bytes = FileLoader::from_bytes(b"hello".to_vec());

Use read() when you only need the content and read_with_path() when you also want each file’s path, for example to label context.

A common use is folding loaded files into an agent’s context:

use rig::loaders::FileLoader;
use rig::providers::openai::{self, OpenAI};
let model = OpenAI::from_env()?.completion(openai::GPT_5_5);
let agent = FileLoader::with_glob("examples/*.rs")?
.read_with_path()
.ignore_errors()
.into_iter()
.fold(AgentBuilder::new(model), |builder, (path, content)| {
builder.context(format!("Rust example {path:?}:\n{content}"))
})
.build();

PdfFileLoader has the same shape as FileLoader and adds page-by-page extraction. Use load() / load_with_path() to get parsed documents, then by_page() to split them into pages:

use rig::loaders::PdfFileLoader;
let pages = PdfFileLoader::with_glob("docs/*.pdf")?
.load_with_path() // yields Result<(PathBuf, Document), _>
.ignore_errors()
.by_page() // each document's pages, numbered
.ignore_errors() // yields (PathBuf, Vec<(usize, String)>)
.into_iter();
for (path, doc_pages) in pages {
for (page_no, text) in doc_pages {
println!("{} page {page_no}: {} chars", path.display(), text.len());
}
}

read() and read_with_path() return each PDF’s full text in one string, which is handy when you chunk the text yourself before embedding it.

EpubFileLoader works the same way and splits books by chapter. Its second type parameter picks how chapter markup is handled: RawTextProcessor (the default) keeps it, StripXmlProcessor strips the XML tags:

use rig::loaders::epub::EpubLoaderError;
use rig::loaders::{EpubFileLoader, StripXmlProcessor};
fn count_chapters() -> Result<(), EpubLoaderError> {
let chapters = EpubFileLoader::<_, StripXmlProcessor>::with_glob("books/*.epub")?
.load_with_path()
.ignore_errors()
.by_chapter() // yields (PathBuf, Vec<(usize, String)>)
.ignore_errors()
.into_iter();
for (path, book) in chapters {
println!("{}: {} chapters", path.display(), book.len());
}
Ok(())
}

EpubLoaderError isn’t Send, so handle it where it occurs rather than boxing it into a thread-safe error type.

Many models accept documents and images directly. Build a user message from UserContent parts and prompt with it. There is one constructor per source:

  • UserContent::document_text(text, media_type) for text formats such as Markdown or CSV,
  • document_base64, document_raw (bytes), and document_url for binary documents such as PDFs,
  • the matching image_*, audio_*, and video_* constructors for other media.
use rig::message::{DocumentMediaType, ImageMediaType, UserContent};
use rig::providers::anthropic::{self, Anthropic};
let agent = AgentBuilder::new(Anthropic::from_env()?.completion(anthropic::CLAUDE_SONNET_5_5))
.build();
let pdf = std::fs::read("report.pdf")?;
let chart = std::fs::read("chart.png")?;
let message = Message::User {
content: vec![
UserContent::text("Summarize the report and explain the chart."),
UserContent::document_raw(pdf, Some(DocumentMediaType::PDF)),
UserContent::image_raw(chart, Some(ImageMediaType::PNG), None),
],
};
let answer = agent.prompt(message).await?.output();
println!("{answer}");

Rig encodes each part the way the provider expects. Support for each media type varies by provider and model; check the provider pages.