Loaders
Loaders read files from disk (or bytes) and turn them into text you can give an agent as context or embed into a vector store. They handle glob matching, directory traversal, and per-file errors, so one bad file doesn’t stop a batch. Rig ships loaders for:
- Any text file:
FileLoader - PDFs:
PdfFileLoader(thepdffeature) - ePub books:
EpubFileLoader(theepubfeature)
If the model can read the file format itself (PDFs and images on most large providers), you can also skip text extraction and send the file in the message.
FileLoader
Section titled “FileLoader”FileLoader handles text files. Point it at a glob, a directory, or raw bytes,
then read the contents. ignore_errors() skips files that fail to load instead
of returning their errors:
use rig::loaders::FileLoader;
// Glob: every Rust file in a directory, with each file's path.let examples = FileLoader::with_glob("examples/*.rs")? .read_with_path() // yields Result<(PathBuf, String), _> .ignore_errors() .into_iter();
// Directory: every file directly inside a folder.let dir_files = FileLoader::with_dir("data/")? .read() // yields Result<String, _> .ignore_errors();
// Bytes: content from any source, e.g. a download.let from_bytes = FileLoader::from_bytes(b"hello".to_vec());Use read() when you only need the content and read_with_path() when you
also want each file’s path, for example to label context.
Loading agent context
Section titled “Loading agent context”A common use is folding loaded files into an agent’s context:
use rig::loaders::FileLoader;use rig::providers::openai::{self, OpenAI};
let model = OpenAI::from_env()?.completion(openai::GPT_5_5);
let agent = FileLoader::with_glob("examples/*.rs")? .read_with_path() .ignore_errors() .into_iter() .fold(AgentBuilder::new(model), |builder, (path, content)| { builder.context(format!("Rust example {path:?}:\n{content}")) }) .build();PdfFileLoader
Section titled “PdfFileLoader”PdfFileLoader has the same shape as FileLoader and adds page-by-page
extraction. Use load() / load_with_path() to get parsed documents, then
by_page() to split them into pages:
use rig::loaders::PdfFileLoader;
let pages = PdfFileLoader::with_glob("docs/*.pdf")? .load_with_path() // yields Result<(PathBuf, Document), _> .ignore_errors() .by_page() // each document's pages, numbered .ignore_errors() // yields (PathBuf, Vec<(usize, String)>) .into_iter();
for (path, doc_pages) in pages { for (page_no, text) in doc_pages { println!("{} page {page_no}: {} chars", path.display(), text.len()); }}read() and read_with_path() return each PDF’s full text in one string,
which is handy when you chunk the text yourself before embedding it.
EpubFileLoader
Section titled “EpubFileLoader”EpubFileLoader works the same way and splits books by chapter. Its second type
parameter picks how chapter markup is handled: RawTextProcessor (the default)
keeps it, StripXmlProcessor strips the XML tags:
use rig::loaders::epub::EpubLoaderError;use rig::loaders::{EpubFileLoader, StripXmlProcessor};
fn count_chapters() -> Result<(), EpubLoaderError> { let chapters = EpubFileLoader::<_, StripXmlProcessor>::with_glob("books/*.epub")? .load_with_path() .ignore_errors() .by_chapter() // yields (PathBuf, Vec<(usize, String)>) .ignore_errors() .into_iter();
for (path, book) in chapters { println!("{}: {} chapters", path.display(), book.len()); } Ok(())}EpubLoaderError isn’t Send, so handle it where it occurs rather than
boxing it into a thread-safe error type.
Sending files to the model
Section titled “Sending files to the model”Many models accept documents and images directly. Build a user message from
UserContent parts and prompt with it. There is one constructor per source:
UserContent::document_text(text, media_type)for text formats such as Markdown or CSV,document_base64,document_raw(bytes), anddocument_urlfor binary documents such as PDFs,- the matching
image_*,audio_*, andvideo_*constructors for other media.
use rig::message::{DocumentMediaType, ImageMediaType, UserContent};use rig::providers::anthropic::{self, Anthropic};
let agent = AgentBuilder::new(Anthropic::from_env()?.completion(anthropic::CLAUDE_SONNET_5_5)) .build();
let pdf = std::fs::read("report.pdf")?;let chart = std::fs::read("chart.png")?;
let message = Message::User { content: vec![ UserContent::text("Summarize the report and explain the chart."), UserContent::document_raw(pdf, Some(DocumentMediaType::PDF)), UserContent::image_raw(chart, Some(ImageMediaType::PNG), None), ],};
let answer = agent.prompt(message).await?.output();println!("{answer}");Rig encodes each part the way the provider expects. Support for each media type varies by provider and model; check the provider pages.
See also
Section titled “See also”- Agents: feed loaded files in as context
- Vector Stores & RAG: embed and index loaded documents for retrieval
