Skip to content

Agents

An Agent is Rig’s primary building block for working with LLMs. It bundles a completion model together with a system prompt, optional context documents, and a set of tools, then runs the agent loop for you: it sends your prompt to the model, executes any tools the model calls, feeds the results back, and repeats until the model produces an answer. Reach for an agent whenever you want to prompt a model without writing that loop by hand — from a simple chatbot to a tool-using assistant or a RAG system.

use rig::prelude::*;
use rig::providers::openai::{self, OpenAI};
#[tokio::main]
async fn main() -> anyhow::Result<()> {
let model = OpenAI::from_env()?.completion(openai::GPT_5_5);
let agent = AgentBuilder::new(model)
.preamble("You are a helpful assistant.")
.temperature(0.7)
.build();
let response = agent.prompt("Hello!").await?;
println!("{}", response.output());
Ok(())
}
Hello! How can I help you today?

You build an agent with AgentBuilder::new(model)...build() and run it with agent.prompt(...).await?. The result is a PromptResponse; output() returns the final answer as a String (the response also implements Display, so println!("{response}") prints the same text). Everything else on this page is optional configuration layered on top.

AgentBuilder collects the reusable configuration that every prompt to the agent shares:

  • Model and request settings — the completion model, the system prompt (preamble), sampling parameters such as temperature and max_tokens, portable generation options such as reasoning(..), and the default turn budget (default_max_turns).
  • Context — documents added to every request. Static context is always sent; dynamic context is retrieved from a vector store per prompt.
  • Tools — capabilities the model can call. Static tools are always offered; retrieved tools are picked from a vector store per prompt.
  • Conversation memory (optional) — a backend that loads prior history before each prompt and saves the new turn after it. See Conversations.
  • Default hooks (optional) — code that observes or steers every run. See Hooks.

An Agent is Clone and safe to share across tasks (Send + Sync + 'static), so build it once and reuse it.

Everything an agent does is one loop. When you prompt it, Rig:

  1. Builds a completion request from the preamble, static context, any dynamic context retrieved from a vector store, the conversation history, and the definitions of every tool on offer.
  2. Sends it to the model. One model call is a turn.
  3. Inspects the response.
    • If the model answered with text, the loop ends and that text is your result.
    • If the model requested tool calls, Rig executes each one — parsing the JSON arguments into your Args type, running your call implementation, and appending the result to the conversation as a tool-result message. A model may request several tool calls in a single turn; Rig runs them and returns all the results together.
  4. Repeats from step 2 with the updated history, so the model can use the tool results — until it produces a text answer or the turn budget runs out.
flowchart TD
P["agent.prompt(...)"] --> M["Send request to model<br/>(preamble + context + history + tools)"]
M --> D{"Model response"}
D -- "text" --> F["Return the answer"]
D -- "tool calls" --> T["Rig executes the tools,<br/>appends results to history"]
T --> B{"Turn budget left?"}
B -- "yes" --> M
B -- "no" --> E["Err(PromptError::MaxTurns)"]

These docs call this the agent loop; you may also see “tool-calling loop” elsewhere — same thing.

The turn budget is the agent’s guard against runaway loops: every turn costs latency and tokens, so the loop refuses to run forever. max_turns is the exact total number of model calls a run may make, counting the first call, every call after a round of tool results, and every retry.

The default budget is 1: the agent makes one model call. That is enough for plain question and answer, but a prompt that calls a tool and then answers needs at least 2. When the budget runs out before the model answers, the prompt fails with PromptError::MaxTurns, which carries the history accumulated so far so you can inspect what happened.

Set a budget per prompt with .max_turns(n), or a default for every prompt with AgentBuilder::default_max_turns(n):

let response = tool_agent
.prompt("Please calculate 2 + 5, then multiply the result by 3")
.max_turns(5) // two rounds of tool calls plus the answer, with room to spare
.await?;
println!("{response}");
2 + 5 = 7, and 7 × 3 = 21.

Rig imposes no upper bound, but pick a number that matches the longest tool chain you actually expect, so a confused model fails fast instead of burning tokens. For handling the error itself, see Error Handling.

Two distinct failure modes, with different behavior:

  • Your tool returns an error. The run doesn’t end: Rig sends a failed tool result to the model and the loop continues, so the model can retry or explain the failure. By default the model only sees generic feedback (the tool failed); to give it a useful message, return a ToolExecutionError or override map_error. See When a tool fails.
  • The model emits a call Rig can’t dispatch — a tool name that doesn’t exist, one the active ToolChoice doesn’t allow, or arguments that aren’t a JSON object. An unknown or disallowed tool fails the prompt with PromptError::UnknownToolCall by default; malformed arguments are answered with corrective feedback so the model can try again. To recover from unknown names instead — retry with feedback, repair the name, or skip the call — implement on_invalid_tool_call on a hook.

To observe, rewrite, or veto individual tool calls yourself — for logging, approval flows, or guardrails — use Hooks.

agent.prompt(p) doesn’t send anything yet: it returns an AgentRunner for one run. Chain per-run settings onto it, then pick how to drive it:

  • .await (or .run().await) runs the loop to completion and returns a PromptResponse.
  • .stream() yields text deltas, tool calls, tool results, and the final response as they happen — see Streaming.

The PromptResponse holds more than the text. Use it for budgets, dashboards, per-user accounting, or to continue the conversation:

let response = agent.prompt("What is 2 + 2?").max_turns(3).await?;
println!("answer: {}", response.output());
println!(
"tokens: {:?} in / {:?} out across {} model calls",
response.usage.input_tokens,
response.usage.output_tokens,
response.requests(),
);
println!("messages produced by this run: {}", response.messages.len());
answer: 2 + 2 = 4.
tokens: Some(28) in / Some(12) out across 1 model calls
messages produced by this run: 2
  • usage aggregates token counts across every turn of the loop. Each counter is an Option: None means the provider didn’t report it.
  • completion_calls breaks usage down per model call (the last entry tells you how large the final request’s context was), along with each call’s finish reason and raw provider response.
  • messages holds the messages this run committed — the prompt, any tool calls and results, and the answer — ready to append to your own history.
  • content holds the final assistant turn as structured content, when you need more than the text.

The same numbers are also recorded on tracing spans — see Observability.

An agent is stateless: each prompt starts from scratch unless you supply the earlier turns. You can:

  • pass history yourself with agent.prompt(p).history(history), and record the new turn from response.messages;
  • use agent.chat(p, &mut history), which sends the history and appends the run’s messages to it on success;
  • attach conversation memory with .memory(backend) on the builder and .conversation(id) on the prompt, so Rig loads and stores each conversation’s history for you.

Memory covers all three, plus durable backends and policies that bound history growth. For a ready-made conversational REPL, see Build a CLI chatbot.

To get a typed value instead of text, use prompt_typed::<T>(). Rig sends T’s JSON schema as the run’s structured-output schema (providers with native structured output constrain the response to it) and deserializes the final answer:

#[derive(Debug, serde::Deserialize, schemars::JsonSchema)]
struct Forecast {
city: String,
temperature_f: f64,
}
let forecast: Forecast = agent
.prompt_typed::<Forecast>("Give me a plausible forecast for New York.")
.retries(1) // re-run once if the output doesn't parse as `Forecast`
.await?
.output;
println!("{} will be {}°F", forecast.city, forecast.temperature_f);

A typed run takes the same per-run settings as prompt. It returns a TypedPromptResponse<T> with the parsed output plus usage, completion_calls, and messages; failures surface as StructuredOutputError. To fix the schema on the agent itself, use AgentBuilder::output_schema::<T>() and choose how it reaches the provider with output_mode(OutputMode::..). For extraction-focused agents, see Extractors.

Agents can add context to every model request automatically:

  • Static context (AgentBuilder::context(text)) is attached to every request: good for a small, fixed set of reference material.
  • Dynamic context (AgentBuilder::dynamic_context(n, index)) retrieves the n most relevant documents from a vector store before each model call and attaches them. This is RAG; see Vector Stores & RAG for building the index and a complete agent.

Agents resolve tool calls, execute the tools, and feed results back to the model automatically, as described in the agent loop. Attach tools on the builder:

  • .tool(t) — a static tool implementing the Tool trait, always offered to the model.
  • .dynamic_tool(t) / .dynamic_tools(ts) — tools defined at runtime as DynamicTool values.
  • .retrieved_tools(n, index, toolset) — tools retrieved from a vector store per prompt, so only the n most relevant are offered.
let agent = AgentBuilder::new(OpenAI::from_env()?.completion(openai::GPT_5_5))
.preamble("You are a calculator. Use the tools to compute results.")
.tool(Add)
.default_max_turns(3) // tool call(s), then the answer
.build();
let response = agent.prompt("What is 2 + 5?").await?;

For how to write tools, pass them runtime context, and set up retrieval, see Tools.

By default the model decides freely whether to call a tool or answer directly. When you need to constrain that, set a ToolChoice on the builder (or per run with .tool_choice(..) on the prompt):

use rig::message::ToolChoice;
let agent = AgentBuilder::new(OpenAI::from_env()?.completion(openai::GPT_5_5))
.preamble("You are a calculator. Always compute with tools, never in your head.")
.tool_choice(ToolChoice::Required) // the model must call a tool before answering
.build();
  • ToolChoice::Auto — the default: the model may call tools or answer directly.
  • ToolChoice::None — tools are visible but must not be called.
  • ToolChoice::Required — the model must call at least one tool.
  • ToolChoice::Specific { function_names } — the model must call one of the named tools.

A ToolChoice set on the agent applies to every turn, so Required keeps forcing a tool call and the model never gets a turn where it can stop and answer. To force only the first step (always search before answering), patch the first turn’s request from a hook — see Request patches.

An agent can be handed to another agent as a tool. This is the building block for the manager-worker pattern: a coordinating agent delegates subtasks to specialised worker agents by calling them like any other tool.

agent.into_tool()? turns an agent into a DynamicTool that takes a prompt string, prompts the agent, and returns its answer. The tool is named after the agent’s name (or agent_tool when it has none), and its description includes the agent’s description and preamble, so give workers both:

let model = OpenAI::from_env()?.completion(openai::GPT_5_5).erase();
// The worker: a specialised sub-agent.
let bob = AgentBuilder::new(model.clone())
.name("bob")
.description("An employee who handles admin tasks at FooBar Inc.")
.preamble("You are Bob, an admin employee. Your manager Alice may ask you to do things.")
.build();
// The manager: prompts Bob by calling him as a tool.
let alice = AgentBuilder::new(model)
.name("alice")
.preamble("You are Alice, a manager in the admin department. You manage Bob.")
.dynamic_tool(bob.into_tool()?)
.build();
let response = alice
.prompt("Ask Bob to draft a welcome email and tell me what he wrote.")
.max_turns(3)
.await?;
println!("{response}");

erase() turns the model into a type-erased DynModel, so one handle can be cloned into several agents. The worker runs with its own turn budget (default_max_turns), and the tool call inherits the manager’s tool context. For swarm-style and actor-based architectures, see Multi-agent systems.

An agent can hold several models and pick one per model call. Register each under a label with AgentBuilder::named_model(label, model) and .model_route(label, model), then select by label:

  • per run, with .using_model("label") on the prompt;
  • per model call, from a hook’s on_model_select, returning ModelSelectionAction::select("label").
use rig::agent::{AgentHook, HookContext, ModelSelection, ModelSelectionAction};
/// A cheap model for the first call; the stronger one once tool results are in.
struct CheapFirst;
impl AgentHook for CheapFirst {
fn on_model_select(&self, ctx: &HookContext, _event: ModelSelection<'_>) -> ModelSelectionAction {
if ctx.turn() == 1 {
ModelSelectionAction::select("fast")
} else {
ModelSelectionAction::select("strong")
}
}
}
let client = OpenAI::from_env()?;
let agent = AgentBuilder::named_model("fast", client.completion(openai::GPT_5_4_MINI))
.model_route("strong", client.completion(openai::GPT_5_5))
.add_hook(CheapFirst)
.build();
// Or skip the hook and pick one model for a whole run:
let response = agent.prompt("Summarize Rust's ownership rules.").using_model("strong").await?;

To change an existing agent’s default model, agent.set_model(model) (or with_model(model) by value) swaps in any model; set_model_label("label") switches to an already registered route. For routing strategies, see Model routing.

Common knobs such as reasoning effort, caching, and seeds have typed setters on the builder (.reasoning(Effort::High), .cache(..), .seed(..), …) that every provider either maps or refuses. For provider parameters Rig doesn’t model, pass raw JSON with AgentBuilder::additional_params(); it’s merged into every completion request:

use rig::completion::Effort;
let agent = AgentBuilder::new(OpenAI::from_env()?.completion(openai::GPT_5_5))
.preamble("You are a helpful agent.")
.reasoning(Effort::High)
.additional_params(json!({ "user": "user-42" }))
.build();

Each prompt is driven by an AgentRunner, which owns everything specific to that run: turn budget, history, memory, tool concurrency and context, per-run request overrides, and hooks. Hooks let your code observe and steer each step of the loop — rewrite the prompt, patch a request, pick a model, retry a response, approve or deny tool calls, rewrite tool results, or stop the run. Install hooks for every run with AgentBuilder::add_hook, or for one run with .add_hook(..) on the prompt.

When a run must outlive the process — waiting hours for a human to approve a tool call, for example — step the serializable AgentRun state machine yourself.

An agent lets the model decide the control flow — the right tool for open-ended tasks where the steps can’t be predicted. When you already know the steps, plain Rust calling agents in sequence is simpler, cheaper, and deterministic: see Agent or workflow? on the Workflows page for how to choose.

  • Completions — the model layer beneath agents
  • AgentRunner — per-run settings and the ways to drive a run
  • Hooks — observe and steer agent runs
  • Durable runs — step, pause, and resume the agent loop yourself
  • Tools — extend an agent with callable functions
  • Memory — conversation history, compaction, and long-term memory
  • Streaming — stream an agent’s responses token by token
  • Error Handling — handle PromptError::MaxTurns and transient failures
  • Build a RAG system — a full retrieval-augmented agent