Agents
An Agent is Rig’s primary building block for working with LLMs. It bundles a completion model together
with a system prompt, optional context documents, and a set of tools, then runs the agent loop for
you: it sends your prompt to the model, executes any tools the model calls, feeds the results back, and
repeats until the model produces an answer. Reach for an agent whenever you want to prompt a model
without writing that loop by hand — from a simple chatbot to a tool-using assistant or a RAG system.
A minimal agent
Section titled “A minimal agent”use rig::prelude::*;use rig::providers::openai::{self, OpenAI};
#[tokio::main]async fn main() -> anyhow::Result<()> { let model = OpenAI::from_env()?.completion(openai::GPT_5_5);
let agent = AgentBuilder::new(model) .preamble("You are a helpful assistant.") .temperature(0.7) .build();
let response = agent.prompt("Hello!").await?; println!("{}", response.output());
Ok(())}Hello! How can I help you today?You build an agent with AgentBuilder::new(model)...build() and run it with agent.prompt(...).await?.
The result is a PromptResponse; output() returns the final answer as a String (the response also
implements Display, so println!("{response}") prints the same text). Everything else on this page
is optional configuration layered on top.
What an agent is made of
Section titled “What an agent is made of”AgentBuilder collects the reusable configuration that every prompt to the agent shares:
- Model and request settings — the completion model, the system prompt (
preamble), sampling parameters such astemperatureandmax_tokens, portable generation options such asreasoning(..), and the default turn budget (default_max_turns). - Context — documents added to every request. Static context is always sent; dynamic context is retrieved from a vector store per prompt.
- Tools — capabilities the model can call. Static tools are always offered; retrieved tools are picked from a vector store per prompt.
- Conversation memory (optional) — a backend that loads prior history before each prompt and saves the new turn after it. See Conversations.
- Default hooks (optional) — code that observes or steers every run. See Hooks.
An Agent is Clone and safe to share across tasks (Send + Sync + 'static), so build it once and
reuse it.
How the agent loop works
Section titled “How the agent loop works”Everything an agent does is one loop. When you prompt it, Rig:
- Builds a completion request from the preamble, static context, any dynamic context retrieved from a vector store, the conversation history, and the definitions of every tool on offer.
- Sends it to the model. One model call is a turn.
- Inspects the response.
- If the model answered with text, the loop ends and that text is your result.
- If the model requested tool calls, Rig executes each one — parsing the JSON arguments into
your
Argstype, running yourcallimplementation, and appending the result to the conversation as a tool-result message. A model may request several tool calls in a single turn; Rig runs them and returns all the results together.
- Repeats from step 2 with the updated history, so the model can use the tool results — until it produces a text answer or the turn budget runs out.
flowchart TD P["agent.prompt(...)"] --> M["Send request to model<br/>(preamble + context + history + tools)"] M --> D{"Model response"} D -- "text" --> F["Return the answer"] D -- "tool calls" --> T["Rig executes the tools,<br/>appends results to history"] T --> B{"Turn budget left?"} B -- "yes" --> M B -- "no" --> E["Err(PromptError::MaxTurns)"]These docs call this the agent loop; you may also see “tool-calling loop” elsewhere — same thing.
The turn budget
Section titled “The turn budget”The turn budget is the agent’s guard against runaway loops: every turn costs latency and tokens, so
the loop refuses to run forever. max_turns is the exact total number of model calls a run may
make, counting the first call, every call after a round of tool results, and every retry.
The default budget is 1: the agent makes one model call. That is enough for plain question and
answer, but a prompt that calls a tool and then answers needs at least 2. When the budget runs out
before the model answers, the prompt fails with PromptError::MaxTurns, which carries the history
accumulated so far so you can inspect what happened.
Set a budget per prompt with .max_turns(n), or a default for every prompt with
AgentBuilder::default_max_turns(n):
let response = tool_agent .prompt("Please calculate 2 + 5, then multiply the result by 3") .max_turns(5) // two rounds of tool calls plus the answer, with room to spare .await?;
println!("{response}");2 + 5 = 7, and 7 × 3 = 21.Rig imposes no upper bound, but pick a number that matches the longest tool chain you actually expect, so a confused model fails fast instead of burning tokens. For handling the error itself, see Error Handling.
When a tool call goes wrong
Section titled “When a tool call goes wrong”Two distinct failure modes, with different behavior:
- Your tool returns an error. The run doesn’t end: Rig sends a failed tool result to the model
and the loop continues, so the model can retry or explain the failure. By default the model only
sees generic feedback (
the tool failed); to give it a useful message, return aToolExecutionErroror overridemap_error. See When a tool fails. - The model emits a call Rig can’t dispatch — a tool name that doesn’t exist, one the active
ToolChoicedoesn’t allow, or arguments that aren’t a JSON object. An unknown or disallowed tool fails the prompt withPromptError::UnknownToolCallby default; malformed arguments are answered with corrective feedback so the model can try again. To recover from unknown names instead — retry with feedback, repair the name, or skip the call — implementon_invalid_tool_callon a hook.
To observe, rewrite, or veto individual tool calls yourself — for logging, approval flows, or guardrails — use Hooks.
Prompting and the response
Section titled “Prompting and the response”agent.prompt(p) doesn’t send anything yet: it returns an AgentRunner
for one run. Chain per-run settings onto it, then pick how to drive it:
.await(or.run().await) runs the loop to completion and returns aPromptResponse..stream()yields text deltas, tool calls, tool results, and the final response as they happen — see Streaming.
The PromptResponse holds more than the text. Use it for budgets, dashboards, per-user accounting, or
to continue the conversation:
let response = agent.prompt("What is 2 + 2?").max_turns(3).await?;
println!("answer: {}", response.output());println!( "tokens: {:?} in / {:?} out across {} model calls", response.usage.input_tokens, response.usage.output_tokens, response.requests(),);println!("messages produced by this run: {}", response.messages.len());answer: 2 + 2 = 4.tokens: Some(28) in / Some(12) out across 1 model callsmessages produced by this run: 2usageaggregates token counts across every turn of the loop. Each counter is anOption:Nonemeans the provider didn’t report it.completion_callsbreaks usage down per model call (the last entry tells you how large the final request’s context was), along with each call’s finish reason and raw provider response.messagesholds the messages this run committed — the prompt, any tool calls and results, and the answer — ready to append to your own history.contentholds the final assistant turn as structured content, when you need more than the text.
The same numbers are also recorded on tracing spans — see Observability.
Conversations
Section titled “Conversations”An agent is stateless: each prompt starts from scratch unless you supply the earlier turns. You can:
- pass history yourself with
agent.prompt(p).history(history), and record the new turn fromresponse.messages; - use
agent.chat(p, &mut history), which sends the history and appends the run’s messages to it on success; - attach conversation memory with
.memory(backend)on the builder and.conversation(id)on the prompt, so Rig loads and stores each conversation’s history for you.
Memory covers all three, plus durable backends and policies that bound history growth. For a ready-made conversational REPL, see Build a CLI chatbot.
Structured output
Section titled “Structured output”To get a typed value instead of text, use prompt_typed::<T>(). Rig sends T’s JSON schema as the
run’s structured-output schema (providers with native structured output constrain the response to it)
and deserializes the final answer:
#[derive(Debug, serde::Deserialize, schemars::JsonSchema)]struct Forecast { city: String, temperature_f: f64,}
let forecast: Forecast = agent .prompt_typed::<Forecast>("Give me a plausible forecast for New York.") .retries(1) // re-run once if the output doesn't parse as `Forecast` .await? .output;
println!("{} will be {}°F", forecast.city, forecast.temperature_f);A typed run takes the same per-run settings as prompt. It returns a TypedPromptResponse<T> with the
parsed output plus usage, completion_calls, and messages; failures surface as
StructuredOutputError. To fix the schema on the agent itself, use AgentBuilder::output_schema::<T>()
and choose how it reaches the provider with output_mode(OutputMode::..). For extraction-focused
agents, see Extractors.
Context
Section titled “Context”Agents can add context to every model request automatically:
- Static context (
AgentBuilder::context(text)) is attached to every request: good for a small, fixed set of reference material. - Dynamic context (
AgentBuilder::dynamic_context(n, index)) retrieves thenmost relevant documents from a vector store before each model call and attaches them. This is RAG; see Vector Stores & RAG for building the index and a complete agent.
Agents resolve tool calls, execute the tools, and feed results back to the model automatically, as described in the agent loop. Attach tools on the builder:
.tool(t)— a static tool implementing theTooltrait, always offered to the model..dynamic_tool(t)/.dynamic_tools(ts)— tools defined at runtime asDynamicToolvalues..retrieved_tools(n, index, toolset)— tools retrieved from a vector store per prompt, so only thenmost relevant are offered.
let agent = AgentBuilder::new(OpenAI::from_env()?.completion(openai::GPT_5_5)) .preamble("You are a calculator. Use the tools to compute results.") .tool(Add) .default_max_turns(3) // tool call(s), then the answer .build();
let response = agent.prompt("What is 2 + 5?").await?;For how to write tools, pass them runtime context, and set up retrieval, see Tools.
Steering tool use with ToolChoice
Section titled “Steering tool use with ToolChoice”By default the model decides freely whether to call a tool or answer directly. When you need to
constrain that, set a ToolChoice on the builder (or per run with .tool_choice(..) on the prompt):
use rig::message::ToolChoice;
let agent = AgentBuilder::new(OpenAI::from_env()?.completion(openai::GPT_5_5)) .preamble("You are a calculator. Always compute with tools, never in your head.") .tool_choice(ToolChoice::Required) // the model must call a tool before answering .build();ToolChoice::Auto— the default: the model may call tools or answer directly.ToolChoice::None— tools are visible but must not be called.ToolChoice::Required— the model must call at least one tool.ToolChoice::Specific { function_names }— the model must call one of the named tools.
A ToolChoice set on the agent applies to every turn, so Required keeps forcing a tool call and the
model never gets a turn where it can stop and answer. To force only the first step (always search before answering), patch the
first turn’s request from a hook — see Request patches.
Agents as tools
Section titled “Agents as tools”An agent can be handed to another agent as a tool. This is the building block for the manager-worker pattern: a coordinating agent delegates subtasks to specialised worker agents by calling them like any other tool.
agent.into_tool()? turns an agent into a DynamicTool that takes a prompt string, prompts the
agent, and returns its answer. The tool is named after the agent’s name (or agent_tool when it has
none), and its description includes the agent’s description and preamble, so give workers both:
let model = OpenAI::from_env()?.completion(openai::GPT_5_5).erase();
// The worker: a specialised sub-agent.let bob = AgentBuilder::new(model.clone()) .name("bob") .description("An employee who handles admin tasks at FooBar Inc.") .preamble("You are Bob, an admin employee. Your manager Alice may ask you to do things.") .build();
// The manager: prompts Bob by calling him as a tool.let alice = AgentBuilder::new(model) .name("alice") .preamble("You are Alice, a manager in the admin department. You manage Bob.") .dynamic_tool(bob.into_tool()?) .build();
let response = alice .prompt("Ask Bob to draft a welcome email and tell me what he wrote.") .max_turns(3) .await?;
println!("{response}");erase() turns the model into a type-erased DynModel, so one handle can be cloned into several
agents. The worker runs with its own turn budget (default_max_turns), and the tool call inherits the
manager’s tool context. For swarm-style and actor-based architectures, see
Multi-agent systems.
Choosing models at runtime
Section titled “Choosing models at runtime”An agent can hold several models and pick one per model call. Register each under a label with
AgentBuilder::named_model(label, model) and .model_route(label, model), then select by label:
- per run, with
.using_model("label")on the prompt; - per model call, from a hook’s
on_model_select, returningModelSelectionAction::select("label").
use rig::agent::{AgentHook, HookContext, ModelSelection, ModelSelectionAction};
/// A cheap model for the first call; the stronger one once tool results are in.struct CheapFirst;
impl AgentHook for CheapFirst { fn on_model_select(&self, ctx: &HookContext, _event: ModelSelection<'_>) -> ModelSelectionAction { if ctx.turn() == 1 { ModelSelectionAction::select("fast") } else { ModelSelectionAction::select("strong") } }}
let client = OpenAI::from_env()?;let agent = AgentBuilder::named_model("fast", client.completion(openai::GPT_5_4_MINI)) .model_route("strong", client.completion(openai::GPT_5_5)) .add_hook(CheapFirst) .build();
// Or skip the hook and pick one model for a whole run:let response = agent.prompt("Summarize Rust's ownership rules.").using_model("strong").await?;To change an existing agent’s default model, agent.set_model(model) (or with_model(model) by
value) swaps in any model; set_model_label("label") switches to an already registered route. For
routing strategies, see Model routing.
Additional parameters
Section titled “Additional parameters”Common knobs such as reasoning effort, caching, and seeds have typed setters on the builder
(.reasoning(Effort::High), .cache(..), .seed(..), …) that every provider either maps or refuses.
For provider parameters Rig doesn’t model, pass raw JSON with AgentBuilder::additional_params(); it’s
merged into every completion request:
use rig::completion::Effort;
let agent = AgentBuilder::new(OpenAI::from_env()?.completion(openai::GPT_5_5)) .preamble("You are a helpful agent.") .reasoning(Effort::High) .additional_params(json!({ "user": "user-42" })) .build();Runs and hooks
Section titled “Runs and hooks”Each prompt is driven by an AgentRunner, which owns everything specific
to that run: turn budget, history, memory, tool concurrency and context, per-run request overrides, and
hooks. Hooks let your code observe and steer each step of the loop — rewrite the
prompt, patch a request, pick a model, retry a response, approve or deny tool calls, rewrite tool
results, or stop the run. Install hooks for every run with AgentBuilder::add_hook, or for one run with
.add_hook(..) on the prompt.
When a run must outlive the process — waiting hours for a human to approve a tool call, for example —
step the serializable AgentRun state machine yourself.
Agent or hand-written workflow?
Section titled “Agent or hand-written workflow?”An agent lets the model decide the control flow — the right tool for open-ended tasks where the steps can’t be predicted. When you already know the steps, plain Rust calling agents in sequence is simpler, cheaper, and deterministic: see Agent or workflow? on the Workflows page for how to choose.
See also
Section titled “See also”- Completions — the model layer beneath agents
- AgentRunner — per-run settings and the ways to drive a run
- Hooks — observe and steer agent runs
- Durable runs — step, pause, and resume the agent loop yourself
- Tools — extend an agent with callable functions
- Memory — conversation history, compaction, and long-term memory
- Streaming — stream an agent’s responses token by token
- Error Handling — handle
PromptError::MaxTurnsand transient failures - Build a RAG system — a full retrieval-augmented agent
