AgentRunner
AgentRunner is what agent.prompt(...) returns: one run of the agent loop. It starts from the
agent’s configuration, takes per-run settings, and does nothing until you drive it with .await,
.stream(), or .run_channel().
Three layers, from most to least convenient:
Agentstores reusable configuration: model, preamble, tools, context, memory, and default hooks.AgentRunnerdrives one prompt through the agent loop. It owns per-run settings such as the turn budget, history, memory behavior, tool concurrency, tool context, request overrides, and extra hooks.AgentRunis the serializable, sans-I/O state machine the runner drives. Step it yourself only when a run has to pause and resume across processes — see Durable runs.
Driving a run
Section titled “Driving a run”use futures::StreamExt;use rig::agent::MultiTurnStreamItem;use rig::prelude::*;use rig::providers::openai::{self, OpenAI};
#[tokio::main]async fn main() -> anyhow::Result<()> { let agent = AgentBuilder::new(OpenAI::from_env()?.completion(openai::GPT_5_5)) .preamble("You are a helpful assistant.") .build();
// Await the runner: run the loop to completion and get a `PromptResponse`. let response = agent .prompt("Find the answer and show your work.") .max_turns(5) .await?;
println!("answer: {}", response.output()); println!("model calls: {}", response.completion_calls.len()); println!("tokens: {:?}", response.usage.total_tokens);
// Or stream the same run: deltas and tool activity, then the final response. let mut stream = agent.prompt("Now explain it to a child.").max_turns(5).stream(); while let Some(item) = stream.next().await { if let MultiTurnStreamItem::FinalResponse(done) = item? { println!("\n{}", done.output()); } }
Ok(())}There are three ways to drive a runner. They share run construction, tool execution, memory handling, tracing spans, and hooks, so a run behaves the same whichever you pick; the streaming paths add the streamed delta events.
| Call | Returns | Use it for |
|---|---|---|
.await or .run().await | Result<PromptResponse, PromptError> | The final answer, usage, and transcript. |
.stream() | StreamingResult, a stream of MultiTurnStreamItem | Live text, reasoning, tool calls, and results. See Streaming. |
.run_channel() | a future plus a RunEvents feed | A host with its own executor or tick loop: spawn the future anywhere and poll events from elsewhere. |
agent.prompt_typed::<T>(p) returns the same runner with typed output — see
Structured output. agent.chat(p, &mut history) is
prompt(p).history(..) plus appending the run’s messages to history.
Per-run settings
Section titled “Per-run settings”The loop
Section titled “The loop”| Method | What it changes |
|---|---|
max_turns(n) | Total model-call budget for this run, including the first call and every retry. Past it, the run fails with PromptError::MaxTurns. Defaults to the agent’s default_max_turns (1 unless set). |
history(messages) | Chat history preceding the prompt. Bypasses conversation memory for this run. |
conversation(id) | Conversation id used to load and save the agent’s memory. |
without_memory() | Disable memory loading and saving for this run. |
add_hook(hook) | Add a hook after the agent’s default hooks. See Hooks. |
tool_concurrency(n) | Run up to n tool calls from one model turn at once. Defaults to 1; 0 is treated as 1. |
tool_context(ctx) | Typed, runtime-only values cloned into every tool call. Never sent to the model. |
max_invalid_tool_call_retries(n) | Budget for Retry recovery of invalid tool calls. Retries also count against max_turns. |
max_consecutive_malformed_tool_calls(n) | Fail after n turns in a row with tool arguments that aren’t a JSON object. No limit by default. |
unhandled_invalid_tool_call(policy) | What happens to an unknown or disallowed tool call no hook resolves: Fail (default) or Ignore. |
using_model(label) / using_model_value(model) | Use a registered model route, or any model, for this run. See Choosing models at runtime. |
The request
Section titled “The request”Every request setting on AgentBuilder can be overridden for one run: preamble(..) /
without_preamble(), document(..) / documents(..) (extra context), temperature(..) /
without_temperature(), max_tokens(..) / without_max_tokens(), tool_choice(..) /
without_tool_choice(), merge_additional_params(..) / replace_additional_params(..) /
without_additional_params(), the typed generation options (reasoning, cache, seed, top_p,
stop, service_tier, verbosity, parallel_tool_calls), and provider_option(..).
let response = agent .prompt("Write a haiku about the borrow checker.") .preamble("You are a poet. Answer with the poem only.") .temperature(1.0) .max_tokens(200) .await?;These apply to every model call in the run. To change a single turn — say, only the first — return a request patch from a hook instead.
Turns and invalid tool calls
Section titled “Turns and invalid tool calls”max_turns is the agent loop’s safety bound: the exact number of model calls the run may make. A
prompt that calls a tool and then answers needs at least 2. If the model keeps calling tools past the
budget, the run fails with PromptError::MaxTurns, which carries the history accumulated so far.
A tool call Rig can’t dispatch — an unknown name, or a tool the active ToolChoice doesn’t allow —
fails the run with PromptError::UnknownToolCall unless a hook resolves it, or you set
unhandled_invalid_tool_call(UnhandledInvalidToolCall::Ignore) to drop such calls and carry on.
max_invalid_tool_call_retries only matters when a hook resolves a call with Retry. For the hook
side of recovery, see Hooks.
Conversation memory
Section titled “Conversation memory”If the agent has a memory backend and the run has a conversation id (from .conversation(id) or the
agent’s default), the runner loads the stored history before the first model call and appends the
run’s messages once it finishes:
let response = agent.prompt("What did I ask earlier?").conversation("user-42").await?;
if let Some(append) = response.memory_append() { if !append.is_acknowledged() { eprintln!("answer is fine, but saving the turn failed: {:?}", append.failure()); }}A failed load fails the run with PromptError::Memory before any model call. A failed save does not
invalidate the answer; it is reported on response.memory_append. Passing .history(..) bypasses
memory completely: Rig neither loads from nor saves to the backend. .without_memory() does the same
without providing history. See Memory.
Tool concurrency
Section titled “Tool concurrency”When a model emits several tool calls in one turn, the runner can execute them concurrently:
let response = agent .prompt("Check inventory, shipping, and pricing.") .max_turns(3) .tool_concurrency(3) .await?;The committed history keeps the calls in the order the model emitted them. With
tool_concurrency > 1, side effects such as logs, spans, and hook callbacks may interleave, so make
them safe to run concurrently.
Tool context
Section titled “Tool context”tool_context passes runtime-only values to tools: tenant ids, auth tokens, request metadata, or other
data the model should not see. Each value type declares a stable key; tools read it from the
ToolContext they receive in call.
use rig::tool::ToolContext;
#[derive(serde::Serialize, serde::Deserialize, rig::ContextValue)]#[context(key = "tenant_id")]struct TenantId(String);
let mut context = ToolContext::new();context.insert(TenantId("acme".into()))?;
let response = agent .prompt("Update my account preferences.") .tool_context(context) .await?;Inside a tool, ctx.get::<TenantId>()? or ctx.require::<TenantId>()? reads it back. Don’t put
secrets in the prompt just so a tool can read them. For the tool side, see
Tool context.
Resuming a persisted run
Section titled “Resuming a persisted run”agent.resume(run) returns a runner that continues a serialized AgentRun
instead of starting from a prompt. The run is authoritative for what it persisted — its prompt,
history, and turn budget — while the agent and runner supply the model, tools, hooks, and request
settings. Pending tool calls execute under the agent’s tools, then the loop continues as usual:
let run: rig::AgentRun = serde_json::from_str(&saved)?;let response = agent.resume(run).tool_concurrency(4).await?;A resumed run neither loads nor saves conversation memory; whoever persisted it owns that. See Durable runs for producing a run to persist.
See also
Section titled “See also”- Agents — build an agent and understand the loop the runner drives.
- Hooks — observe and steer runner events.
- Durable runs — step, pause, and resume the loop yourself.
- Tools — define the tools the runner executes.
- Memory — conversation memory used by
conversation(..). - Streaming — consume a streamed run.
