Skip to content

AgentRunner

AgentRunner is what agent.prompt(...) returns: one run of the agent loop. It starts from the agent’s configuration, takes per-run settings, and does nothing until you drive it with .await, .stream(), or .run_channel().

Three layers, from most to least convenient:

  • Agent stores reusable configuration: model, preamble, tools, context, memory, and default hooks.
  • AgentRunner drives one prompt through the agent loop. It owns per-run settings such as the turn budget, history, memory behavior, tool concurrency, tool context, request overrides, and extra hooks.
  • AgentRun is the serializable, sans-I/O state machine the runner drives. Step it yourself only when a run has to pause and resume across processes — see Durable runs.
use futures::StreamExt;
use rig::agent::MultiTurnStreamItem;
use rig::prelude::*;
use rig::providers::openai::{self, OpenAI};
#[tokio::main]
async fn main() -> anyhow::Result<()> {
let agent = AgentBuilder::new(OpenAI::from_env()?.completion(openai::GPT_5_5))
.preamble("You are a helpful assistant.")
.build();
// Await the runner: run the loop to completion and get a `PromptResponse`.
let response = agent
.prompt("Find the answer and show your work.")
.max_turns(5)
.await?;
println!("answer: {}", response.output());
println!("model calls: {}", response.completion_calls.len());
println!("tokens: {:?}", response.usage.total_tokens);
// Or stream the same run: deltas and tool activity, then the final response.
let mut stream = agent.prompt("Now explain it to a child.").max_turns(5).stream();
while let Some(item) = stream.next().await {
if let MultiTurnStreamItem::FinalResponse(done) = item? {
println!("\n{}", done.output());
}
}
Ok(())
}

There are three ways to drive a runner. They share run construction, tool execution, memory handling, tracing spans, and hooks, so a run behaves the same whichever you pick; the streaming paths add the streamed delta events.

CallReturnsUse it for
.await or .run().awaitResult<PromptResponse, PromptError>The final answer, usage, and transcript.
.stream()StreamingResult, a stream of MultiTurnStreamItemLive text, reasoning, tool calls, and results. See Streaming.
.run_channel()a future plus a RunEvents feedA host with its own executor or tick loop: spawn the future anywhere and poll events from elsewhere.

agent.prompt_typed::<T>(p) returns the same runner with typed output — see Structured output. agent.chat(p, &mut history) is prompt(p).history(..) plus appending the run’s messages to history.

MethodWhat it changes
max_turns(n)Total model-call budget for this run, including the first call and every retry. Past it, the run fails with PromptError::MaxTurns. Defaults to the agent’s default_max_turns (1 unless set).
history(messages)Chat history preceding the prompt. Bypasses conversation memory for this run.
conversation(id)Conversation id used to load and save the agent’s memory.
without_memory()Disable memory loading and saving for this run.
add_hook(hook)Add a hook after the agent’s default hooks. See Hooks.
tool_concurrency(n)Run up to n tool calls from one model turn at once. Defaults to 1; 0 is treated as 1.
tool_context(ctx)Typed, runtime-only values cloned into every tool call. Never sent to the model.
max_invalid_tool_call_retries(n)Budget for Retry recovery of invalid tool calls. Retries also count against max_turns.
max_consecutive_malformed_tool_calls(n)Fail after n turns in a row with tool arguments that aren’t a JSON object. No limit by default.
unhandled_invalid_tool_call(policy)What happens to an unknown or disallowed tool call no hook resolves: Fail (default) or Ignore.
using_model(label) / using_model_value(model)Use a registered model route, or any model, for this run. See Choosing models at runtime.

Every request setting on AgentBuilder can be overridden for one run: preamble(..) / without_preamble(), document(..) / documents(..) (extra context), temperature(..) / without_temperature(), max_tokens(..) / without_max_tokens(), tool_choice(..) / without_tool_choice(), merge_additional_params(..) / replace_additional_params(..) / without_additional_params(), the typed generation options (reasoning, cache, seed, top_p, stop, service_tier, verbosity, parallel_tool_calls), and provider_option(..).

let response = agent
.prompt("Write a haiku about the borrow checker.")
.preamble("You are a poet. Answer with the poem only.")
.temperature(1.0)
.max_tokens(200)
.await?;

These apply to every model call in the run. To change a single turn — say, only the first — return a request patch from a hook instead.

max_turns is the agent loop’s safety bound: the exact number of model calls the run may make. A prompt that calls a tool and then answers needs at least 2. If the model keeps calling tools past the budget, the run fails with PromptError::MaxTurns, which carries the history accumulated so far.

A tool call Rig can’t dispatch — an unknown name, or a tool the active ToolChoice doesn’t allow — fails the run with PromptError::UnknownToolCall unless a hook resolves it, or you set unhandled_invalid_tool_call(UnhandledInvalidToolCall::Ignore) to drop such calls and carry on. max_invalid_tool_call_retries only matters when a hook resolves a call with Retry. For the hook side of recovery, see Hooks.

If the agent has a memory backend and the run has a conversation id (from .conversation(id) or the agent’s default), the runner loads the stored history before the first model call and appends the run’s messages once it finishes:

let response = agent.prompt("What did I ask earlier?").conversation("user-42").await?;
if let Some(append) = response.memory_append() {
if !append.is_acknowledged() {
eprintln!("answer is fine, but saving the turn failed: {:?}", append.failure());
}
}

A failed load fails the run with PromptError::Memory before any model call. A failed save does not invalidate the answer; it is reported on response.memory_append. Passing .history(..) bypasses memory completely: Rig neither loads from nor saves to the backend. .without_memory() does the same without providing history. See Memory.

When a model emits several tool calls in one turn, the runner can execute them concurrently:

let response = agent
.prompt("Check inventory, shipping, and pricing.")
.max_turns(3)
.tool_concurrency(3)
.await?;

The committed history keeps the calls in the order the model emitted them. With tool_concurrency > 1, side effects such as logs, spans, and hook callbacks may interleave, so make them safe to run concurrently.

tool_context passes runtime-only values to tools: tenant ids, auth tokens, request metadata, or other data the model should not see. Each value type declares a stable key; tools read it from the ToolContext they receive in call.

use rig::tool::ToolContext;
#[derive(serde::Serialize, serde::Deserialize, rig::ContextValue)]
#[context(key = "tenant_id")]
struct TenantId(String);
let mut context = ToolContext::new();
context.insert(TenantId("acme".into()))?;
let response = agent
.prompt("Update my account preferences.")
.tool_context(context)
.await?;

Inside a tool, ctx.get::<TenantId>()? or ctx.require::<TenantId>()? reads it back. Don’t put secrets in the prompt just so a tool can read them. For the tool side, see Tool context.

agent.resume(run) returns a runner that continues a serialized AgentRun instead of starting from a prompt. The run is authoritative for what it persisted — its prompt, history, and turn budget — while the agent and runner supply the model, tools, hooks, and request settings. Pending tool calls execute under the agent’s tools, then the loop continues as usual:

let run: rig::AgentRun = serde_json::from_str(&saved)?;
let response = agent.resume(run).tool_concurrency(4).await?;

A resumed run neither loads nor saves conversation memory; whoever persisted it owns that. See Durable runs for producing a run to persist.

  • Agents — build an agent and understand the loop the runner drives.
  • Hooks — observe and steer runner events.
  • Durable runs — step, pause, and resume the loop yourself.
  • Tools — define the tools the runner executes.
  • Memory — conversation memory used by conversation(..).
  • Streaming — consume a streamed run.