Skip to content

Completions

Completions are the layer beneath agents: one request to one model, and the normalized response that comes back. There is no tool loop, no memory, and no hooks at this level. Use it when you want full control over a single call, or to write your own loop.

A completion model is a Model built by a client. Its call method sends a request and returns a CompletionResponse:

use rig::completion::CompletionRequest;
use rig::providers::openai::{self, OpenAI};
let model = OpenAI::from_env()?.completion(openai::GPT_5_5);
let request = CompletionRequest::new("What is the Rust programming language?")
.preamble("You are a helpful assistant.")
.temperature(0.7)
.max_tokens(1000);
let response = model.call(request).await?;
println!("{}", response.text());

call also takes a plain prompt (model.call("Hello").await?) or a Vec<Message> conversation. Models are cheap handles over a shared client: build them once and reuse them.

CompletionRequest::new(prompt) starts a conversation whose last message is the prompt. Its setters add the rest:

SetterSets
.preamble(text)a leading system message
.messages(history) / .message(m)earlier turns, inserted before the prompt
.documents(docs) / .document(d)context documents
.tools(defs) / .tool(def)tool definitions the model may call
.tool_choice(choice)whether the model must, may, or must not call tools
.output_schema(schema)a JSON schema the answer must follow
.temperature(t) / .max_tokens(n)sampling temperature, output-token cap
.model(id)a different model id for this request only
.additional_params(json)raw JSON merged into the provider’s request body

Carrying a conversation forward is a matter of appending each response to the history:

let mut history = vec![
Message::user("My name is Ferris."),
Message::assistant("Nice to meet you, Ferris!"),
];
let prompt = Message::user("What's my name?");
let response = model
.call(CompletionRequest::new(prompt.clone()).messages(history.clone()))
.await?;
history.push(prompt);
history.extend(response.message());

A request with tool definitions can come back with tool calls instead of text. At this level you run them yourself and send the results back; the Tools page shows that loop. An agent does it for you.

Settings that every provider spells differently (reasoning effort, prompt caching, service tier) are typed and portable. Set them on the request, and each provider maps them to its own JSON:

use rig::completion::{CacheRetention, CompletionRequest, Effort, ServiceTier};
let request = CompletionRequest::new("Plan the refactor of the parser module.")
.reasoning(Effort::High)
.cache(CacheRetention::Long)
.service_tier(ServiceTier::Flex)
.seed(7);
let response = model.call(request).await?;
println!("{}", response.reasoning());
OptionValues
reasoningEffort::{Minimal, Low, Medium, High, XHigh, Max}, Reasoning::Budget { tokens }, or Reasoning::Off
cacheCacheRetention::{None, Short, Long}: how long the provider keeps the prompt prefix cached
service_tierServiceTier::{Auto, Default, Flex, Priority}
verbosityVerbosity::{Low, Medium, High}
parallel_tool_calls, top_p, seed, stopthe usual sampling and tool settings

A provider never drops an option silently. If the provider or model can’t honour one, the call fails with ProviderError::UnsupportedOption before anything is sent. Add .on_unsupported(OnUnsupported::Ignore) to skip unsupported options with a warning instead. A request that sets no options is sent as built.

GenerationOptions holds the same settings as one reusable value: pass it to a request with .options(..). AgentBuilder has the same setters, so an agent can apply them to every call it makes.

For settings only one provider has, each provider module has an extension module with a typed Options type for requests and an Extras type for the fields of a reply that Rig doesn’t normalize:

use rig::providers::anthropic::{self, Anthropic};
use rig::providers::anthropic::extension::{AnthropicExt, AnthropicOptions};
let model = Anthropic::from_env()?.completion(anthropic::CLAUDE_SONNET_5_5);
let request = CompletionRequest::new("Name three Rust web frameworks.")
.provider_option(AnthropicOptions::default().top_k(40));
let response = model.call(request).await?;
let extras = response.extras_lossy::<AnthropicExt>();
println!("stop reason: {:?}, tier: {:?}", extras.stop_reason, extras.service_tier);

Each provider only reads its own options, so one request can carry options for several providers and be sent to whichever model you pick. additional_params still accepts raw JSON for anything that has no typed field. The request body is built in this order, later layers winning: the provider’s encoding, generation options, provider options, then additional_params.

CompletionResponse is the same shape for every provider:

Field or methodWhat it holds
choicea Vec<AssistantContent>, one block per output item, in provider order (possibly empty)
text() / reasoning()the text blocks, or the reasoning blocks, joined
tool_calls()the tool calls the model asked for
message()the response as an assistant Message, ready to append to history
usagetoken counts (see below)
finish_reason()why the model stopped
provider(), model(), response_id()where the reply came from
rawthe provider’s response document as JSON

AssistantContent is Text, ToolCall, Reasoning, Image, or Opaque (a provider item Rig keeps so it can be replayed). The enum is #[non_exhaustive], so a match needs a _ arm.

finish_reason() returns a normalized FinishReason: Stop, Length, ToolCalls, ContentFilter, or Other(raw) for a provider value outside that set.

let response = model.call(CompletionRequest::new("Write a long poem.").max_tokens(50)).await?;
match response.finish_reason() {
Some(FinishReason::Length) => println!("cut off: raise max_tokens"),
Some(FinishReason::ContentFilter) => println!("filtered by the provider"),
Some(FinishReason::Other(raw)) => println!("provider said: {raw}"),
_ => println!("{}", response.text()),
}

An agent treats an unknown (Other) reason, filtered content, or a provider-reported failure as a failed turn rather than a normal answer. If a provider you use sends reasons outside the normalized set that you want to accept, set .accept_unknown_finish_reasons(true) on the request or on the agent.

Usage has a field per counter, each an Option<u64> so “not reported” is distinct from zero: input_tokens, output_tokens, total_tokens, cached_input_tokens, cache_creation_input_tokens, tool_use_prompt_tokens, and reasoning_tokens. cost holds a Cost when the provider reports one.

let response = model.call(CompletionRequest::new("Hi")).await?;
let usage = &response.usage;
println!(
"in: {}, out: {}, cached: {:?}",
usage.input_tokens.unwrap_or(0),
usage.output_tokens.unwrap_or(0),
usage.cached_input_tokens,
);

An agent run adds up the usage of every model call it makes; read it from the usage field of the run’s PromptResponse. The catalog can price a Usage for a known model.

model.stream(request) returns the reply as it arrives. Each item is an event (text, reasoning, or tool-call fragments), and finish() folds whatever is left into the same CompletionResponse that call returns:

use futures::StreamExt;
use rig::completion::CompletionRequest;
use rig::streaming::{Item, StreamEvent};
let mut stream = model.stream(CompletionRequest::new("Tell me a short story."))?;
while let Some(item) = stream.next().await {
if let Item::Event(StreamEvent::Text { text, .. }) = item? {
print!("{text}");
}
}

Nothing is sent until the stream is first polled, and dropping the stream cancels the request. See Streaming for streaming an agent run.

Every model built by a client is a Model, typed by its provider’s wire. To write code that takes a model of any provider, accept a DynModel<Completion>; any completion model converts into one with .into():

use rig::completion::{CompletionRequest, CompletionResponse};
use rig::operation::Completion;
async fn summarize(model: &DynModel<Completion>, text: &str) -> Result<String, ProviderError> {
let request = CompletionRequest::new(format!("Summarize:\n\n{text}"));
let response: CompletionResponse = model.call(request).await?;
Ok(response.text())
}

AgentBuilder::new takes any model that converts into a DynModel<Completion>, which is how the same agent code works across providers.

A conversation is a Vec<Message>:

  • Message::System { content }: a system instruction. Message::system(text) builds one.
  • Message::User { content }: a list of UserContent blocks, which may be text, images, audio, video, documents, or tool results. Message::user(text) builds a text message.
  • Message::Assistant(AssistantMessage): the model’s blocks, plus where they came from and how the turn ended. Message::assistant(text) builds one by hand.

A failed call returns ProviderError. Its variants separate transport failures (Http), malformed replies (Json, Response), errors the provider reported (ProviderResponse, InvalidAuthentication), and refused options (UnsupportedOption). is_retryable() tells transient failures from permanent ones. See Error Handling.