Skip to content

Model Providers

A model provider is a service (or a local runtime) that serves completion, embedding, or other models: OpenAI, Anthropic, Gemini, Ollama, and so on. Rig gives each one a small client with the same shape, so switching providers is usually a change to the import and the model id.

  • Provider module: rig::providers::<name>, holding the client, model-id constants, and typed options for that provider.
  • Configuration: the provider’s settings (API key, base URL, API version), such as openai::OpenAIConfig or anthropic::AnthropicConfig. A configuration is plain, serializable data; the key is never serialized.
  • Client: a configuration on an HTTP transport, such as openai::OpenAI or anthropic::Anthropic. Clients are cheap to clone and build models.
  • Model: a Model<W> pairs a wire (what to send and how to read the reply) with the client’s transport. You get one from a client method such as client.completion(id) or client.embedding(id, None), and hand it to AgentBuilder::new, EmbeddingsBuilder::new, a vector store, and so on. model.erase() turns it into a DynModel when you need one type for several providers.

The flow is always: client → model → agent / extractor / index.

use rig::prelude::*;
use rig::providers::openai::{self, OpenAI};
let client = OpenAI::from_env()?;
// A completion model, used to build an agent...
let agent = AgentBuilder::new(client.completion(openai::GPT_5_5))
.preamble("You are Gandalf the White, discussing the fate of Middle Earth.")
.build();
// ...and an embedding model, for embeddings and vector search.
let embedder = client.embedding(openai::TEXT_EMBEDDING_3_SMALL, None);

See Providers & Clients for how these types fit into the rest of Rig.

Every built-in client has the same constructors: from_env()? reads the provider’s environment variables (and returns an EnvError naming a missing one), new(key) takes an explicit key, and a configuration such as OpenAIConfig::new(key).with_base_url(..).client() covers settings beyond the key. .with_http(http) swaps the HTTP transport. OpenAI-compatible vendors (DeepSeek, Groq, Together, OpenRouter, and others) have no client type of their own: their module’s from_env() and new(key) return an OpenAI client configured for that vendor, such as deepseek::from_env()?.

Providers & Clients covers these constructors, custom transports, and choosing a provider at runtime from configuration data (ProviderRef::parse("deepseek:deepseek-v4-flash")).

These live in rig::providers and need no extra feature.

ProviderConstructNotes
OpenAIopenai::OpenAI::from_env()? (OPENAI_API_KEY)completion uses the Responses API; also chat, embedding, transcription, image_generation (image feature), audio_generation (audio feature), Responses websocket sessions (websocket feature).
Anthropicanthropic::Anthropic::from_env()? (ANTHROPIC_API_KEY)Messages API: tools, extended thinking, prompt caching. No embeddings.
Azure OpenAIazure::from_env()? (AZURE_ENDPOINT, AZURE_API_VERSION, AZURE_API_KEY or AZURE_TOKEN)The model id is your deployment name.
ChatGPT (subscription)chatgpt::from_env()? (CHATGPT_ACCESS_TOKEN)Responses format; sign in with chatgpt::auth::Authenticator and OpenAI::authenticate.
Coherecohere::Cohere::from_env()? (COHERE_API_KEY)Chat through Cohere’s OpenAI-compatible API, or its native API (document grounding, citations) with cohere::ChatRoute. Text and image embeddings.
GitHub Copilotcopilot::Copilot::from_env()? (GITHUB_COPILOT_API_KEY or COPILOT_API_KEY)Takes an exchanged session token; Copilot::authenticate runs the device-code login. Chat, Responses (picked by model), embeddings.
DeepSeekdeepseek::from_env()? (DEEPSEEK_API_KEY)OpenAI-compatible.
Doubleworddoubleword::from_env()? (DOUBLEWORD_API_KEY)OpenAI-compatible chat and embeddings.
Google Geminigemini::Gemini::from_env()? (GEMINI_API_KEY)completion (generateContent), interactions (Interactions API), embedding, transcription, image_generation, context caching.
Groqgroq::from_env()? (GROQ_API_KEY)OpenAI-compatible; Groq-only fields in groq::extension.
Hugging Facehuggingface::from_env()? (HUGGINGFACE_API_KEY)Inference router; choose a backend with OpenAIConfig::with_sub_route.
Hyperbolichyperbolic::from_env()? (HYPERBOLIC_API_KEY)OpenAI-compatible chat, image generation, speech.
llama.cppllamacpp::from_env()? or llamacpp::new("")Local llama-server, default http://localhost:8080/v1.
MiniMaxminimax::from_env()? (MINIMAX_API_KEY)Chat Completions; minimax::anthropic_from_env()? for its Anthropic-format endpoint.
Miramira::from_env()? (MIRA_API_KEY)Gateway; model ids come from list_models().
Mistralmistral::from_env()? (MISTRAL_API_KEY)OpenAI-compatible chat and embeddings.
Moonshot (Kimi)moonshot::from_env()? (MOONSHOT_API_KEY)Chat Completions; moonshot::anthropic_from_env()? for the Anthropic-format endpoint.
Ollamaollama::Ollama::new() or Ollama::from_env()?Local daemon. completion uses /v1, native_completion uses /api/chat; embeddings.
OpenRouteropenrouter::from_env()? (OPENROUTER_API_KEY)Routing preferences, fallbacks and reply cost in openrouter::extension.
Perplexityperplexity::from_env()? (PERPLEXITY_API_KEY)OpenAI-compatible Sonar models.
Together AItogether::from_env()? (TOGETHER_API_KEY)OpenAI-compatible chat and embeddings.
Venicevenice::from_env()? (VENICE_API_KEY)venice_parameters and reply cost in venice::extension.
Voyage AIvoyageai::VoyageAi::from_env()? (VOYAGE_API_KEY)Embeddings and reranking only.
xAIxai::from_env()? (XAI_API_KEY)Grok models; completion uses xAI’s Responses endpoint.
Xiaomi MiMoxiaomimimo::from_env()? (XIAOMI_MIMO_API_KEY)Chat Completions; xiaomimimo::anthropic_from_env()? for the Anthropic-format endpoint.
Z.AIzai::from_env()? (ZAI_API_KEY)Chat Completions (ZAI_CODING dialect for the coding endpoint); zai::anthropic_from_env()? for the Anthropic-format endpoint.

For an OpenAI-shaped client, completion(id) builds the vendor’s primary endpoint: the Responses API for OpenAI, xAI and ChatGPT, Chat Completions for everyone else. chat(id) and responses(id) pick one explicitly.

These ship as separate crates and are enabled through a feature on rig. Each one appears as a module on the facade:

[dependencies]
rig = { version = "0.44.0", features = ["bedrock"] }
ProviderFeatureConstructNotes
AWS Bedrockbedrockrig::bedrock::client::BedrockRuntime::from_env().completion(id)Converse API with streaming and tools; .embedding(id, None), .image_generation(id). Uses the AWS SDK’s credential chain (AWS_DEFAULT_REGION, keys or a profile).
Google Vertex AIvertexairig::vertexai::VertexAi::from_env()?.completion(id)Gemini on Vertex. Uses Application Default Credentials; call it inside a Tokio runtime. VertexAi::builder() takes explicit credentials or your own PredictionService.
Gemini over gRPCgemini-grpcrig::gemini_grpc::GeminiGrpc::from_env().await?.completion(id)Gemini’s gRPC API (GEMINI_API_KEY); completion with streaming, tools, reasoning, image input, and .embedding(id, None).
Candlecandlerig::candle::CandleModel::from_gguf(data)?.completion()Local CPU inference for validated Llama 3, SmolLM2 and Qwen3 checkpoints, plus YOLOv8 pose estimation. SmolLM2 also runs in the browser through WASM.
fastembedfastembedrig::fastembed::Fastembed::load(&m)?.embedding(&m, None)?Local embedding models in-process. Native targets only.
TypeSafe Jevtypesafeairig::typesafeai::Jev::from_env()?.evaluation() (JEV_TOKEN)Typed judgments (choices, scores, yes/no) for routing and rubric evaluation, not general chat.

Portable knobs (reasoning effort, prompt-cache retention, service tier, top_p, stop sequences, …) are fields of rig::completion::GenerationOptions, which every provider maps to its own request fields; an option a model can’t honour fails the request with ProviderError::UnsupportedOption unless you set .on_unsupported(OnUnsupported::Ignore). Fields only one provider has live in its extension module (openai::extension::OpenAiOptions, anthropic::extension::AnthropicOptions, openrouter::extension::OpenRouterOptions, …) and are attached with .provider_option(options). See Generation options and Provider extensions; the OpenAI and Anthropic pages show each provider’s fields.