Model Providers
A model provider is a service (or a local runtime) that serves completion, embedding, or other models: OpenAI, Anthropic, Gemini, Ollama, and so on. Rig gives each one a small client with the same shape, so switching providers is usually a change to the import and the model id.
Vocabulary
Section titled “Vocabulary”- Provider module:
rig::providers::<name>, holding the client, model-id constants, and typed options for that provider. - Configuration: the provider’s settings (API key, base URL, API version), such as
openai::OpenAIConfigoranthropic::AnthropicConfig. A configuration is plain, serializable data; the key is never serialized. - Client: a configuration on an HTTP transport, such as
openai::OpenAIoranthropic::Anthropic. Clients are cheap to clone and build models. - Model: a
Model<W>pairs a wire (what to send and how to read the reply) with the client’s transport. You get one from a client method such asclient.completion(id)orclient.embedding(id, None), and hand it toAgentBuilder::new,EmbeddingsBuilder::new, a vector store, and so on.model.erase()turns it into aDynModelwhen you need one type for several providers.
The flow is always: client → model → agent / extractor / index.
use rig::prelude::*;use rig::providers::openai::{self, OpenAI};
let client = OpenAI::from_env()?;
// A completion model, used to build an agent...let agent = AgentBuilder::new(client.completion(openai::GPT_5_5)) .preamble("You are Gandalf the White, discussing the fate of Middle Earth.") .build();
// ...and an embedding model, for embeddings and vector search.let embedder = client.embedding(openai::TEXT_EMBEDDING_3_SMALL, None);See Providers & Clients for how these types fit into the rest of Rig.
Constructing a client
Section titled “Constructing a client”Every built-in client has the same constructors: from_env()? reads the provider’s environment variables (and returns an EnvError naming a missing one), new(key) takes an explicit key, and a configuration such as OpenAIConfig::new(key).with_base_url(..).client() covers settings beyond the key. .with_http(http) swaps the HTTP transport. OpenAI-compatible vendors (DeepSeek, Groq, Together, OpenRouter, and others) have no client type of their own: their module’s from_env() and new(key) return an OpenAI client configured for that vendor, such as deepseek::from_env()?.
Providers & Clients covers these constructors, custom transports, and choosing a provider at runtime from configuration data (ProviderRef::parse("deepseek:deepseek-v4-flash")).
Supported providers
Section titled “Supported providers”Built into rig
Section titled “Built into rig”These live in rig::providers and need no extra feature.
| Provider | Construct | Notes |
|---|---|---|
| OpenAI | openai::OpenAI::from_env()? (OPENAI_API_KEY) | completion uses the Responses API; also chat, embedding, transcription, image_generation (image feature), audio_generation (audio feature), Responses websocket sessions (websocket feature). |
| Anthropic | anthropic::Anthropic::from_env()? (ANTHROPIC_API_KEY) | Messages API: tools, extended thinking, prompt caching. No embeddings. |
| Azure OpenAI | azure::from_env()? (AZURE_ENDPOINT, AZURE_API_VERSION, AZURE_API_KEY or AZURE_TOKEN) | The model id is your deployment name. |
| ChatGPT (subscription) | chatgpt::from_env()? (CHATGPT_ACCESS_TOKEN) | Responses format; sign in with chatgpt::auth::Authenticator and OpenAI::authenticate. |
| Cohere | cohere::Cohere::from_env()? (COHERE_API_KEY) | Chat through Cohere’s OpenAI-compatible API, or its native API (document grounding, citations) with cohere::ChatRoute. Text and image embeddings. |
| GitHub Copilot | copilot::Copilot::from_env()? (GITHUB_COPILOT_API_KEY or COPILOT_API_KEY) | Takes an exchanged session token; Copilot::authenticate runs the device-code login. Chat, Responses (picked by model), embeddings. |
| DeepSeek | deepseek::from_env()? (DEEPSEEK_API_KEY) | OpenAI-compatible. |
| Doubleword | doubleword::from_env()? (DOUBLEWORD_API_KEY) | OpenAI-compatible chat and embeddings. |
| Google Gemini | gemini::Gemini::from_env()? (GEMINI_API_KEY) | completion (generateContent), interactions (Interactions API), embedding, transcription, image_generation, context caching. |
| Groq | groq::from_env()? (GROQ_API_KEY) | OpenAI-compatible; Groq-only fields in groq::extension. |
| Hugging Face | huggingface::from_env()? (HUGGINGFACE_API_KEY) | Inference router; choose a backend with OpenAIConfig::with_sub_route. |
| Hyperbolic | hyperbolic::from_env()? (HYPERBOLIC_API_KEY) | OpenAI-compatible chat, image generation, speech. |
| llama.cpp | llamacpp::from_env()? or llamacpp::new("") | Local llama-server, default http://localhost:8080/v1. |
| MiniMax | minimax::from_env()? (MINIMAX_API_KEY) | Chat Completions; minimax::anthropic_from_env()? for its Anthropic-format endpoint. |
| Mira | mira::from_env()? (MIRA_API_KEY) | Gateway; model ids come from list_models(). |
| Mistral | mistral::from_env()? (MISTRAL_API_KEY) | OpenAI-compatible chat and embeddings. |
| Moonshot (Kimi) | moonshot::from_env()? (MOONSHOT_API_KEY) | Chat Completions; moonshot::anthropic_from_env()? for the Anthropic-format endpoint. |
| Ollama | ollama::Ollama::new() or Ollama::from_env()? | Local daemon. completion uses /v1, native_completion uses /api/chat; embeddings. |
| OpenRouter | openrouter::from_env()? (OPENROUTER_API_KEY) | Routing preferences, fallbacks and reply cost in openrouter::extension. |
| Perplexity | perplexity::from_env()? (PERPLEXITY_API_KEY) | OpenAI-compatible Sonar models. |
| Together AI | together::from_env()? (TOGETHER_API_KEY) | OpenAI-compatible chat and embeddings. |
| Venice | venice::from_env()? (VENICE_API_KEY) | venice_parameters and reply cost in venice::extension. |
| Voyage AI | voyageai::VoyageAi::from_env()? (VOYAGE_API_KEY) | Embeddings and reranking only. |
| xAI | xai::from_env()? (XAI_API_KEY) | Grok models; completion uses xAI’s Responses endpoint. |
| Xiaomi MiMo | xiaomimimo::from_env()? (XIAOMI_MIMO_API_KEY) | Chat Completions; xiaomimimo::anthropic_from_env()? for the Anthropic-format endpoint. |
| Z.AI | zai::from_env()? (ZAI_API_KEY) | Chat Completions (ZAI_CODING dialect for the coding endpoint); zai::anthropic_from_env()? for the Anthropic-format endpoint. |
For an OpenAI-shaped client, completion(id) builds the vendor’s primary endpoint: the Responses API for OpenAI, xAI and ChatGPT, Chat Completions for everyone else. chat(id) and responses(id) pick one explicitly.
Companion crates
Section titled “Companion crates”These ship as separate crates and are enabled through a feature on rig. Each one appears as a module on the facade:
[dependencies]rig = { version = "0.44.0", features = ["bedrock"] }| Provider | Feature | Construct | Notes |
|---|---|---|---|
| AWS Bedrock | bedrock | rig::bedrock::client::BedrockRuntime::from_env().completion(id) | Converse API with streaming and tools; .embedding(id, None), .image_generation(id). Uses the AWS SDK’s credential chain (AWS_DEFAULT_REGION, keys or a profile). |
| Google Vertex AI | vertexai | rig::vertexai::VertexAi::from_env()?.completion(id) | Gemini on Vertex. Uses Application Default Credentials; call it inside a Tokio runtime. VertexAi::builder() takes explicit credentials or your own PredictionService. |
| Gemini over gRPC | gemini-grpc | rig::gemini_grpc::GeminiGrpc::from_env().await?.completion(id) | Gemini’s gRPC API (GEMINI_API_KEY); completion with streaming, tools, reasoning, image input, and .embedding(id, None). |
| Candle | candle | rig::candle::CandleModel::from_gguf(data)?.completion() | Local CPU inference for validated Llama 3, SmolLM2 and Qwen3 checkpoints, plus YOLOv8 pose estimation. SmolLM2 also runs in the browser through WASM. |
| fastembed | fastembed | rig::fastembed::Fastembed::load(&m)?.embedding(&m, None)? | Local embedding models in-process. Native targets only. |
| TypeSafe Jev | typesafeai | rig::typesafeai::Jev::from_env()?.evaluation() (JEV_TOKEN) | Typed judgments (choices, scores, yes/no) for routing and rubric evaluation, not general chat. |
Provider-specific options
Section titled “Provider-specific options”Portable knobs (reasoning effort, prompt-cache retention, service tier, top_p, stop sequences, …) are fields of rig::completion::GenerationOptions, which every provider maps to its own request fields; an option a model can’t honour fails the request with ProviderError::UnsupportedOption unless you set .on_unsupported(OnUnsupported::Ignore). Fields only one provider has live in its extension module (openai::extension::OpenAiOptions, anthropic::extension::AnthropicOptions, openrouter::extension::OpenRouterOptions, …) and are attached with .provider_option(options). See Generation options and Provider extensions; the OpenAI and Anthropic pages show each provider’s fields.
See also
Section titled “See also”- Providers & Clients: how clients and models fit into Rig’s architecture
- Completions: calling a completion model directly
- Embeddings: using embedding models
- Write your own provider
rig::providerson docs.rs
