Media (image, audio & transcription)
Besides text completions and embeddings, Rig can generate images, generate speech from text, and transcribe audio. Each is a model you get from a provider client, just like a completion model, and you call it with a request:
| Task | Client method | Request builder | Response field |
|---|---|---|---|
| Image generation | client.image_generation(model) | ImageGenerationRequestBuilder | image: Vec<u8> |
| Audio generation (text-to-speech) | client.audio_generation(model) | AudioGenerationRequestBuilder | audio: Vec<u8> |
| Transcription (speech-to-text) | client.transcription(model) | TranscriptionRequestBuilder | text: String |
Every response also carries usage, the provider name, the model the provider reported, and the raw provider response for anything Rig doesn’t normalize.
Image generation
Section titled “Image generation”Requires the image feature.
use rig::image_generation::ImageGenerationRequestBuilder;use rig::providers::openai::{self, OpenAI};
let model = OpenAI::from_env()?.image_generation(openai::GPT_IMAGE_1);
let response = model .call( ImageGenerationRequestBuilder::new("A futuristic city at sunset") .width(1024) .height(1024) .build(), ) .await?;
std::fs::write("city.png", &response.image)?;response.image holds the decoded image bytes. Provider-specific options (quality, style, background, …) go in .additional_params(json!({ ... })).
The same code works with Gemini’s image models:
use rig::providers::gemini::{self, Gemini};
let model = Gemini::from_env()?.image_generation(gemini::GEMINI_2_5_FLASH_IMAGE);let response = model .call(ImageGenerationRequestBuilder::new("A flat icon of a banana").build()) .await?;Audio generation (text-to-speech)
Section titled “Audio generation (text-to-speech)”Requires the audio feature. The builder takes the text and a voice name:
use rig::audio_generation::AudioGenerationRequestBuilder;use rig::providers::openai::{self, OpenAI};
let model = OpenAI::from_env()?.audio_generation(openai::TTS_1);
let response = model .call(AudioGenerationRequestBuilder::new("Hello, how can I help you today?", "alloy").build()) .await?;
std::fs::write("hello.mp3", &response.audio)?;Use .speed(1.25) to change the speaking rate.
Transcription (speech-to-text)
Section titled “Transcription (speech-to-text)”Build a request from a file (its name tells the provider the audio format) or from bytes:
use rig::providers::openai::{self, OpenAI};use rig::transcription::TranscriptionRequestBuilder;
let model = OpenAI::from_env()?.transcription(openai::WHISPER_1);
let response = model .call( TranscriptionRequestBuilder::from_file("audio.mp3")? .language("en".to_string()) .build(), ) .await?;
println!("Transcription: {}", response.text);Transcription: Hey, just calling to confirm our meeting tomorrow at ten. Talk soon.For audio already in memory, use TranscriptionRequestBuilder::new(bytes).filename(Some("clip.wav".to_string())). You can also pass a prompt to guide spelling of names and jargon, and a temperature.
Provider support
Section titled “Provider support”| Provider | Image generation | Audio generation | Transcription |
|---|---|---|---|
| OpenAI | Yes (gpt-image-*, DALL·E) | Yes (tts-1, tts-1-hd) | Yes (Whisper) |
| Gemini | Yes | No | Yes |
| OpenAI-compatible vendors (Azure, Groq, Hugging Face, Mistral, …) | Where the vendor serves it | Where the vendor serves it | Where the vendor serves it |
OpenAI-compatible vendors share the OpenAI client type, so they expose the same methods; a call fails with a provider error if the vendor has no such endpoint. Check the provider pages for model names.
Media as input
Section titled “Media as input”To send images, audio, video, or documents to a model rather than generate them, add them to a user message with the UserContent constructors. See Sending files to the model.
See also
Section titled “See also”- Completions: text generation
- Providers & Clients: how clients create models
