Skip to content

Media (image, audio & transcription)

Besides text completions and embeddings, Rig can generate images, generate speech from text, and transcribe audio. Each is a model you get from a provider client, just like a completion model, and you call it with a request:

TaskClient methodRequest builderResponse field
Image generationclient.image_generation(model)ImageGenerationRequestBuilderimage: Vec<u8>
Audio generation (text-to-speech)client.audio_generation(model)AudioGenerationRequestBuilderaudio: Vec<u8>
Transcription (speech-to-text)client.transcription(model)TranscriptionRequestBuildertext: String

Every response also carries usage, the provider name, the model the provider reported, and the raw provider response for anything Rig doesn’t normalize.

Requires the image feature.

use rig::image_generation::ImageGenerationRequestBuilder;
use rig::providers::openai::{self, OpenAI};
let model = OpenAI::from_env()?.image_generation(openai::GPT_IMAGE_1);
let response = model
.call(
ImageGenerationRequestBuilder::new("A futuristic city at sunset")
.width(1024)
.height(1024)
.build(),
)
.await?;
std::fs::write("city.png", &response.image)?;

response.image holds the decoded image bytes. Provider-specific options (quality, style, background, …) go in .additional_params(json!({ ... })).

The same code works with Gemini’s image models:

use rig::providers::gemini::{self, Gemini};
let model = Gemini::from_env()?.image_generation(gemini::GEMINI_2_5_FLASH_IMAGE);
let response = model
.call(ImageGenerationRequestBuilder::new("A flat icon of a banana").build())
.await?;

Requires the audio feature. The builder takes the text and a voice name:

use rig::audio_generation::AudioGenerationRequestBuilder;
use rig::providers::openai::{self, OpenAI};
let model = OpenAI::from_env()?.audio_generation(openai::TTS_1);
let response = model
.call(AudioGenerationRequestBuilder::new("Hello, how can I help you today?", "alloy").build())
.await?;
std::fs::write("hello.mp3", &response.audio)?;

Use .speed(1.25) to change the speaking rate.

Build a request from a file (its name tells the provider the audio format) or from bytes:

use rig::providers::openai::{self, OpenAI};
use rig::transcription::TranscriptionRequestBuilder;
let model = OpenAI::from_env()?.transcription(openai::WHISPER_1);
let response = model
.call(
TranscriptionRequestBuilder::from_file("audio.mp3")?
.language("en".to_string())
.build(),
)
.await?;
println!("Transcription: {}", response.text);
Transcription: Hey, just calling to confirm our meeting tomorrow at ten. Talk soon.

For audio already in memory, use TranscriptionRequestBuilder::new(bytes).filename(Some("clip.wav".to_string())). You can also pass a prompt to guide spelling of names and jargon, and a temperature.

ProviderImage generationAudio generationTranscription
OpenAIYes (gpt-image-*, DALL·E)Yes (tts-1, tts-1-hd)Yes (Whisper)
GeminiYesNoYes
OpenAI-compatible vendors (Azure, Groq, Hugging Face, Mistral, …)Where the vendor serves itWhere the vendor serves itWhere the vendor serves it

OpenAI-compatible vendors share the OpenAI client type, so they expose the same methods; a call fails with a provider error if the vendor has no such endpoint. Check the provider pages for model names.

To send images, audio, video, or documents to a model rather than generate them, add them to a user message with the UserContent constructors. See Sending files to the model.