Step 1
Open the studio
Go to FairStack and open the ai voice workspace.
Text-to-speech and speech-to-speech on eight selectable voice models. Pay per generation with a 20% margin. FairStack never trains on your content.
ElevenLabs TTS Multilingual V2 · Prompt: Product UI recording — FairStack docs voice API walkthrough
FairStack AI voice covers text-to-speech and speech transformation models under the same balance as image and video. You pay the model cost plus 20%.
8 selectable ai voice models · 113 on the platform
Workflows
Every model below is non-hidden, non-fallback, generation-category — part of the 113 customers can actually run.
$0.00002/char
Fish Audio S2.1 Pro is the synthesis engine behind FairStack's customer voice cloning. Unlike ElevenLabs (hard-capped at 10 custom voices per account), Fish cloning is cap-free: every user can create and keep their own persistent voice model with no slot wall. Clone a voice from a short reference sample via POST /v1/voice/clone (engine: fish), then generate speech in that voice by passing the returned voice_id to voice generation with this model. Customer-facing generation runs on the free s2.1-pro-free tier at $0 marginal cost. Voice cloning is consent-gated and discloses that voice data is processed by Fish Audio, which may use it to train their models.
$0.0001/char
Qwen3-TTS is FairStack's default text-to-speech model, offering a unique dual capability: voice design from text descriptions and voice cloning from reference audio. In voice design mode, users describe the desired voice characteristics in natural language, and the model generates matching speech without any reference audio needed. The voice design capability is particularly distinctive. Rather than choosing from a preset library, users can describe a voice as they imagine it, and the model creates speech matching that description. Clone mode requires both reference audio and a reference text transcript for best results. Self-hosting provides cost advantages and privacy. Compared to IndexTTS2 which offers higher cloning fidelity, Qwen3-TTS's voice design mode provides a capability no other model on the platform offers. Against preset-based models like Kokoro with 54 fixed voices, voice design offers unlimited vocal variety. As the default TTS model, it is well-tested across diverse use cases. Best suited for custom voice creation, voice design experimentation, and general TTS generation where the flexibility to either design a voice from description or clone from audio provides maximum creative control. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.0001/char
ElevenLabs Turbo 2.5 is ElevenLabs' fastest text-to-speech model, optimized for sub-second latency in voice generation. The Turbo 2.5 variant is designed for real-time and near-real-time applications where generation speed directly impacts user experience, delivering premium voice quality with the lowest possible delay. At $0.05 per generation on FairStack, compared to $0.30 or more through ElevenLabs direct at comparable usage levels, it offers substantial savings. The model achieves sub-second latency while maintaining the voice quality that ElevenLabs is known for, making it viable for interactive applications, live demos, and conversational AI where users cannot wait for generation. Compared to ElevenLabs Multilingual V2, the Turbo variant trades some audio quality and language breadth for significantly faster generation. Against non-ElevenLabs speed-focused models like MiniMax Turbo, it benefits from ElevenLabs' extensive voice model training. Best suited for real-time voice generation, low-latency TTS applications, and interactive use cases where sub-second generation speed is a requirement. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.0002/char
ElevenLabs Multilingual V2 is ElevenLabs' multi-language text-to-speech model supporting 29 languages with natural voice synthesis. The model produces high-quality speech in each supported language with native-sounding pronunciation, prosody, and intonation, making it the broadest multilingual TTS option available on the platform. At $0.05 per generation on FairStack, compared to $0.30 or more through ElevenLabs direct, it delivers the same premium voice quality at a significant discount. The 29-language coverage includes major world languages and many regional languages, with natural-sounding output in each. Voice quality is consistent across languages, with good handling of language-specific phonemes and rhythm patterns. Compared to single-language TTS models, Multilingual V2 provides access to a wide range of languages from a single model with consistent quality. Against other multilingual options like Chatterbox Multilingual with 23 languages, ElevenLabs offers broader coverage and generally higher audio quality. Best suited for multilingual content creation, international voice production, and multi-language TTS workflows where broad language coverage at premium quality matters. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.0001/char
IndexTTS2 is FairStack's self-hosted voice cloning model delivering the highest fidelity to reference audio available on the platform. The model produces natural-sounding speech that closely matches the tone, cadence, timbre, and unique character of the provided voice sample, making it the best choice when preserving the distinctive qualities of a specific voice is critical. The model works with all 168 voices in FairStack's library and accepts custom reference audio for cloning. Unlike some voice cloning models that require a reference text transcript, IndexTTS2 works from audio alone, simplifying the cloning workflow. Self-hosting ensures low latency when the model is warm and eliminates third-party API costs. Compared to Qwen3-TTS which offers both voice design and cloning, IndexTTS2 focuses exclusively on cloning with superior reference fidelity. Against cloud-based voice cloning services, self-hosting provides cost advantages and privacy benefits. Best suited for voice cloning, audiobook narration in a specific voice, and character voice preservation where maximum fidelity to the reference voice is the priority. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.0060/sec
ElevenLabs Voice Changer is ElevenLabs' real-time voice transformation model that modifies voice characteristics including pitch, tone, and speaking style while preserving speech content. With per-second billing, it processes audio and outputs a transformed version with the desired voice characteristics applied. With per-second billing at $0.005 per second, costs remain proportional to audio duration. The model handles a range of transformations from subtle tone adjustments to significant voice character changes. ElevenLabs' expertise in voice synthesis ensures that transformations sound natural rather than mechanically processed. Compared to speech-to-speech models like Chatterbox S2S at $0.05 per generation, the Voice Changer offers more granular per-second pricing and ElevenLabs' established voice quality. Against manual pitch-shifting and processing in audio editing software, AI-driven voice changing produces more natural results. Best suited for voice transformation, character voice creation, and voice modification workflows where changing voice characteristics while maintaining natural sound quality matters. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.018/sec
ElevenLabs AI Dubbing is ElevenLabs' automated dubbing model that translates and re-voices audio or video content into different languages while preserving the original speaker's voice characteristics, emotional tone, and timing. The model handles the full dubbing pipeline: transcription, translation, and voice synthesis in the target language. With per-second billing at $0.015 per second, it makes professional dubbing accessible at a predictable per-use price of human dubbing services, which typically charge hundreds of dollars per minute. The model preserves the speaker's voice timbre and emotional delivery across languages, producing dubbed content that sounds like the original speaker in a new language. Compared to human dubbing services, AI dubbing is dramatically faster and lower-priced while maintaining reasonable quality for most content types. Against manual translation plus TTS approaches, the integrated pipeline preserves speaker identity and timing automatically. Best suited for multi-language dubbing, content localization, and international distribution workflows where automated dubbing at low cost enables global reach. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.060/req
Chatterbox Speech-to-Speech is Resemble AI's voice transformation model that changes the style, tone, and characteristics of existing speech while preserving the spoken content. Users provide source audio, and the model re-voices it with different characteristics, effectively performing voice conversion without requiring the speaker to re-record. At $0.05 per generation, it provides access to Resemble AI's voice conversion technology. The model handles transformations including tone changes, pitch adjustments, speaking style modifications, and voice character shifts. The content preservation ensures that the words, timing, and intent of the original speech remain intact through the transformation. Compared to re-recording speech with a different speaker or voice actor, speech-to-speech conversion preserves the original delivery's timing and emotional nuance while changing the voice character. Against TTS models that generate speech from text, S2S maintains the natural cadence and performance of the original speaker. Best suited for voice style conversion, speech transformation, and voice modification workflows where preserving original speech content while changing voice characteristics is needed. Available on FairStack at infrastructure cost plus a 20% platform fee.
One balance. Pick a model. Pay the model cost plus 20%.
Step 1
Go to FairStack and open the ai voice workspace.
Step 2
Choose from the selectable list on this page — capability flags decide who appears.
Step 3
Run the job. Credits never expire if you stop mid-project.
Step 4
The same balance covers image, video, voice, and music.
FairStack never trains on your content. Upstream voice providers may train under their own terms — our ToS discloses that separately.
Still have questions? We're here to help.