Step 1
Record a clean sample
Quiet room, natural pace, 30–60 seconds of speech.
Clone a voice once, then generate TTS on FairStack. Cloning is a platform capability paired with selectable TTS models — not a fake model slug.
Fish Audio S2.1 Pro · Prompt: FairStack landing — voice product entry
Voice cloning captures a speaker once so TTS models can speak as that voice later. FairStack exposes cloning in the voice product; the models below are the selectable TTS engines you generate with afterward.
Clone a voice →5 selectable models — filtered from config capability flags, not a hand-maintained list.
$0.00002/char
Fish Audio S2.1 Pro is the synthesis engine behind FairStack's customer voice cloning. Unlike ElevenLabs (hard-capped at 10 custom voices per account), Fish cloning is cap-free: every user can create and keep their own persistent voice model with no slot wall. Clone a voice from a short reference sample via POST /v1/voice/clone (engine: fish), then generate speech in that voice by passing the returned voice_id to voice generation with this model. Customer-facing generation runs on the free s2.1-pro-free tier at $0 marginal cost. Voice cloning is consent-gated and discloses that voice data is processed by Fish Audio, which may use it to train their models.
$0.0001/char
Qwen3-TTS is FairStack's default text-to-speech model, offering a unique dual capability: voice design from text descriptions and voice cloning from reference audio. In voice design mode, users describe the desired voice characteristics in natural language, and the model generates matching speech without any reference audio needed. The voice design capability is particularly distinctive. Rather than choosing from a preset library, users can describe a voice as they imagine it, and the model creates speech matching that description. Clone mode requires both reference audio and a reference text transcript for best results. Self-hosting provides cost advantages and privacy. Compared to IndexTTS2 which offers higher cloning fidelity, Qwen3-TTS's voice design mode provides a capability no other model on the platform offers. Against preset-based models like Kokoro with 54 fixed voices, voice design offers unlimited vocal variety. As the default TTS model, it is well-tested across diverse use cases. Best suited for custom voice creation, voice design experimentation, and general TTS generation where the flexibility to either design a voice from description or clone from audio provides maximum creative control. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.0001/char
ElevenLabs Turbo 2.5 is ElevenLabs' fastest text-to-speech model, optimized for sub-second latency in voice generation. The Turbo 2.5 variant is designed for real-time and near-real-time applications where generation speed directly impacts user experience, delivering premium voice quality with the lowest possible delay. At $0.05 per generation on FairStack, compared to $0.30 or more through ElevenLabs direct at comparable usage levels, it offers substantial savings. The model achieves sub-second latency while maintaining the voice quality that ElevenLabs is known for, making it viable for interactive applications, live demos, and conversational AI where users cannot wait for generation. Compared to ElevenLabs Multilingual V2, the Turbo variant trades some audio quality and language breadth for significantly faster generation. Against non-ElevenLabs speed-focused models like MiniMax Turbo, it benefits from ElevenLabs' extensive voice model training. Best suited for real-time voice generation, low-latency TTS applications, and interactive use cases where sub-second generation speed is a requirement. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.0002/char
ElevenLabs Multilingual V2 is ElevenLabs' multi-language text-to-speech model supporting 29 languages with natural voice synthesis. The model produces high-quality speech in each supported language with native-sounding pronunciation, prosody, and intonation, making it the broadest multilingual TTS option available on the platform. At $0.05 per generation on FairStack, compared to $0.30 or more through ElevenLabs direct, it delivers the same premium voice quality at a significant discount. The 29-language coverage includes major world languages and many regional languages, with natural-sounding output in each. Voice quality is consistent across languages, with good handling of language-specific phonemes and rhythm patterns. Compared to single-language TTS models, Multilingual V2 provides access to a wide range of languages from a single model with consistent quality. Against other multilingual options like Chatterbox Multilingual with 23 languages, ElevenLabs offers broader coverage and generally higher audio quality. Best suited for multilingual content creation, international voice production, and multi-language TTS workflows where broad language coverage at premium quality matters. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.0001/char
IndexTTS2 is FairStack's self-hosted voice cloning model delivering the highest fidelity to reference audio available on the platform. The model produces natural-sounding speech that closely matches the tone, cadence, timbre, and unique character of the provided voice sample, making it the best choice when preserving the distinctive qualities of a specific voice is critical. The model works with all 168 voices in FairStack's library and accepts custom reference audio for cloning. Unlike some voice cloning models that require a reference text transcript, IndexTTS2 works from audio alone, simplifying the cloning workflow. Self-hosting ensures low latency when the model is warm and eliminates third-party API costs. Compared to Qwen3-TTS which offers both voice design and cloning, IndexTTS2 focuses exclusively on cloning with superior reference fidelity. Against cloud-based voice cloning services, self-hosting provides cost advantages and privacy benefits. Best suited for voice cloning, audiobook narration in a specific voice, and character voice preservation where maximum fidelity to the reference voice is the priority. Available on FairStack at infrastructure cost plus a 20% platform fee.
Step 1
Quiet room, natural pace, 30–60 seconds of speech.
Step 2
Upload in the FairStack voice workspace.
Step 3
Generate with any selectable type=tts model that accepts your clone.
Step 4
Credits never expire — come back when the script is ready.
We do not train on uploads. Read provider terms for upstream engines.
Ship weekly updates without re-recording every line.
Keep one voice across episodes.
Cloning is a product capability. This page lists selectable TTS models (type tts) you generate with after cloning — derived from config, not invented slugs.
Still have questions? We're here to help.