Step 1
Open Studio and pick a t2v model
Go to Studio → Video and choose a model from the text-to-video list on this page.
Turn a written prompt into a finished clip. FairStack only lists models whose config type or supportedModes include text-to-video — image-only models like Sora 2 never appear here.
Seedance 2.0 · Prompt: Morning dew on alpine grass, slow lateral track, soft backlight
Text-to-video AI generates a short clip from a written prompt — subject, camera, and mood — without a start frame. On FairStack, only models with a t2v type or t2v supportedModes are listed.
Try text to video →14 selectable models — filtered from config capability flags, not a hand-maintained list.
$0.123/sec
Seedance 2.0 Mini is ByteDance’s most affordable Seedance 2 model, built for high-volume video generation where cost matters more than maximum resolution. Like the rest of the Seedance 2 family it is multi-mode from a single endpoint: drive it with a text prompt, a single first frame, first+last frames for interpolation, or up to nine reference images plus video/audio references — FairStack auto-detects the mode from your input. It renders 480p or 720p clips from 4 to 15 seconds with native audio generation, and supports seven aspect ratios (including 21:9 and adaptive). At roughly $0.0625 per second it is the value pick of the Seedance 2 lineup, trading the 1080p/4K ceiling of the standard model for materially lower cost. Best for social-media shorts, rapid iteration, storyboard animation, and any workflow that generates a lot of video. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.198/sec
Seedance 2.0 Fast is a text to video model on FairStack. Pay $0.198/sec with a transparent 20% margin — no subscription.
$0.246/sec
Seedance 2.0 is a text to video model on FairStack. Pay $0.246/sec with a transparent 20% margin — no subscription.
$0.084/sec
WAN 2.6 T2V is Alibaba's latest text-to-video generation model, representing the most advanced entry in the WAN family. It generates 5-second video clips at 720p with notably strong artistic style and visual quality, making it particularly well-suited for stylized and creative content rather than strict photorealism. The model benefits from Alibaba's extensive research in video generation, delivering improved temporal coherence and motion quality over earlier WAN versions. It handles artistic prompts with distinctive flair, producing outputs with painterly qualities, rich color grading, and creative visual interpretations that set it apart from purely photorealistic competitors. Compared to Kling 3.0 which focuses on cinematic realism, WAN 2.6 T2V leans toward artistic expression and stylized output. Against Veo 3 which costs $0.20 per second, WAN 2.6 at $0.50 per generation offers a different aesthetic that may better serve creative and artistic content needs. The model is hosted on RunPod for reliable delivery. Best suited for artistic video content, music video concepts, stylized social media clips, and creative projects where visual artistry matters more than photorealism. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.162/sec
Kling 3.0 Pro from Kuaishou is the highest-rated video generation model globally by Elo score, delivering the best motion quality (0.90), excellent camera control (0.85), and support for clips up to 15 seconds. When video quality is the top priority and cost is secondary, this is the definitive choice. The model produces smooth, naturalistic movement that avoids the common artifacts of AI video — jitter, ghosting, and unnatural physics. Camera control follows complex instructions including pan, zoom, dolly, and orbital movements with high fidelity. Up to 15 seconds of continuous video is supported, the longest duration among top-tier models. Visual quality (0.88 score) is consistently high across subjects and scenes. Compared to Veo 3.1 Quality (the only model with native audio), Kling 3.0 Pro offers longer maximum duration and lower cost per second. Against Grok Imagine T2V (Elo #3), it delivers noticeably superior motion quality and camera control at a higher price. Direct pricing from Kling would cost approximately $0.40 per generation — the FairStack price of approximately $0.135 per second represents significant savings at infrastructure cost plus 20%. Best suited for premium video production where quality is paramount, long-form clips up to 15 seconds, cinematic content requiring smooth camera work, and professional video workflows. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.018/sec
Grok Imagine T2V from xAI is ranked Elo #3 for video generation, delivering strong overall quality at an exceptional price point of $0.10 per generation. It represents the best value proposition in the premium video category — priced below Kling Pro or Veo while still producing high-quality output suitable for professional use. The model generates up to 10 seconds of video with good motion quality (0.84) and strong visual quality (0.85). The balanced capability profile means it handles diverse subjects and scenes without significant weaknesses in any particular area. As a newer entrant in the video generation space, it benefits from xAI's latest research while building a growing track record in production workflows. Compared to Kling 3.0 Pro (Elo #1), Grok T2V trades some motion smoothness and camera control precision for a lower cost per generation. Against Veo 3.1 Fast (Elo #2), it offers similar quality at a competitive price with up to 10-second duration. Versus Runway Gen-4 Turbo ($0.06), it costs more but delivers meaningfully higher quality. The Elo #3 ranking at $0.10 per generation makes it the clear choice when quality and value must both be strong. Ideal for general-purpose video generation, best-value premium video, 10-second clips, and any budget-conscious video workflow that still demands high quality. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.024/sec
P-Video (Pruna) is a text to video model on FairStack. Pay $0.024/sec with a transparent 20% margin — no subscription.
$0.042/req
Vidu Q3 Turbo T2V is Vidu's budget text-to-video model optimized for fast generation at significantly lower cost than the standard Q3 tier. The model supports up to 16 seconds of video duration with built-in audio generation, offering one of the longest maximum durations among budget video models. With per-second billing at $0.077 per second, it is substantially priced below most video generation alternatives. A 5-second clip costs approximately $0.39, while a full 16-second clip costs about $1.23. Multiple resolution options are available, with 720p providing the best quality-to-cost balance. Compared to standard Vidu Q3 at $0.154 per second, the Turbo variant halves the per-second cost while maintaining usable quality for drafts and social media. Against premium models like Kling 2.6 at $0.275 flat, it costs less per generation for longer content. Best suited for draft video previews, social media content, and high-volume video production where budget-friendly per-second pricing enables cost-effective generation at scale. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.084/req
Vidu Q3 T2V is Vidu's text-to-video model with flexible per-second billing, generating balanced-quality video from text prompts at proportional cost. The model delivers reliable output with good motion quality and scene coherence, providing a solid option for general-purpose video generation with the advantage of duration-flexible pricing. With per-second billing at $0.154 per second, users pay only for the exact duration they need. This pricing structure is advantageous for short clips where flat-rate models may overprice the output, and remains competitive for longer sequences. The model handles a variety of prompt types with consistent quality. Compared to flat-rate models like Kling 2.6 at $0.275, Vidu Q3's per-second pricing is more economical for shorter clips but can add up for longer sequences. Against other per-second models like Kling V3 Standard at $0.168 per second, Vidu Q3 offers a modest cost advantage. The model provides a good balance of quality and pricing flexibility. Best suited for flexible-duration video generation and general text-to-video workflows where per-second billing provides cost control. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.120/sec
Kling 3.0 Standard is the cost-optimized variant of Kuaishou's video generation model, sharing the same architecture and maximum 15-second duration as the Pro variant while delivering good quality at a lower price point. It bridges the gap between budget video models and the premium Kling Pro tier. The model produces solid video with good motion quality (0.82 score), maintaining natural-looking movement and reasonable camera following. Supporting the same 15-second maximum as Pro means users can generate longer clips without the premium cost. The Kling architecture provides consistent visual coherence throughout the clip duration. Compared to Kling 3.0 Pro, the Standard variant shows a noticeable quality gap in motion smoothness, fine detail, and camera precision. Against Runway Gen-4 Turbo ($0.06/5s), it offers longer duration and higher overall quality at a higher price. Versus Hailuo 2.3 Fast ($0.08/6s), it provides the longer 15-second option with better motion quality. It occupies the practical middle ground for users who need more than budget quality but less than premium pricing. Best suited for budget premium video work, longer clips where 15-second duration matters, good-enough quality workflows, and projects where Kling Pro's cost premium is not justified. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.174/sec
HappyHorse 1.1 is Alibaba's flagship video model, a strict upgrade over HappyHorse 1.0: lower price, built-in native audio with Foley, and multilingual lip-sync. The text-to-video mode renders 3-15 second clips at 720p or 1080p across nine aspect ratios (16:9 through 9:21). Served from Kie.ai at $0.1125/s (720p) or $0.145/s (1080p) with an automatic fal.ai fallback. Best all-rounder when synchronized audio and speech matter.
$0.180/req
PixVerse V5.6 T2V is PixVerse's latest text-to-video model, known for producing high-quality video with a distinctive creative and artistic sensibility. The model excels at stylized content, generating video with strong visual aesthetics, rich color palettes, and compositional choices that lean toward artistic expression rather than strict photorealism. At $0.45 per generation, it sits in the premium tier alongside models from Runway and Luma. The price reflects the model's strong visual quality and its particular strength in creative and artistic content. The V5.6 architecture represents significant improvements over earlier PixVerse versions in both quality and consistency. Compared to photorealism-focused models like Kling 2.6, PixVerse V5.6 is stronger at stylized, artistic, and creative content but may not match the natural look of dedicated photorealistic generators. Against Runway Gen-4 at $0.50, it offers similar pricing with a different creative strength profile. Best suited for creative video content, artistic video production, and visually appealing output where stylistic quality matters. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.288/sec
HappyHorse 1.0 T2V is a text to video model on FairStack. Pay $0.288/sec with a transparent 20% margin — no subscription.
$1.53/req
Google Veo 3.1 Quality is the premium tier of Google's video generation model and the only video model that generates both 1080p video and synchronized audio in a single generation pass. No other model combines video and audio generation into one step, eliminating the need for separate audio production and synchronization. Top-tier visual quality (0.92 score) and motion quality (0.90) place it at the pinnacle of AI video generation capability. The 1080p output resolution is the highest available from any video model, delivering detail that is suitable for professional presentations and large-screen viewing. Native audio generation produces synchronized sound effects, ambient audio, and environmental sounds that match the visual content. Compared to Kling 3.0 Pro (Elo #1), Veo 3.1 Quality matches or exceeds visual quality while adding native audio and higher resolution. At approximately $1.25 per generation, it is the most expensive video model — roughly 10x the cost of mid-tier options. Against a workflow combining a separate video model plus audio generation, the one-step approach is simpler but may cost more. Maximum duration is 8 seconds per generation. Best suited for professional video production where both 1080p resolution and synchronized audio are required, premium marketing content, and workflows where the convenience of one-step video-plus-audio generation justifies the premium price. Available on FairStack at infrastructure cost plus a 20% platform fee.
Step 1
Go to Studio → Video and choose a model from the text-to-video list on this page.
Step 2
Lead with camera movement, then subject action, then lighting. Keep one idea per shot.
Step 3
Pick length and resolution the model supports. Price scales per second.
Step 4
Run the job, review the MP4, then iterate the prompt or switch models without changing accounts.
Step 5
Download the clip or feed it into image-to-video / video-to-video for a second pass.
Start with dolly, pan, orbit, or static locked-off — models latch onto the first motion cue.
Crowded prompts smear identity. Name one hero subject and one secondary.
35mm, 85mm, anamorphic — focal length cues stabilize framing.
Price is per second. Draft at shorter durations, then spend on the keeper.
Prompt atmosphere plates instead of licensing stock for every cut.
Show a moving board before you commit to a shoot day.
Generate brand-safe mood clips, then composite real product plates.
Every model on this page is filtered from config: type t2v, or supportedModes including t2v. The list updates when config changes — it is not hand-maintained.
Sora 2 on FairStack is image-to-video only (type i2v). It appears on /video/image-to-video and is structurally excluded from this page.
Per second at each model's charged rate. Example: Kling 3.0 Standard is $0.12/sec; Seedance 2 at 720p t2v is $0.246/sec. No invented clip pricing.
Still have questions? We're here to help.