text-to-videoExplore models.Know what runs.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
All model endpoints
text-to-video
audio-to-videoLTX-2.3 22B Distilled
fal-ai/ltx-2.3-22b/distilled/audio-to-videoGenerate video with audio from audio, text and images using LTX-2 Distilled
text-to-imageIdeogram V2 Turbo
fal-ai/ideogram/v2/turboAccelerated image generation with Ideogram V2 Turbo. Create high-quality visuals, posters, and logos with enhanced speed while maintaining Ideogram's signature quality.
image-to-videoVidu
fal-ai/vidu/q1/reference-to-videoGenerate video clips from your multiple image references using Vidu Q1
text-to-imageAuraFlow
fal-ai/aura-flowAuraFlow v0.3 is an open-source flow-based text-to-image generation model that achieves state-of-the-art results on GenEval. The model is currently in beta.
image-to-imageStable Diffusion V3
fal-ai/stable-diffusion-v3-medium/image-to-imageStable Diffusion 3 Medium (Image to Image) is a Multimodal Diffusion Transformer (MMDiT) model that improves image quality, typography, prompt understanding, and efficiency.
video-to-videoVideo Background Removal
veed/video-background-removal/green-screenRemove background from videos filmed using chromakey, with automatic green spill suppression for clean, professional edges.
text-to-audioDiffRhythm: Lyrics to Song
fal-ai/diffrhythmDiffRhythm is a blazing fast model for transforming lyrics into full songs. It boasts the capability to generate full songs in less than 30 seconds.
text-to-imageQwen Image Max
fal-ai/qwen-image-max/text-to-imageText-to-Image endpoint for Qwen-Image-Max. Qwen Image Max improves upon the Qwen Image Plus series by enhancing the realism and naturalness of images.
audio-to-textSilero VAD
fal-ai/silero-vadDetect speech presence and timestamps with accuracy and speed using the ultra-lightweight Silero VAD model
image-to-imageIP Adapter Face ID
fal-ai/ip-adapter-face-idHigh quality zero-shot personalization
image-to-imageReplace Background
bria/replace-backgroundGenerate professional, eCommerce-ready product shots by replacing backgrounds with realistic lighting and accurate perspective from a simple text prompt. Trained exclusively on licensed data for safe commercial use.
text-to-imageIdeogram V2A Turbo
fal-ai/ideogram/v2a/turboAccelerated image generation with Ideogram V2A Turbo. Create high-quality visuals, posters, and logos with enhanced speed while maintaining Ideogram's signature quality.
3d-to-3dTripo3D Segment
tripo3d/tripo/segmentAutomatically splits a 3D model into semantic parts for editing, texturing, and rigging.
text-to-speechVibeVoice 1.5B
fal-ai/vibevoiceGenerate long, expressive multi-voice speech using Microsoft's powerful TTS
trainingRecraft V4 Styles Create Style
recraft/v4/create-styleCreates a reusable style from your reference images and returns a style ID you can pass to Recraft V4 Styles image and vector generation.
video-to-videoThinkSound
fal-ai/thinksoundGenerate realistic audio for a video with an optional text prompt and combine
text-to-imageHidream I1 Dev
fal-ai/hidream-i1-devHiDream-I1 dev is a new open-source image generative foundation model with 17B parameters that achieves state-of-the-art image generation quality within seconds.
text-to-imageFLUX.2 [klein] 4B LoRA
fal-ai/flux-2/klein/4b/loraText-to-image generation with FLUX.2 [klein] 4B from Black Forest Labs and custom LoRA. Enhanced realism, crisper text generation, and native editing capabilities.
text-to-videoHunyuan Video V1.5
fal-ai/hunyuan-video-v1.5/text-to-videoHunyuan Video 1.5 is Tencent's latest and best video model
text-to-videoHeygen Video Agent
fal-ai/heygen/v3/video-agentGenerate videos with a single prompt. Describe what you want in plain text, and the agent handles avatar selection, scripting, scene composition - all in one.
text-to-videoCogVideoX-5B
fal-ai/cogvideox-5bGenerate videos from prompts using CogVideoX-5B
text-to-imagePony V7
fal-ai/pony-v7Pony V7 is a finetuned text to image for superior aesthetics and prompt following.
text-to-audioKokoro TTS (Japanese)
fal-ai/kokoro/japaneseA fast and natural-sounding Japanese text-to-speech model optimized for smooth pronunciation.
image-to-videoFlux 3 Keyframes To Video Draft
blackforestlabs/flux-3/keyframes-to-video/draftFLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews pinned to your keyframe images, with a reusable draft cache for full-quality enhancement.
image-to-imageLeffa Virtual TryOn
fal-ai/leffa/virtual-tryonLeffa Virtual TryOn is a high quality image based Try-On endpoint which can be used for commercial try on.
text-to-videoLTX-2.3 22B Distilled
fal-ai/ltx-2.3-22b/distilled/text-to-videoGenerate video with audio from text using LTX-2.3 Distilled
text-to-video