EXPLORE / MODEL DIRECTORY

Explore models.Know what runs.

Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.

3Kaista-ready endpointsPublic catalog connected
PUBLIC REFERENCE CATALOG

All model endpoints

Page 6 · 28 per page
image-to-video
Black Forest LabsREVIEW REQUIRED

FLUX 3 Image to Video

blackforestlabs/flux-3/image-to-video

FLUX 3 is Black Forest Labs' frontier video model. This endpoint animates a single still image into video, extending one frame into coherent, natural motion.

stylizedtransformlipsync
video-to-video
falREVIEW REQUIRED

LatentSync

fal-ai/latentsync

LatentSync is a video-to-video model that generates lip sync animations from audio using advanced algorithms for high-quality synchronization.

animationlip sync
image-to-image
AlibabaREVIEW REQUIRED

Qwen Image 3 Image Editing

alibaba/qwen-image-3/edit

Edits images from one to three reference images and a natural-language instruction, preserving key details such as facial features and identity while applying the requested changes

stylizedtransformtypography
image-to-video
ByteDanceREVIEW REQUIRED

Seedance 1.0 Pro

fal-ai/bytedance/seedance/v1/pro/image-to-video

Seedance 1.0 Pro, a high quality video generation model developed by Bytedance.

video-to-video
falREVIEW REQUIRED

FFmpeg API Compose

fal-ai/ffmpeg-api/compose

Compose videos from multiple media sources using FFmpeg API.

ffmpeg
text-to-image
Black Forest LabsREVIEW REQUIRED

FLUX 2 Turbo

fal-ai/flux-2/turbo

Text-to-image generation with FLUX.2 [dev] from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities—all at turbo speed.

text-to-video
KlingREVIEW REQUIRED

Kling Video v3 Text to Video [Standard]

fal-ai/kling-video/v3/standard/text-to-video

Kling 3.0 Standard: Top-tier text-to-video with cinematic visuals, fluid motion, and native audio generation, with multi-shot support.

text-to-video
image-to-image
Black Forest LabsREVIEW REQUIRED

Flux 2 Max

fal-ai/flux-2-max/edit

FLUX.2 [max] delivers state-of-the-art image generation and advanced image editing with exceptional realism, precision, and consistency.

flux2image-editinghigh-quality
image-to-video
xAIREVIEW REQUIRED

Grok Imagine Video 1.5 Lite Image to Video

xai/grok-imagine-video/v1.5/lite/image-to-video

Generate videos from images using xAI's Grok Imagine Video 1.5 Lite model.

image-to-video
image-to-image
AlibabaREVIEW REQUIRED

Qwen Image Edit 2511 Multiple Angles

fal-ai/qwen-image-edit-2511-multiple-angles

Generates same scene from different angles (azimuth/elevation) with Qwen image Edit 2511 and the Lora Multiple Angles

stylizedtransformloramulti-angles
image-to-3d
tripo3dREVIEW REQUIRED

Tripo H3.1 Image to 3D

tripo3d/h3.1/image-to-3d

Generate high-quality 3D models from a single image using Tripo H3.1.

3dimage-to-3d3d-generationtripo
image-to-image
topazREVIEW REQUIRED

Topaz Upscale Image Generative

topaz/upscale/image/generative

Professional generative image upscaling powered by Topaz Labs. Wonder 3.5 leads the range, with Redefine for prompt-guided detail and Recovery for extreme low-resolution sources. Best for rebuilding sharp detail in small or blurry images.

upscaleimage
video-to-video
KlingREVIEW REQUIRED

Kling Video v2.6 Motion Control [Standard]

fal-ai/kling-video/v2.6/standard/motion-control

Transfer movements from a reference video to any character image. Cost-effective mode for motion transfer, perfect for portraits and simple animations.

text-to-video
MiniMaxREVIEW REQUIRED

MiniMax H3 Text to Video

minimax/h3/text-to-video

MiniMax H3 is a frontier video model. This endpoint generates video from a text prompt alone, rendering at 2K in durations from 5 to 15 seconds across seven aspect ratios.

stylizedtransformlipsync
image-to-video
veedREVIEW REQUIRED

Fabric 1.0

veed/fabric-1.0

VEED Fabric 1.0 is an image-to-video API that turns any image into a talking video

lipsyncavatar
text-to-video
GoogleREVIEW REQUIRED

Veo 3.1

fal-ai/veo3.1

Veo 3.1 by Google, the most advanced AI video generation model in the world. With sound on!

text-to-video
AlibabaREVIEW REQUIRED

Wan Text to Video

alibaba/wan-3.0/text-to-video

Wan 3.0 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.

stylizedtransformlipsync
text-to-video
KlingREVIEW REQUIRED

Kling LipSync Audio-to-Video

fal-ai/kling-video/lipsync/audio-to-video

Kling LipSync is an audio-to-video model that generates realistic lip movements from audio input.

audio to videolipsync
text-to-audio
falREVIEW REQUIRED

Lyria2

fal-ai/lyria2

Lyria 2 is Google's latest music generation model, you can generate any type of music with this model.

musicstylized
video-to-video
KlingREVIEW REQUIRED

Kling Video

fal-ai/kling-video/v3/standard/motion-control

Transfer movements from a reference video to any character image. Cost-effective mode for motion transfer, perfect for portraits and simple animations.

stylizedtransformediting
text-to-speech
MiniMaxREVIEW REQUIRED

MiniMax Voice Cloning

fal-ai/minimax/voice-clone

Clone a voice from a sample audio and generate speech from text prompts using the MiniMax model, which leverages advanced AI techniques to create high-quality text-to-speech.

speech
image-to-image
RecraftREVIEW REQUIRED

Recraft

fal-ai/recraft/vectorize

Converts a given raster image to SVG format using Recraft model.

stylizedtransform
image-to-video
AlibabaREVIEW REQUIRED

Wan 3.0 Prime

alibaba/wan-3.0-prime/image-to-video

Wan 3.0 Prime Image-to-Video turns still images into dynamic, cinematic sequences with rapid turnaround, natural motion, and excellent visual continuity. It preserves the identity, composition, and atmosphere of the source image while introducing expressive movement, camera dynamics, and richly detailed animation.

imagevideo
text-to-audio
cassetteaiREVIEW REQUIRED

music generator

cassetteai/music-generator

CassetteAI’s model generates a 30-second sample in under 2 seconds and a full 3-minute track in under 10 seconds. At 44.1 kHz stereo audio, expect a level of professional consistency with no breaks, no squeaks, and no random interruptions in your creations.

musiccassetteai
image-to-video
ByteDanceREVIEW REQUIRED

Bytedance Seedance V1 Pro Fast Image To Video

fal-ai/bytedance/seedance/v1/pro/fast/image-to-video

Image to Video endpoint for Seedance 1.0 Pro Fast, a next-generation video model designed to deliver maximum performance at minimal cost

bytedanceseedanceprofast
text-to-image
OpenAIREVIEW REQUIRED

GPT-Image 1.5

fal-ai/gpt-image-1.5

GPT Image 1.5 generates high-fidelity images with strong prompt adherence, preserving composition, lighting, and fine-grained detail.

openaigpt-image
text-to-image
xAIREVIEW REQUIRED

Grok Imagine Image 2.0

xai/grok-imagine-image/v2.0/text-to-image

Generate images from text using xAi's Grok Imagine 2.0 model.

text-to-imagexaigrok
image-to-video
KlingREVIEW REQUIRED

Kling Video V3 Standard Turbo Image to Video

fal-ai/kling-video/v3/turbo/standard/image-to-video

Kling 3.0 Turbo Standard animates a first and last frame reference image into 720P video with native audio, delivering quick, affordable image-driven motion for fast turnaround

stylizedtransformlipsync