EXPLORE / MODEL DIRECTORY

Explore models.Know what runs.

Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.

3Kaista-ready endpointsPublic catalog connected
PUBLIC REFERENCE CATALOG

All model endpoints

Page 11 · 28 per page
image-to-video
KlingREVIEW REQUIRED

Kling O1 First Frame Last Frame to Video [Pro]

fal-ai/kling-video/o1/image-to-video

Generate a video by taking a start frame and an end frame, animating the transition between them while following text-driven style and scene guidance.

video-to-video
falREVIEW REQUIRED

Depth Anything Video

fal-ai/depth-anything-video

Generates depth maps from video using Video Depth Anything (CVPR 2025). Produces per-frame depth estimation with temporal consistency across frames. Supports 3 model sizes (Small, Base, Large), 5 colormaps including grayscale, side-by-side comparison with the original video, and raw depth export as .npz. Useful for 3D reconstruction, video effects, compositing, and scene understanding.

video to videomotionedit
image-to-image
AlibabaREVIEW REQUIRED

Qwen Image Edit Plus

fal-ai/qwen-image-edit-plus

Endpoint for Qwen's Image Editing Plus model also known as Qwen-Image-Edit-2509. Has superior text editing capabilities and multi-image support.

image-editingimage-to-imagehigh-quality-text
text-to-image
Black Forest LabsREVIEW REQUIRED

Flux 2 Flex

fal-ai/flux-2-flex

Text-to-image generation with FLUX.2 [flex] from Black Forest Labs. Features adjustable inference steps and guidance scale for fine-tuned control. Enhanced typography and text rendering capabilities.

stylizedtransform
text-to-video
GoogleREVIEW REQUIRED

Gemini Omni Flash

google/gemini-omni-flash

Creates video with synchronized audio from text input. Grounded in Gemini's real-world knowledge, with improved physics understanding for more coherent motion and interaction.

stylizedtransformlipsync
text-to-audio
soniloREVIEW REQUIRED

V1.1 Text to Sound Effects

sonilo/v1.1/text-to-sound-effects

Generates high-quality, commercial-use-safe sound effects from a text prompt, with full control over type, texture, intensity, and exact duration.

sfxaudioeffects
video-to-video
KlingREVIEW REQUIRED

Kling O3 Reference Video to Video [Pro]

fal-ai/kling-video/o3/pro/video-to-video/reference

Kling O3 Omni generates new shots guided by an input reference video, preserving cinematic language such as motion, and camera style to produce seamless scene continuity.

video-to-video
image-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.1 Kontext [pro]

fal-ai/flux-pro/kontext/multi

Experimental version of FLUX.1 Kontext [pro] with multi image handling capabilities

speech-to-text
ElevenLabsREVIEW REQUIRED

Elevenlabs - Forced Alignment

fal-ai/elevenlabs/forced-alignment

Align the transcript and your audio recording using Elevenlab's forced alignment feature!

forced-alignmentspeech-to-text
image-to-video
xAIREVIEW REQUIRED

Grok Imagine Reference to Video

xai/grok-imagine-video/reference-to-video

Generate videos using multiple reference images with xAI's Grok Imagine video model

video-editv2vgrokxai
image-to-video
MiniMaxREVIEW REQUIRED

MiniMax Hailuo 2.3 Fast [Standard] (Image to Video)

fal-ai/minimax/hailuo-2.3-fast/standard/image-to-video

MiniMax Hailuo-2.3-Fast Image To Video API (Standard, 768p): Advanced fast image-to-video generation model with 768p resolution

image-to-video
video-to-video
MiniMaxREVIEW REQUIRED

H3 Max Turbo Extend Video

minimax/h3-max-turbo/extend-video

Extend an existing video with H3 Max Turbo: add 1 to 15 seconds of prompt-guided footage, with optional audio references and output resolutions from 480p to 2K.

extendvideocontinuation
image-to-image
Black Forest LabsREVIEW REQUIRED

Flux 2 Flex

fal-ai/flux-2-flex/edit

Image editing with FLUX.2 [flex] from Black Forest Labs. Supports multi-reference editing with customizable inference steps and enhanced text rendering.

text-to-speech
MiniMaxREVIEW REQUIRED

MiniMax Speech-02 Turbo

fal-ai/minimax/speech-02-turbo

Generate fast speech from text prompts and different voices using the MiniMax Speech-02 Turbo model, which leverages advanced AI techniques to create high-quality text-to-speech.

speech
vision
falREVIEW REQUIRED

NSFW Checker

fal-ai/x-ailab/nsfw

Predict whether an image is NSFW or SFW.

filtersafetyutility
video-to-video
briaREVIEW REQUIRED

Bria's VRMBG 3.0

bria/video/background-removal/v3

Remove backgrounds from any video with Bria's VRMBG 3.0. Fast, accurate background removal across talking heads, podcasts, product videos, commercials, and cinematic footage.

video-to-video
text-to-video
PixVerseREVIEW REQUIRED

PixVerse V6 Text To Video

fal-ai/pixverse/v6/text-to-video

Pixverse's latest v6 Model.

text-to-video
image-to-image
KlingREVIEW REQUIRED

Kling Image

fal-ai/kling-image/o3/image-to-image

Kling Omni 3: Top-tier image-to-image with flawless consistency.

image-to-image
image-to-image
falREVIEW REQUIRED

Z Image Turbo Image To Image

fal-ai/z-image/turbo/image-to-image

Generate images from text and images using Z-Image Turbo, Tongyi-MAI's super-fast 6B model.

turboz-imagefast
audio-to-audio
ElevenLabsREVIEW REQUIRED

ElevenLabs Voice Changer

fal-ai/elevenlabs/voice-changer

Change the voices in your audios with voices in ElevenLabs!

voice-changeaudio-to-audio
image-to-3d
tripo3dREVIEW REQUIRED

Tripo3D

tripo3d/tripo/v2.5/image-to-3d

State of the art Image to 3D Object generation. Generate 3D model from a single image!

image-to-3dstylized
unknown
openrouterREVIEW REQUIRED

OpenRouter [Audio]

openrouter/router/audio

Run any audio capable LLM with fal. Process audio files — transcription, analysis, understanding, understand— using Gemini (Google) models. Supports wav, mp3, aiff, aac, ogg, flac, m4a. Powered by OpenRouter.

image-to-3d
falREVIEW REQUIRED

Sam 3

fal-ai/sam-3/3d-objects

SAM 3D enables precise 3D reconstruction of objects from real images, while accurately reconstructing their geometry and texture.

3dobject
image-to-image
Black Forest LabsREVIEW REQUIRED

Flux Kontext Lora

fal-ai/flux-kontext-lora

Fast endpoint for the FLUX.1 Kontext [dev] model with LoRA support, enabling rapid and high-quality image editing using pre-trained LoRA adaptations for specific styles, brand identities, and product-specific outputs.

image-editingimage-to-image
video-to-video
falREVIEW REQUIRED

Birefnet

fal-ai/birefnet/v2/video

Video background removal version of bilateral reference framework (BiRefNet) for high-resolution dichotomous image segmentation (DIS)

utilityediting
text-to-audio
ElevenLabsREVIEW REQUIRED

Elevenlabs

fal-ai/elevenlabs/text-to-dialogue/eleven-v3

Generate realistic audio dialogues using Eleven-v3 from ElevenLabs.

audio
video-to-video
veedREVIEW REQUIRED

Subtitles

veed/subtitles

VEED’s Subtitles API transforms raw footage into polished, publish-ready content with professional burned-in subtitles starting at a base rate of $0.10 per minute.

training
falREVIEW REQUIRED

Krea 2 Trainer

fal-ai/krea-2-trainer

Train a custom LoRA on your own images to teach Krea 2 a new subject, character, or style. Provide a set of training images (and an optional trigger word), and the trainer outputs LoRA weights you can use for inference with the Krea 2 LoRA endpoint.

lorapersonalization