EXPLORE / MODEL DIRECTORY

Explore models.Know what runs.

Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.

3Kaista-ready endpointsPublic catalog connected
PUBLIC REFERENCE CATALOG

All model endpoints

Page 19 · 28 per page
audio-to-audio
falREVIEW REQUIRED

ACE Step Audio To Audio

fal-ai/ace-step/audio-to-audio

Generate music from a lyrics and example audio using ACE-Step

audio-to-audioaudio-edit
image-to-image
falREVIEW REQUIRED

Marigold V2 Depth

fal-ai/marigold-v2

Estimate depth from a single image with Marigold V2, a diffusion-based depth model built on Qwen-Image-Edit, returning a colorized depth map.

depthutility
text-to-text
openrouterREVIEW REQUIRED

Router

openrouter/router/decisions

Run any decision model with fal, powered by OpenRouter.

video-to-video
KlingREVIEW REQUIRED

Kling Video 4K Video to Video

fal-ai/kling-video/o3/4k/video-to-video/reference

Kling's Native 4K is a video generation model that directly outputs professional-grade 4K video in one step, eliminating the need for post-production upscaling

utilityediting
image-to-image
falREVIEW REQUIRED

Feynobg Background Remover

fal-ai/feynobg

FeyNobg is a state of the art AI model for background removal from feyninc

utilityediting
image-to-image
briaREVIEW REQUIRED

Fibo Edit 1.5 Image Editing

bria/fibo-edit-1.5/edit

Commercially safe, multi-reference image editing model. Follows natural language instructions alone or with up to 4 reference images, purpose-built for complex object and character combinations, virtual try-on, background replacement, style transfer, and more.

stylizedtransformediting
image-to-video
falREVIEW REQUIRED

LTX Video (preview)

fal-ai/ltx-video/image-to-video

Generate videos from images using LTX Video

image-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.2 [klein] 4B Base

fal-ai/flux-2/klein/4b/base/edit

Image-to-image editing with FLUX.2 [klein] 4B Base from Black Forest Labs. Precise modifications using natural language descriptions and hex color control.

text-to-video
AlibabaREVIEW REQUIRED

Happy Horse

alibaba/happy-horse/text-to-video

Generate 1080p video with synchronized native audio from a text prompt. Aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4. Duration: 3–15s.

happy-horse
image-to-3d
tripo3dREVIEW REQUIRED

Triposplat

tripo3d/triposplat

TripoSplat is an open-source model from TripoAI / VAST AI Research that converts a single 2D image into high-quality 3D Gaussians using a novel learned density-control approach

3Dgaussian-splat
text-to-video
AlibabaREVIEW REQUIRED

Wan

fal-ai/wan/v2.2-a14b/text-to-video/turbo

Wan-2.2 turbo text-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts.

text to videomotion
video-to-video
falREVIEW REQUIRED

Sam 3

fal-ai/sam-3/video-rle

SAM 3 is a unified foundation model for promptable segmentation in images and videos. It can detect, segment, and track objects using text or visual prompts such as points, boxes, and masks.

segmentationmaskreal-timerle
text-to-image
AlibabaREVIEW REQUIRED

Qwen Image 2512

fal-ai/qwen-image-2512/lora

LoRA inference endpoint for Qwen Image 2512, an improved version of Qwen Image with better text rendering, finer natural textures, and more realistic human generation.

qwen2512lora
vision
falREVIEW REQUIRED

Florence-2 Large

fal-ai/florence-2-large/more-detailed-caption

Florence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks

captioningmultimodalvision
image-to-image
falREVIEW REQUIRED

DDColor

fal-ai/ddcolor

Bring colors into old or new black and white photos with DDColor.

image-recolorizationfacesutility
image-to-video
falREVIEW REQUIRED

Pika Image to Video (v2.2)

fal-ai/pika/v2.2/image-to-video

Turn photos into mind-blowing, dynamic videos in up to 1080p. Experience better image clarity and crisper, sharper visuals.

editingeffectsanimation
video-to-audio
KlingREVIEW REQUIRED

Kling Video

fal-ai/kling-video/video-to-audio

Generate audio from input videos using Kling

llm
OpenAIREVIEW REQUIRED

OpenRouter Embeddings [OpenAI Compatible]

openrouter/router/openai/v1/embeddings

Generate text embeddings using OpenAI-compatible API. Access embedding models like text-embedding-3-small, text-embedding-3-large (OpenAI), and other embedding models available through OpenRouter. Drop-in replacement for the OpenAI embeddings API. Powered by OpenRouter.

image-to-video
AlibabaREVIEW REQUIRED

Happy Horse 1.1 Reference to Video

alibaba/happy-horse/v1.1/reference-to-video

Happy Horse 1.1 is Alibaba's #1-ranked video model. This reference-to-video endpoint turns up to 9 reference images into 1080p video with synchronized native audio and multilingual lip-sync for consistent characters.

happy-horsevideoreference
vision
falREVIEW REQUIRED

Moondream2

fal-ai/moondream2/visual-query

Moondream2 is a highly efficient open-source vision language model that combines powerful image understanding capabilities with a remarkably small footprint.

Vision
text-to-video
MiniMaxREVIEW REQUIRED

MiniMax Hailuo 2.3 [Standard] (Text to Video)

fal-ai/minimax/hailuo-2.3/standard/text-to-video

MiniMax Hailuo-2.3 Text To Video API (Standard, 768p): Advanced text-to-video generation model with 768p resolution

text-to-video
image-to-video
AlibabaREVIEW REQUIRED

Wan 2.7 Reference to Video

fal-ai/wan/v2.7/reference-to-video

Wan 2.7 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.

stylizedtransformlipsync
image-to-image
falREVIEW REQUIRED

Smart Resize

fal-ai/smart-resize

Smart image resize to arbitrary dimensions, powered by Nano Banana Pro with vision-LLM-guided prompting for composition-aware recomposition. Crop, cropping, resize ads.

realismtypographyvisualads
image-to-image
falREVIEW REQUIRED

Florence-2 Large

fal-ai/florence-2-large/caption-to-phrase-grounding

Florence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks

multimodalvision
audio-to-video
ElevenLabsREVIEW REQUIRED

ElevenLabs Dubbing

fal-ai/elevenlabs/dubbing

Generate dubbed videos or audios using ElevenLabs Dubbing feature!

dubbingaudio-to-audio
video-to-video
topazREVIEW REQUIRED

Topaz Upscale Video Creative

topaz/upscale/video/creative

Professional creative video upscaling powered by Topaz Labs. Astra 2 reimagines fine detail and typically delivers 4K output. Best for cinematic shots that need maximum visual impact.

upscalevideo
video-to-video
mirelo-aiREVIEW REQUIRED

Mirelo SFX1.6

mirelo-ai/sfx1.6/video-to-video

Generate synced sounds for any video, and return it with its new sound track (like MMAudio). Now up to 60 seconds!

video-to-videosfx