EXPLORE / MODEL DIRECTORY

Explore models.Know what runs.

Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.

3Kaista-ready endpointsPublic catalog connected
PUBLIC REFERENCE CATALOG

All model endpoints

Page 33 · 28 per page
audio-to-video
falREVIEW REQUIRED

LTX-2.3 22B Distilled

fal-ai/ltx-2.3-22b/distilled/audio-to-video

Generate video with audio from audio, text and images using LTX-2 Distilled

text-to-image
IdeogramREVIEW REQUIRED

Ideogram V2 Turbo

fal-ai/ideogram/v2/turbo

Accelerated image generation with Ideogram V2 Turbo. Create high-quality visuals, posters, and logos with enhanced speed while maintaining Ideogram's signature quality.

realismtypography
image-to-video
falREVIEW REQUIRED

Vidu

fal-ai/vidu/q1/reference-to-video

Generate video clips from your multiple image references using Vidu Q1

stylizedtransform
text-to-image
falREVIEW REQUIRED

AuraFlow

fal-ai/aura-flow

AuraFlow v0.3 is an open-source flow-based text-to-image generation model that achieves state-of-the-art results on GenEval. The model is currently in beta.

typographystyle
image-to-image
falREVIEW REQUIRED

Stable Diffusion V3

fal-ai/stable-diffusion-v3-medium/image-to-image

Stable Diffusion 3 Medium (Image to Image) is a Multimodal Diffusion Transformer (MMDiT) model that improves image quality, typography, prompt understanding, and efficiency.

diffusioneditingstyle
video-to-video
veedREVIEW REQUIRED

Video Background Removal

veed/video-background-removal/green-screen

Remove background from videos filmed using chromakey, with automatic green spill suppression for clean, professional edges.

text-to-audio
falREVIEW REQUIRED

DiffRhythm: Lyrics to Song

fal-ai/diffrhythm

DiffRhythm is a blazing fast model for transforming lyrics into full songs. It boasts the capability to generate full songs in less than 30 seconds.

music
text-to-image
AlibabaREVIEW REQUIRED

Qwen Image Max

fal-ai/qwen-image-max/text-to-image

Text-to-Image endpoint for Qwen-Image-Max. Qwen Image Max improves upon the Qwen Image Plus series by enhancing the realism and naturalness of images.

qwen-imagemax
audio-to-text
falREVIEW REQUIRED

Silero VAD

fal-ai/silero-vad

Detect speech presence and timestamps with accuracy and speed using the ultra-lightweight Silero VAD model

vadsilerovoice-activity-detection
image-to-image
falREVIEW REQUIRED

IP Adapter Face ID

fal-ai/ip-adapter-face-id

High quality zero-shot personalization

ip-adapterpersonalizationcustomizationediting
image-to-image
briaREVIEW REQUIRED

Replace Background

bria/replace-background

Generate professional, eCommerce-ready product shots by replacing backgrounds with realistic lighting and accurate perspective from a simple text prompt. Trained exclusively on licensed data for safe commercial use.

briareplace-background
text-to-image
IdeogramREVIEW REQUIRED

Ideogram V2A Turbo

fal-ai/ideogram/v2a/turbo

Accelerated image generation with Ideogram V2A Turbo. Create high-quality visuals, posters, and logos with enhanced speed while maintaining Ideogram's signature quality.

realismtypography
3d-to-3d
tripo3dREVIEW REQUIRED

Tripo3D Segment

tripo3d/tripo/segment

Automatically splits a 3D model into semantic parts for editing, texturing, and rigging.

stylizedtransform
text-to-speech
falREVIEW REQUIRED

VibeVoice 1.5B

fal-ai/vibevoice

Generate long, expressive multi-voice speech using Microsoft's powerful TTS

text-to-speechmulti-speakerpodcast
training
RecraftREVIEW REQUIRED

Recraft V4 Styles Create Style

recraft/v4/create-style

Creates a reusable style from your reference images and returns a style ID you can pass to Recraft V4 Styles image and vector generation.

transformutilitystylized
video-to-video
falREVIEW REQUIRED

ThinkSound

fal-ai/thinksound

Generate realistic audio for a video with an optional text prompt and combine

audio-generationvideo-to-audio
text-to-image
falREVIEW REQUIRED

Hidream I1 Dev

fal-ai/hidream-i1-dev

HiDream-I1 dev is a new open-source image generative foundation model with 17B parameters that achieves state-of-the-art image generation quality within seconds.

text-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.2 [klein] 4B LoRA

fal-ai/flux-2/klein/4b/lora

Text-to-image generation with FLUX.2 [klein] 4B from Black Forest Labs and custom LoRA. Enhanced realism, crisper text generation, and native editing capabilities.

text-to-video
falREVIEW REQUIRED

Hunyuan Video V1.5

fal-ai/hunyuan-video-v1.5/text-to-video

Hunyuan Video 1.5 is Tencent's latest and best video model

hunyuan-videotext-to-video
text-to-video
falREVIEW REQUIRED

Heygen Video Agent

fal-ai/heygen/v3/video-agent

Generate videos with a single prompt. Describe what you want in plain text, and the agent handles avatar selection, scripting, scene composition - all in one.

text-to-image
falREVIEW REQUIRED

Pony V7

fal-ai/pony-v7

Pony V7 is a finetuned text to image for superior aesthetics and prompt following.

diffusionstyle
text-to-audio
falREVIEW REQUIRED

Kokoro TTS (Japanese)

fal-ai/kokoro/japanese

A fast and natural-sounding Japanese text-to-speech model optimized for smooth pronunciation.

speech
image-to-video
Black Forest LabsREVIEW REQUIRED

Flux 3 Keyframes To Video Draft

blackforestlabs/flux-3/keyframes-to-video/draft

FLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews pinned to your keyframe images, with a reusable draft cache for full-quality enhancement.

stylizedtransformlipsync
image-to-image
falREVIEW REQUIRED

Leffa Virtual TryOn

fal-ai/leffa/virtual-tryon

Leffa Virtual TryOn is a high quality image based Try-On endpoint which can be used for commercial try on.

try-onfashionclothing
text-to-video
falREVIEW REQUIRED

LTX-2.3 22B Distilled

fal-ai/ltx-2.3-22b/distilled/text-to-video

Generate video with audio from text using LTX-2.3 Distilled

text-to-video
veedREVIEW REQUIRED

Fabric 1.0

veed/fabric-1.0/text

VEED Fabric 1.0 text-to-video API

lipsyncavatartext-to-video