EXPLORE / MODEL DIRECTORY

Explore models.Know what runs.

Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.

3Kaista-ready endpointsPublic catalog connected
PUBLIC REFERENCE CATALOG

All model endpoints

Page 3 · 28 per page
image-to-video
MiniMaxREVIEW REQUIRED

MiniMax H3 Image to Video

minimax/h3/image-to-video

MiniMax H3 is a frontier video model. This endpoint animates a supplied image into 2K video, using it as the opening frame or pairs a first and last frame to control a transition between two images with the aspect ratio following the input.

stylizedtransformlipsync
text-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.1 [dev] with LoRAs

fal-ai/flux-lora

Super fast endpoint for the FLUX.1 [dev] model with LoRA support, enabling rapid and high-quality image generation using pre-trained LoRA adaptations for personalization, specific styles, brand identities, and product-specific outputs.

lorapersonalization
image-to-video
ByteDanceREVIEW REQUIRED

Seedance 2 Reference to Video

bytedance/seedance-2.0/reference-to-video

ByteDance's most advanced reference-to-video model. Generate video from up to 9 images, 3 videos, and 3 audio clips with native audio and cinematic camera control.

stylizedtransformlipsync
text-to-image
ByteDanceREVIEW REQUIRED

Bytedance Seedream V4 Text To Image

fal-ai/bytedance/seedream/v4/text-to-image

A new-generation image creation model ByteDance, Seedream 4.0 integrates image generation and image editing capabilities into a single, unified architecture.

stylizedtransform
text-to-audio
ElevenLabsREVIEW REQUIRED

Elevenlabs Music v2.5

elevenlabs/music/v2.5

Generate high quality, realistic music with fine controls using Elevenlabs Music v2.5!

musictext-to-music
text-to-video
ByteDanceREVIEW REQUIRED

Seedance 2.5 Text to Video

bytedance/seedance-2.5/text-to-video

Dreamina Seedance 2.5 generates native 30-second single-shot video at up to 720p from a single text prompt, reasoning about the whole shot at once so motion, lighting, and subject identity stay coherent from first frame to last.

stylizedtransformlipsync
image-to-image
falREVIEW REQUIRED

Birefnet Background Removal

fal-ai/birefnet

bilateral reference framework (BiRefNet) for high-resolution dichotomous image segmentation (DIS)

background removalsegmentationhigh-resutility
image-to-video
GoogleREVIEW REQUIRED

Veo 3.1

fal-ai/veo3.1/image-to-video

Veo 3.1 is the latest state-of-the art video generation model from Google DeepMind

speech-to-text
falREVIEW REQUIRED

Wizper (Whisper v3 -- fal.ai edition)

fal-ai/wizper

[Experimental] Whisper v3 Large -- but optimized by our inference wizards. Same WER, double the performance!

transcriptionspeech
text-to-image
GoogleREVIEW REQUIRED

Nano Banana 2 Lite

google/nano-banana-2-lite

Nano banana lite is the efficiency-focused model in the image generation family. Sub-2 second latency with cost-effective generation and editing, fast multi-turn local edits, and 14 supported aspect ratios.

image-to-image
metaREVIEW REQUIRED

Meta Muse Image Edit

meta/muse-image/edit

Meta's Muse Image model does precise edits that change only what you ask, stay coherent across turns, and compose from multiple reference images.

realismtypographystylizedediting
text-to-video
MiniMaxREVIEW REQUIRED

H3 Max Turbo Text to Video

minimax/h3-max-turbo/text-to-video

fal's H3 Max Turbo is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

stylizedtransformlipsync
image-to-video
GoogleREVIEW REQUIRED

Veo3.1 Lite Image to Video

fal-ai/veo3.1/lite/image-to-video

Veo 3.1 Lite balances practical utility with professional capabilities, supporting Text-to-Video and Image-to-Video

stylizedtransformlipsync
image-to-video
AlibabaREVIEW REQUIRED

Wan 3.0

alibaba/wan-3.0/reference-to-video

Wan 3.0 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.

stylizedtransformlipsync
text-to-speech
MiniMaxREVIEW REQUIRED

MiniMax Speech-02 HD

fal-ai/minimax/speech-02-hd

Generate speech from text prompts and different voices using the MiniMax Speech-02 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.

speech
text-to-image
Black Forest LabsREVIEW REQUIRED

FLUX 2

fal-ai/flux-2

Text-to-image generation with FLUX.2 [dev] from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities.

image-to-image
GoogleREVIEW REQUIRED

Nano Banana Lite Edit

google/nano-banana-lite/edit

Nano banana lite is the efficiency-focused model in the image generation family. Sub-2 second latency with cost-effective generation and editing, fast multi-turn local edits, and 14 supported aspect ratios.

text-to-speech
MiniMaxREVIEW REQUIRED

MiniMax Speech 2.8 [HD]

fal-ai/minimax/speech-2.8-hd

Generate speech from text prompts and different voices using the MiniMax Speech-2.8 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.

image-to-video
ByteDanceREVIEW REQUIRED

Bytedance Seedance V1.5 Pro Image To Video

fal-ai/bytedance/seedance/v1.5/pro/image-to-video

Generate videos with audio with Seedance 1.5 (supports start & end frame)

bytedanceseedanceaudio
text-to-image
Black Forest LabsREVIEW REQUIRED

Flux 3 Image

blackforestlabs/flux-3/text-to-image

FLUX 3 Image is Black Forest Labs' newest image model. Generate detailed, well-composed images in native 2K and 4K with precise layout control and improved text rendering.

flux-3-imageblack-forest-labstypography4k
image-to-video
GoogleREVIEW REQUIRED

Gemini Omni Flash 1.1 Image to Video

google/gemini-omni-flash/v1.1/image-to-video

Gemini Omni Flash 1.1 is Google's multimodal video model. This endpoint animates a still image into video with synchronized audio, extending a single frame into coherent motion that reflects the logic of the real world.

stylizedtransformlipsync
video-to-video
KlingREVIEW REQUIRED

Kling O3 Edit Video [Pro]

fal-ai/kling-video/o3/pro/video-to-video/edit

Edit videos using Kling O3 from Kling Team!

video-to-video
text-to-image
IdeogramREVIEW REQUIRED

Ideogram V4.5 Text to Image

ideogram/v4.5

Generate high-quality images, posters, and logos with Ideogram 4.5, with accurate text rendering and low, medium, or high quality tiers.

realismtypographystylized
image-to-video
AlibabaREVIEW REQUIRED

Wan 3.0

alibaba/wan-3.0/image-to-video

Wan 3.0 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.

stylizedtransformlipsync
video-to-video
AlibabaREVIEW REQUIRED

Wan 3.0 Prime

alibaba/wan-3.0-prime/reference-to-video

Wan 3.0 Prime Reference-to-Video combines reference images, videos, and audio into a unified video with fast generation and strong multimodal coherence. It follows character identity, visual style, movement, and sound cues across references to create controlled, consistent, and production-ready results.

referencevideo
video-to-video
falREVIEW REQUIRED

sync-3 Lipsync

fal-ai/sync-lipsync/v3

sync-3 most powerful lipsync model yet, featuring native visual intelligence for professional-quality video.

stylizedtransformlipsync
image-to-image
falREVIEW REQUIRED

Clarity Upscaler

fal-ai/clarity-upscaler

Clarity upscaler for upscaling images with high very fidelity.

upscaling
image-to-video
ByteDanceREVIEW REQUIRED

Bytedance Omnihuman V1.5

fal-ai/bytedance/omnihuman/v1.5

Omnihuman v1.5 is a new and improved version of Omnihuman. It generates video using an image of a human figure paired with an audio file. It produces vivid, high-quality videos where the character’s emotions and movements maintain a strong correlation with the audio.

image-to-videolipsync