EXPLORE / MODEL DIRECTORY

Explore models.Know what runs.

Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.

3Kaista-ready endpointsPublic catalog connected
PUBLIC REFERENCE CATALOG

All model endpoints

Page 2 · 28 per page
image-to-image
falREVIEW REQUIRED

SeedVR2

fal-ai/seedvr/upscale/image

Use SeedVR2 to upscale your images

upscaleimage-to-image
image-to-image
Black Forest LabsREVIEW REQUIRED

FLUX 2 Pro Edit

fal-ai/flux-2-pro/edit

Text-to-image generation with FLUX.2 [pro] from Black Forest Labs. Optimized for maximum quality, exceptional photorealism and artistic images.

text-to-image
Black Forest LabsREVIEW REQUIRED

FLUX1.1 [pro]

fal-ai/flux-pro/v1.1

FLUX1.1 [pro] is an enhanced version of FLUX.1 [pro], improved image generation capabilities, delivering superior composition, detail, and artistic fidelity compared to its predecessor.

text-to-image
ByteDanceREVIEW REQUIRED

Seedream 5.0 Pro Text to Image

bytedance/seedream/v5/pro/text-to-image

ByteDance's Seedream 5.0 Pro is flagship text-to-image model, with deep-thinking prompt understanding, native text in 14 languages, and precise control over dense layouts and structured designs.

realismtypographystylized
image-to-image
IdeogramREVIEW REQUIRED

Ideogram V4.5 Edit

ideogram/v4.5/edit

Edit images with Ideogram 4.5: prompt edits with up to 4 reference images, optional mask, and a high-precision mode that keeps unchanged pixels intact.

realismtypographyediting
image-to-video
KlingREVIEW REQUIRED

Kling Video v3 Image to Video [Standard]

fal-ai/kling-video/v3/standard/image-to-video

Kling 3.0 Standard: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation, with custom element support.

image-to-video
image-to-image
falREVIEW REQUIRED

Bria RMBG 2.0

fal-ai/bria/background/remove

Bria RMBG 2.0 enables seamless removal of backgrounds from images, ideal for professional editing tasks. Trained exclusively on licensed data for safe and risk-free commercial use. Model weights for commercial use are available here: https://share-eu1.hsforms.com/2GLpEVQqJTI2Lj7AMYwgfIwf4e04

background removalimage segmentationhigh resolutionutility
vision
openrouterREVIEW REQUIRED

OpenRouter [Vision]

openrouter/router/vision

Run any Vision Language Model with fal. Analyze and understand images using Claude (Anthropic), GPT-5 / GPT-4o (OpenAI), Gemini (Google), Grok (xAI), Llama (Meta), Qwen, Pixtral (Mistral), and more. Send one or multiple images for captioning, analysis, OCR, or visual Q&A. Powered by OpenRouter.

text-to-video
MiniMaxREVIEW REQUIRED

MiniMax H3 Max Text to Video

minimax/h3-max/text-to-video

fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

stylizedtransformlipsync
llm
OpenAIREVIEW REQUIRED

OpenRouter Chat Completions [OpenAI Compatible]

openrouter/router/openai/v1/chat/completions

OpenAI-compatible chat completions API. Drop-in replacement for the OpenAI API — use any OpenAI SDK or client to access Claude, Gemini, Grok, DeepSeek, Llama, Qwen, Mistral, and all OpenAI models (GPT-5, GPT-4o, o3) through fal. Powered by OpenRouter.

text-to-image
Black Forest LabsREVIEW REQUIRED

FLUX1.1 [pro] ultra

fal-ai/flux-pro/v1.1-ultra

FLUX1.1 [pro] ultra is the newest version of FLUX1.1 [pro], maintaining professional-grade image quality while delivering up to 2K resolution with improved photo realism.

high-resrealism
speech-to-text
ElevenLabsREVIEW REQUIRED

ElevenLabs Speech to Text - Scribe V2

fal-ai/elevenlabs/speech-to-text/scribe-v2

Use Scribe-V2 from ElevenLabs to do blazingly fast speech to text inferences!

speech-to-text
image-to-image
falREVIEW REQUIRED

Segment Anything Model 3

fal-ai/sam-3/image

SAM 3 is a unified foundation model for promptable segmentation in images and videos. It can detect, segment, and track objects using text or visual prompts such as points, boxes, and masks.

segmentationmaskreal-time
image-to-image
Black Forest LabsREVIEW REQUIRED

Flux 3 Image

blackforestlabs/flux-3/edit-image

FLUX 3 Image Edit from Black Forest Labs makes precise local edits without changing the rest of the image, and combines up to 10 references into one balanced, well-composed result.

flux-3-imageblack-forest-labsmulti-reference
llm
openrouterREVIEW REQUIRED

OpenRouter

openrouter/router

Run any LLM with fal. Access Claude (Anthropic), ChatGPT / GPT-5 / GPT-4o (OpenAI), Gemini (Google), Grok (xAI), DeepSeek, Llama (Meta), Qwen (Alibaba), Mistral, and 200+ more models through a single API. Supports reasoning, structured output, and streaming. Powered by OpenRouter.

text-to-audio
ElevenLabsREVIEW REQUIRED

ElevenLabs TTS Multilingual v2

fal-ai/elevenlabs/tts/multilingual-v2

Generate multilingual text-to-speech audio using ElevenLabs TTS Multilingual v2.

audio
text-to-image
falREVIEW REQUIRED

Z Image Turbo

fal-ai/z-image/turbo

Z-Image Turbo is a super fast text-to-image model of 6B parameters developed by Tongyi-MAI.

turboz-imagefast
image-to-video
ByteDanceREVIEW REQUIRED

Seedance 2 Image to Video

bytedance/seedance-2.0/image-to-video

ByteDance's most advanced image-to-video model. Animate still images into cinematic video with synchronized audio, start and end frame control, and motion prompts.

stylizedtransformlipsync
image-to-video
KlingREVIEW REQUIRED

Kling Video v2.6 Image to Video

fal-ai/kling-video/v2.6/pro/image-to-video

Kling 2.6 Pro: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation.

text-to-audio
ElevenLabsREVIEW REQUIRED

Elevenlabs Music

fal-ai/elevenlabs/music

Generate high quality, realistic music with fine controls using Elevenlabs Music!

musictext-to-music
text-to-speech
ElevenLabsREVIEW REQUIRED

Eleven v4

elevenlabs/tts/eleven-v4

Generate expressive speech with Eleven v4 from ElevenLabs. Control delivery with audio tags, voice stability, similarity settings, and IPA pronunciation.

audiotext-to-speechvoiceover
image-to-video
GoogleREVIEW REQUIRED

Veo 3.1 Fast

fal-ai/veo3.1/fast/image-to-video

Generate videos from your image prompts using Veo 3.1 fast.

image-to-image
ByteDanceREVIEW REQUIRED

Bytedance Seedream V4 Edit

fal-ai/bytedance/seedream/v4/edit

A new-generation image creation model ByteDance, Seedream 4.0 integrates image generation and image editing capabilities into a single, unified architecture.

stylizedtransformediting
image-to-video
MiniMaxREVIEW REQUIRED

MiniMax H3 Reference to Video

minimax/h3/reference-to-video

MiniMax H3 is a frontier video model. This endpoint generates 2K video from multimodal references up to 9 images for subject and style, 3 video clips for motion, and 3 audio clips each cited in the prompt by order, keeping subjects consistent while following the referenced motion and audio.

stylizedtransformlipsync
text-to-image
ByteDanceREVIEW REQUIRED

Bytedance Seedream V4.5 Text To Image

fal-ai/bytedance/seedream/v4.5/text-to-image

A new-generation image creation model ByteDance, Seedream 4.5 integrates image generation and image editing capabilities into a single, unified architecture.

stylizedtransform
text-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.2 [klein] 9B

fal-ai/flux-2/klein/9b

Text-to-image generation with FLUX.2 [klein] 9B from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities.

image-to-image
xAIREVIEW REQUIRED

Grok Imagine Image

xai/grok-imagine-image/edit

Edit images precisely with xAI's Grok Imagine model

grokxaiimage-editing