image-to-imageExplore models.Know what runs.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
All model endpoints
image-to-image
image-to-imageFLUX 2 Pro Edit
fal-ai/flux-2-pro/editText-to-image generation with FLUX.2 [pro] from Black Forest Labs. Optimized for maximum quality, exceptional photorealism and artistic images.
FLUX1.1 [pro]
fal-ai/flux-pro/v1.1FLUX1.1 [pro] is an enhanced version of FLUX.1 [pro], improved image generation capabilities, delivering superior composition, detail, and artistic fidelity compared to its predecessor.
text-to-imageSeedream 5.0 Pro Text to Image
bytedance/seedream/v5/pro/text-to-imageByteDance's Seedream 5.0 Pro is flagship text-to-image model, with deep-thinking prompt understanding, native text in 14 languages, and precise control over dense layouts and structured designs.
image-to-imageIdeogram V4.5 Edit
ideogram/v4.5/editEdit images with Ideogram 4.5: prompt edits with up to 4 reference images, optional mask, and a high-precision mode that keeps unchanged pixels intact.
image-to-videoKling Video v3 Image to Video [Standard]
fal-ai/kling-video/v3/standard/image-to-videoKling 3.0 Standard: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation, with custom element support.
image-to-imageBria RMBG 2.0
fal-ai/bria/background/removeBria RMBG 2.0 enables seamless removal of backgrounds from images, ideal for professional editing tasks. Trained exclusively on licensed data for safe and risk-free commercial use. Model weights for commercial use are available here: https://share-eu1.hsforms.com/2GLpEVQqJTI2Lj7AMYwgfIwf4e04
visionOpenRouter [Vision]
openrouter/router/visionRun any Vision Language Model with fal. Analyze and understand images using Claude (Anthropic), GPT-5 / GPT-4o (OpenAI), Gemini (Google), Grok (xAI), Llama (Meta), Qwen, Pixtral (Mistral), and more. Send one or multiple images for captioning, analysis, OCR, or visual Q&A. Powered by OpenRouter.
text-to-videoMiniMax H3 Max Text to Video
minimax/h3-max/text-to-videofal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
llmOpenRouter Chat Completions [OpenAI Compatible]
openrouter/router/openai/v1/chat/completionsOpenAI-compatible chat completions API. Drop-in replacement for the OpenAI API — use any OpenAI SDK or client to access Claude, Gemini, Grok, DeepSeek, Llama, Qwen, Mistral, and all OpenAI models (GPT-5, GPT-4o, o3) through fal. Powered by OpenRouter.
text-to-imageFLUX1.1 [pro] ultra
fal-ai/flux-pro/v1.1-ultraFLUX1.1 [pro] ultra is the newest version of FLUX1.1 [pro], maintaining professional-grade image quality while delivering up to 2K resolution with improved photo realism.
speech-to-textElevenLabs Speech to Text - Scribe V2
fal-ai/elevenlabs/speech-to-text/scribe-v2Use Scribe-V2 from ElevenLabs to do blazingly fast speech to text inferences!
image-to-image
image-to-imageSegment Anything Model 3
fal-ai/sam-3/imageSAM 3 is a unified foundation model for promptable segmentation in images and videos. It can detect, segment, and track objects using text or visual prompts such as points, boxes, and masks.
image-to-imageFlux 3 Image
blackforestlabs/flux-3/edit-imageFLUX 3 Image Edit from Black Forest Labs makes precise local edits without changing the rest of the image, and combines up to 10 references into one balanced, well-composed result.
llmOpenRouter
openrouter/routerRun any LLM with fal. Access Claude (Anthropic), ChatGPT / GPT-5 / GPT-4o (OpenAI), Gemini (Google), Grok (xAI), DeepSeek, Llama (Meta), Qwen (Alibaba), Mistral, and 200+ more models through a single API. Supports reasoning, structured output, and streaming. Powered by OpenRouter.
text-to-audioElevenLabs TTS Multilingual v2
fal-ai/elevenlabs/tts/multilingual-v2Generate multilingual text-to-speech audio using ElevenLabs TTS Multilingual v2.
text-to-imageZ Image Turbo
fal-ai/z-image/turboZ-Image Turbo is a super fast text-to-image model of 6B parameters developed by Tongyi-MAI.
image-to-videoSeedance 2 Image to Video
bytedance/seedance-2.0/image-to-videoByteDance's most advanced image-to-video model. Animate still images into cinematic video with synchronized audio, start and end frame control, and motion prompts.
image-to-videoKling Video v2.6 Image to Video
fal-ai/kling-video/v2.6/pro/image-to-videoKling 2.6 Pro: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation.
text-to-audioElevenlabs Music
fal-ai/elevenlabs/musicGenerate high quality, realistic music with fine controls using Elevenlabs Music!
text-to-speechEleven v4
elevenlabs/tts/eleven-v4Generate expressive speech with Eleven v4 from ElevenLabs. Control delivery with audio tags, voice stability, similarity settings, and IPA pronunciation.
image-to-videoVeo 3.1 Fast
fal-ai/veo3.1/fast/image-to-videoGenerate videos from your image prompts using Veo 3.1 fast.
image-to-imageBytedance Seedream V4 Edit
fal-ai/bytedance/seedream/v4/editA new-generation image creation model ByteDance, Seedream 4.0 integrates image generation and image editing capabilities into a single, unified architecture.
image-to-videoMiniMax H3 Reference to Video
minimax/h3/reference-to-videoMiniMax H3 is a frontier video model. This endpoint generates 2K video from multimodal references up to 9 images for subject and style, 3 video clips for motion, and 3 audio clips each cited in the prompt by order, keeping subjects consistent while following the referenced motion and audio.
text-to-imageBytedance Seedream V4.5 Text To Image
fal-ai/bytedance/seedream/v4.5/text-to-imageA new-generation image creation model ByteDance, Seedream 4.5 integrates image generation and image editing capabilities into a single, unified architecture.
text-to-imageFLUX.2 [klein] 9B
fal-ai/flux-2/klein/9bText-to-image generation with FLUX.2 [klein] 9B from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities.
image-to-imageGrok Imagine Image
xai/grok-imagine-image/editEdit images precisely with xAI's Grok Imagine model