text-to-imageFLUX.2 [klein] 4B
fal-ai/flux-2/klein/4bText-to-image generation with FLUX.2 [klein] 4B from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
text-to-imagefal-ai/flux-2/klein/4bText-to-image generation with FLUX.2 [klein] 4B from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities.
text-to-videogoogle/gemini-omni-flash/v1.1/text-to-videoGemini Omni Flash 1.1 is Google's multimodal video model. This endpoint generates video with synchronized native audio from a text prompt, grounded in Gemini's real-world knowledge and physics understanding, with cinematic camera control expressed in natural language.
text-to-imagealibaba/qwen-image-3/text-to-imageGenerates images from a text prompt at resolutions up to 2048×2048, with automatic prompt rewriting and prompt-guided resolution selection, building on Qwen's strength in complex text rendering and precise prompt adherence
image-to-imageclarityai/crystal-upscalerAn advanced image enhancement tool designed specifically for facial details and portrait photography, utilizing Clarity AI's upscaling technology.
text-to-speechelevenlabs/tts/eleven-v4-turboGenerate speech with Eleven v4 Turbo from ElevenLabs. Choose a voice and control delivery with audio tags, stability, similarity settings, and IPA pronunciation.
text-to-imagefal-ai/krea-2/turboGenerate high-fidelity images from text in seconds with Krea 2 Turbo, the speed-optimized open-source version of Krea 2, preserving its aesthetic range for rapid ideation.
image-to-image
text-to-audiofal-ai/lyria3/proLyria 3 Pro is the latest music model from Google
image-to-imagefal-ai/sam2/imageSAM 2 is a model for segmenting images and videos in real-time.
text-to-videobytedance/seedance-2.0/fast/text-to-videoByteDance's most advanced text-to-video model, fast tier. Lower latency and cost with cinematic output, native audio, multi-shot editing, and director-level camera control.
image-to-imagefal-ai/flux-2/klein/4b/editImage-to-image editing with FLUX.2 [klein] 4B from Black Forest Labs. Precise modifications using natural language descriptions and hex color control.
video-to-videogoogle/gemini-omni-flash/v1.1/editGemini Omni Flash 1.1 is Google's multimodal video model. This endpoint edits video through natural-language instruction, applying the requested change while preserving the parts of the scene you want kept, and carrying character and scene consistency across successive edits.
image-to-videofal-ai/wan/v2.7/image-to-videoWan 2.7 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.
text-to-imageideogram/v4Generate high-quality images, posters, and logos with Ideogram's latest V4.0q — producing crisp visuals with accurate text rendering, fine detail, and full creative control for polished, ready-to-use designs.
text-to-videofal-ai/veo3.1/liteVeo 3.1 Lite balances practical utility with professional capabilities, supporting Text-to-Video and Image-to-Video
image-to-3dfal-ai/hyper3d/rodin/v2.5Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images.
text-to-audiofal-ai/minimax-music/v2Generate music from text prompts using the MiniMax Music 2.0 model, which leverages advanced AI techniques to create high-quality, diverse musical compositions.
text-to-imagefal-ai/qwen-imageQwen-Image is an image generation foundation model in the Qwen series that achieves significant advances in complex text rendering and precise image editing.
trainingfal-ai/flux-lora-fast-trainingTrain styles, people and other subjects at blazing speeds.
image-to-imagebytedance/seedream/v5/pro/layerizeSplits a finished image into independent, editable transparent-PNG layers — background plus separate elements, from a text description, returning 2 to 17 layers per call for non-destructive reuse in design tools.
text-to-imagebytedance/seedream/v5/flash/text-to-imageSeedream 5.0 Flash is a fast image generation and editing model, built for workflows where speed and budget matter.
text-to-imagefal-ai/gemini-3-pro-image-previewGemini 3 Pro Image (a.k.a Nano Banana Pro) is Google's state-of-the-art high-fidelity image generation and editing model
image-to-videobytedance/seedance-2.0/mini/reference-to-videoSeedance 2.0 Mini is a faster version of Seedance 2.0 that brings great performance and high generation speed at a lower cost.
image-to-imagexai/grok-imagine-image/quality/editGrok Imagine Pro is an advanced AI model from xAI that creates high-quality visuals from text prompts and allows you to edit or analyze existing images.
image-to-videobytedance/seedance-2.0/mini/image-to-videoSeedance 2.0 Mini is a faster version of Seedance 2.0 that brings great performance and high generation speed at a lower cost.
text-to-videoxai/grok-imagine-video/text-to-videoGenerate videos with audio from text using Grok Imagine Video.
image-to-3dmeshy/v7.1/image-to-3dMeshy 7.1 generates 3D models from a single image, with standard, low-poly, and Smart Topology modes, optional textures and PBR maps, and geometry resolution up to 4K.
visionfal-ai/imageutils/nsfwPredict the probability of an image being NSFW.