text-to-imageFLUX 2 Lora
fal-ai/flux-2/loraText-to-image generation with LoRA support for FLUX.2 [dev] from Black Forest Labs. Custom style adaptation and fine-tuned model variations.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
text-to-imagefal-ai/flux-2/loraText-to-image generation with LoRA support for FLUX.2 [dev] from Black Forest Labs. Custom style adaptation and fine-tuned model variations.
video-to-videominimax/h3-max/extend-videoH3 Max Extend Video adds a text-guided continuation to an existing video. It supports prompt expansion, adjustable duration and aspect ratio, and output resolutions from 480p to 2K, returning either the full extended video or only the new footage.
audio-to-audiofal-ai/sam-audio/separateAudio separation with SAM Audio. Isolate any sound using natural language—professional-grade audio editing made simple for creators, researchers, and accessibility applications.
text-to-speechfal-ai/minimax/voice-designDesign a personalized voice from a text description, and generate speech from text prompts using the MiniMax model, which leverages advanced AI techniques to create high-quality text-to-speech.
text-to-imagexai/grok-imagine-image/quality/text-to-imageGrok Imagine Pro is an advanced AI model from xAI that creates high-quality visuals from text prompts and allows you to edit or analyze existing images.
image-to-imagefal-ai/qwen-image-layeredQwen-Image-Layered is a model capable of decomposing an image into multiple RGBA layers.
text-to-imagefal-ai/gemini-3.1-flash-image-previewGemini 3.1 Flash Image (a.k.a Nano Banana 2) is Google's new state-of-the-art fast image generation and editing model
visionfal-ai/video-understandingA video understanding model to analyze video content and answer questions about what's happening in the video based on user prompts.
text-to-audiofal-ai/ace-step/prompt-to-audioGenerate music from a simple prompt using ACE-Step
image-to-videofal-ai/pixverse/swapGenerate high quality video clips by swapping person, objects and background using Pixverse Swap.
text-to-speechfal-ai/minimax/speech-2.6-hdGenerate speech from text prompts and different voices using the MiniMax Speech-2.6 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.
video-to-videofal-ai/sync-lipsyncGenerate realistic lipsync animations from audio using advanced algorithms for high-quality synchronization.
text-to-audiocassetteai/sound-effects-generatorCreate stunningly realistic sound effects in seconds - CassetteAI's Sound Effects Model generates high-quality SFX up to 30 seconds long in just 1 second of processing time
image-to-videolightricks/ltx-2.5/image-to-video/fastLTX-2.5 is Lightricks' open-source audio-video model. This endpoint animates a still image into video with synchronized audio in a single pass, in a speed-optimized mode for quick iteration.
audio-to-audiofal-ai/qwen-3-tts/clone-voice/1.7bClone your voices using Qwen3-TTS Clone-Voice model with zero shot cloning capabilities and use it on text-to-speech models to create speeches of yours!
text-to-imagefal-ai/z-image/turbo/loraText-to-Image endpoint with LoRA support for Z-Image Turbo, a super fast text-to-image model of 6B parameters developed by Tongyi-MAI.
image-to-video
text-to-speechxai/tts/v1Generate speech with expressive and realistic voices from xAI
text-to-speechfal-ai/chatterbox/text-to-speechWhether you're working on memes, videos, games, or AI agents, Chatterbox brings your content to life. Use the first tts from resemble ai.
text-to-imagefal-ai/qwen-image-2512Qwen Image 2512 is an improved version of Qwen Image with better text rendering, finer natural textures, and more realistic human generation.
image-to-imagefal-ai/wan/v2.7/editTransform and edit existing images with text-guided instructions using the WAN 2.7 model for creative image manipulation.
3d-to-3dfal-ai/meshy/rigging/multi-animationMeshy auto-rigs a humanoid 3D model fitting a skeleton and binding the mesh, then applies several motion presets from its animation library
image-to-imagefal-ai/flux-pro/v1/eraseLatest object erasing model from Black Forest Labs. Remove undesired objects, texts from images.
speech-to-speechresemble-ai/chatterboxhd/speech-to-speechTransform voices using Resemble AI's Chatterbox. Convert audio to new voices or your own samples, with expressive results and built-in perceptual watermarking.
fal-ai/codeformerFix distorted or blurred photos of people with CodeFormer.
text-to-videofal-ai/kling-video/o3/pro/text-to-videoGenerate realistic videos using Kling O3 from Kling Team!
image-to-videoblackforestlabs/flux-3/first-last-frame-to-videoFLUX 3 is Black Forest Labs' frontier video model. This endpoint generates the video between a defined start and end frame, interpolating a smooth, coherent transition from the first image to the last.
video-to-videofal-ai/bytedance/dreamactor/v2Transfer motion from a video to characters in an image using Dreamactor v2. Great performance for non-human and multiple characters