video-to-videoV2.6
wan/v2.6/reference-to-video/flashWan 2.6 reference-to-video flash model.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
video-to-videowan/v2.6/reference-to-video/flashWan 2.6 reference-to-video flash model.
speech-to-textfal-ai/speech-to-text/turboLeverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.
image-to-imagehitem3d/hi3d/image-to-reliefGenerate a 3D relief depth map with Hi3D from a single image.
image-to-videofal-ai/pixverse/v5.6/image-to-videoUse the latest pixverse v5.6 model to turn your texts and images into amazing videos.
image-to-imagefal-ai/joyai-image-editAll-in-one image AI with JoyAI-Image. Understand, create, and edit images through natural language—the model's deep visual understanding powers more accurate generation and precise editing in a unified system.
video-to-videobria/video/background-removal/green-screen-despillRemove background from videos filmed using chromakey, with automatic green spill suppression for clean, professional edges.
image-to-imagefal-ai/z-image/turbo/inpaint/loraGenerate images from text, an image, a mask and custom LoRA using Z-Image Turbo, Tongyi-MAI's super-fast 6B model.
image-to-imagefal-ai/flux-control-lora-canny/image-to-imageFLUX Control LoRA Canny is a high-performance endpoint that uses a control image using a Canny edge map to transfer structure to the generated image and another initial image to guide color.
text-to-videofal-ai/ltxv-13b-098-distilledGenerate long videos from prompts using LTX Video-0.9.8 13B Distilled and custom LoRA
image-to-imagefal-ai/fast-sdxl/inpaintingRun SDXL at the speed of light
image-to-imagebria/fibo-edit/erase_by_textRemove unwanted objects from images with a text prompt - fast, precise editing that seamlessly blends results. Built for production scale and trained on licensed data for safe commercial use.
image-to-videofal-ai/vidu/start-end-to-videoVidu Start-End to Video generates smooth transition videos between specified start and end images.
image-to-imagefal-ai/z-image/turbo/controlnet/loraGenerate images from text and edge, depth or pose images using custom LoRA and Z-Image Turbo, Tongyi-MAI's super-fast 6B model.
text-to-audiofal-ai/stable-audio-3/small/music/base/text-to-audioStable Audio 3 Small Music Base is the foundational 459 million parameter checkpoint generating full music compositions up to 2 minutes from text prompts, intended as the unmodified base for fine-tuning.
text-to-imagefal-ai/flux/srpoFLUX.1 SRPO [dev] is a 12 billion parameter flow transformer that generates high-quality images from text with incredible aesthetics. It is suitable for personal and commercial use.
trainingminimax/h3/ref2va/trainerTrain a MiniMax H3 LoRA with reference conditioning, so different modalities animate into video with audio; captions optional.
image-to-imagefal-ai/flux/srpo/image-to-imageFLUX.1 SRPO [dev] is a 12 billion parameter flow transformer that generates high-quality images from text with incredible aesthetics. It is suitable for personal and commercial use.
text-to-imagefal-ai/lumina-image/v2Lumina-Image-2.0 is a 2 billion parameter flow-based diffusion transforer which features improved performance in image quality, typography, complex prompt understanding, and resource-efficiency.
audio-to-audiofal-ai/workflow-utilities/audio-compressorFFMPEG Utility for Audio Compression
image-to-imagefal-ai/qwen-image-edit-2509-lora-gallery/multiple-anglesPrecise camera position and angle control (rotation, zoom, vertical movement)
image-to-videofal-ai/pixverse/v4/image-to-videoGenerate high quality video clips from text and image prompts using PixVerse v4
image-to-videofal-ai/wan-pro/image-to-videoWan-2.1 Pro is a premium image-to-video model that generates high-quality 1080p videos at 30fps with up to 6 seconds duration, delivering exceptional visual quality and motion diversity from images
image-to-videofal-ai/vidu/q1/start-end-to-videoVidu Q1 Start-End to Video generates smooth transition 1080p videos between specified start and end images.
text-to-imagefal-ai/bria/text-to-image/baseBria's Text-to-Image model, trained exclusively on licensed data for safe and risk-free commercial use. Available also as source code and weights. For access to weights: https://bria.ai/contact-us
video-to-videofal-ai/bernini-r/reference-edit-videoEdit a video guided by reference images with Bernini-R, bringing an object, material, background, style, or weather from a reference image into your video.
text-to-speechfal-ai/minimax/preview/speech-2.5-turboGenerate fast speech from text prompts and different voices using the MiniMax Speech-02 Turbo model, which leverages advanced AI techniques to create high-quality text-to-speech.
text-to-imagefal-ai/cogview4Generate high quality images from text prompts using CogView4. Longer text prompts will result in better quality images.
image-to-videofal-ai/flashheadSoulX-FlashHead is a unified 1.3B-parameter framework designed for high-fidelity, infinite-length, and real-time streaming portrait video generation.