audio-to-audioSam Audio
fal-ai/sam-audio/span-separateAudio separation with SAM Audio. Isolate any sound using natural language—professional-grade audio editing made simple for creators, researchers, and accessibility applications.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
audio-to-audiofal-ai/sam-audio/span-separateAudio separation with SAM Audio. Isolate any sound using natural language—professional-grade audio editing made simple for creators, researchers, and accessibility applications.
speech-to-speechxai/grok-voice/realtimeBuild real-time voice applications powered by Grok. Stream audio and text bidirectionally via WebSocket for voice assistants, phone agents, and interactive voice systems.
audio-to-videopixverse/music-video/vibemvPixVerse VibeMV generates music videos from audio, with optional character references and lyric subtitles. It supports visual style presets, custom style references, five aspect ratios, and output at 720p or 1080p.
text-to-imagebria/fibo-gen-1.5/text-to-imageText-to-image model with high-fidelity outputs, accurate typography, and style preset, strong in photorealism, textures, and beyond. JSON-structured prompts give enterprise and agentic workflows production-ready control. Trained on licensed data.
image-to-videofal-ai/ovi/image-to-videoOvi can generate videos with audio from image and text inputs.
text-to-image
text-to-image
image-to-imagepixelcut/product-photoPixelcut's Background Remover produces fast, high-quality cutouts built for e-commerce product imagery
visionperceptron/isaac-01Isaac-01 is a multimodal vision-language model from Perceptron for various vision language tasks.
image-to-videofal-ai/minimax/video-01-live/image-to-videoGenerate video clips from your images using MiniMax Video model
text-to-audiofal-ai/kokoro/frenchAn expressive and natural French text-to-speech model for both European and Canadian French.
trainingfal-ai/turbo-flux-trainerA blazing fast FLUX dev LoRA trainer for subjects and styles.
text-to-audiofal-ai/kokoro/brazilian-portugueseA natural and expressive Brazilian Portuguese text-to-speech model optimized for clarity and fluency.
text-to-imagefal-ai/glm-imageCreate high-quality images with accurate text rendering and rich knowledge details—supports editing, style transfer, and maintaining consistent characters across multiple images.
image-to-videofal-ai/pixverse/v4.5/effectsGenerate high quality video clips with different effects using PixVerse v4.5
image-to-imagefal-ai/flux-krea-lora/image-to-imageFLUX LoRA Image-to-Image is a high-performance endpoint that transforms existing images using FLUX models, leveraging LoRA adaptations to enable rapid and precise image style transfer, modifications, and artistic variations.
workflowfal-ai/workflow-utilities/pick-image-by-indexChoose the Nth image from an image URL list for workflows.
video-to-videofal-ai/infinitalk/video-to-videoInfinitalk model generates a talking avatar video from an image and audio file. The avatar lip-syncs to the provided audio with natural facial expressions.
text-to-videofal-ai/heygen/avatar5/digital-twinCreate natural HeyGen Avatar V digital twin videos from text or audio, with lip-sync, optional backgrounds, captions, and MP4/WebM output.
image-to-3dhitem3d/hi3d/image-to-3dGenerate 3D models from a single image with Hi3D.
text-to-imagebria/fibo/generateSOTA open-source text-to-image model delivering high-fidelity outputs with accurate typography. JSON-structured prompts provide production-ready controllability for enterprise and agentic workflows. Trained exclusively on licensed data.
video-to-videofal-ai/workflow-utilities/blend-videoFFMPEG Utility for Blending Videos
text-to-imagefal-ai/longcat-imageLongCat image is a 6B parameter model excelling at multilingual text rendering, photorealism and deployment efficiency.
video-to-videofal-ai/wan-22-vace-fun-a14b/inpaintingVACE Fun for Wan 2.2 A14B from Alibaba-PAI
image-to-imagefal-ai/longcat-image/editLongCat image Edit is a 6B parameter image editing model excelling at multilingual text rendering, photorealism and deployment efficiency.
text-to-3dfal-ai/meshy/v6-preview/text-to-3dMeshy-6-Preview is the latest model from Meshy. It generates realistic and production ready 3D models.
text-to-videofal-ai/minimax/video-01-liveGenerate video clips from your prompts using MiniMax model
image-to-imagefal-ai/image-editing/background-changeReplace your photo's background with any scene you desire, from beach sunsets to urban landscapes, with perfect lighting and shadows