text-to-audioMMAudio V2 Text to Audio
fal-ai/mmaudio-v2/text-to-audioMMAudio generates synchronized audio given text inputs. It can generate sounds described by a prompt.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
text-to-audiofal-ai/mmaudio-v2/text-to-audioMMAudio generates synchronized audio given text inputs. It can generate sounds described by a prompt.
video-to-videofal-ai/heygen/v3/lipsync/precisionReplace or dub audio on an existing video with high-accuracy avatar-inference lip-sync.
text-to-imagefal-ai/wan/v2.7/text-to-imageGenerate high-quality images from text prompts using the WAN 2.7 model with advanced prompt understanding and detailed output.
video-to-videofal-ai/infinitalkInfinitalk model generates a talking avatar video from an image and audio file. The avatar lip-syncs to the provided audio with natural facial expressions.
image-to-imagebria/increase-resolutionUpscale any image 2x or 4x, up to 8192×8192, with Bria Increase Resolution. Preserves the original content — no regeneration, no altered details. Commercial-safe
image-to-imagefal-ai/florence-2-large/ocr-with-regionFlorence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks
video-to-audiosonilo/v1.1/video-to-musicAnalyzes your video’s pacing, mood, and timing to generate a frame-synced, licensed, commercial-use-safe soundtrack in seconds.
llmopenrouter/router/openai/v1/responsesThe OpenRouter Responses API with fal, powered by OpenRouter, provides unified access to a wide range of large language models - including GPT, Claude, Gemini, and many others through a single API interface.
image-to-imagefal-ai/bria/product-shotPlace any product in any scenery with just a prompt or reference image while maintaining high integrity of the product. Trained exclusively on licensed data for safe and risk-free commercial use and optimized for eCommerce.
image-to-3dfal-ai/trellis/multiGenerate 3D models from multiple images using Trellis. A native 3D generative model enabling versatile and high-quality 3D asset creation.
image-to-imagefal-ai/image-editing/photo-restorationRestore and enhance old or damaged photos by removing imperfections, adding color while preserving the original character and details of the image.
audio-to-audiofal-ai/ffmpeg-api/merge-audiosMerge audios into a single audio using FFmpeg API!
image-to-imagefal-ai/flux-2/klein/9b/edit/loraImage-to-image editing with FLUX.2 [klein] 9B from Black Forest Labs and custom LoRA. Precise modifications using natural language descriptions and hex color control.
image-to-3dfal-ai/hyper3d/rodin/v2.5/fastRodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images. Do fast prototyping using the fast model.
audio-to-audiofal-ai/deepfilternet3Enhance speech audio by removing background noise and upsampling to 48KHz
image-to-imagefal-ai/sam2/auto-segmentSAM 2 is a model for segmenting images automatically. It can return individual masks or a single mask for the entire image.
image-to-videofal-ai/kling-video/o1/reference-to-videoTransform images, elements, and text into consistent, high-quality video scenes, ensuring stable character identity, object details, and environments.
image-to-imagefal-ai/imageutils/depthCreate depth maps using Midas depth estimation.
image-to-3dfal-ai/hunyuan-3d/v3.1/rapid/image-to-3dRapidly generate 3D models from images using Hunyuan 3D.
text-to-imagefal-ai/flux-pro/kontext/max/text-to-imageFLUX.1 Kontext [max] text-to-image is a new premium model brings maximum performance across all aspects – greatly improved prompt adherence.
image-to-videofal-ai/kling-video/o1/standard/image-to-videoGenerate a video by taking a start frame and an end frame, animating the transition between them while following text-driven style and scene guidance.
image-to-imagefal-ai/flux-general/inpaintingFLUX General Inpainting is a versatile endpoint that enables precise image editing and completion, supporting multiple AI extensions including LoRA, ControlNet, and IP-Adapter for enhanced control over inpainting results and sophisticated image modifications.
text-to-audio
image-to-videofal-ai/bytedance/omnihumanOmniHuman generates video using an image of a human figure paired with an audio file. It produces vivid, high-quality videos where the character’s emotions and movements maintain a strong correlation with the audio.
image-to-imagefal-ai/flux-2/klein/9b/base/editImage-to-image editing with Flux 2 [klein] 9B Base from Black Forest Labs. Precise modifications using natural language descriptions and hex color control.
image-to-imagebria/fibo-edit-1.5/virtual-try-onBria Virtual Try-On edits a person photo to show the subject wearing garments or accessories from one to three reference images, guided by optional text instructions. Built on FIBO-Edit-1.5, it supports multi-garment changes and preserves the source aspect ratio by default.
text-to-imagefal-ai/ideogram/v2Generate high-quality images, posters, and logos with Ideogram V2. Features exceptional typography handling and realistic outputs optimized for commercial and creative use.
text-to-videofal-ai/bytedance/seedance/v1/pro/text-to-videoSeedance 1.0 Pro, a high quality video generation model developed by Bytedance.