text-to-audioMinimax Music 2.6
fal-ai/minimax-music/v2.6MiniMax Music 2.6 creates complete tracks with singing, backing music, and detailed arrangements from lyrics and a style description.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
text-to-audiofal-ai/minimax-music/v2.6MiniMax Music 2.6 creates complete tracks with singing, backing music, and detailed arrangements from lyrics and a style description.
text-to-image
text-to-videofal-ai/kling-video/v2.5-turbo/pro/text-to-videoKling 2.5 Turbo Pro: Top-tier text-to-video generation with unparalleled motion fluidity, cinematic visuals, and exceptional prompt precision.
image-to-videofal-ai/minimax/hailuo-02/standard/image-to-videoMiniMax Hailuo-02 Image To Video API (Standard, 768p, 512p): Advanced image-to-video generation model with 768p and 512p resolutions
image-to-videobytedance/seedance-2.5/us/image-to-videoUS-hosted ByteDance Seedance 2.5 animates still images with synchronized audio and optional end-frame control. Generate videos up to 30 seconds at 480p, 720p or 1080p.
image-to-imagefal-ai/flux/dev/image-to-imageFLUX.1 Image-to-Image is a high-performance endpoint for the FLUX.1 [dev] model that enables rapid transformation of existing images, delivering high-quality style transfers and image modifications with the core FLUX capabilities.
image-to-videofal-ai/veo3.1/first-last-frame-to-videoGenerate videos from a first and last framed using Google's Veo 3.1
jsonfal-ai/ffmpeg-api/metadataGet encoding metadata from video and audio files using FFmpeg API.
image-to-imagefal-ai/bria/eraserBria Eraser enables precise removal of unwanted objects from images while maintaining high-quality outputs. Trained exclusively on licensed data for safe and risk-free commercial use. Access the model's source code and weights: https://bria.ai/contact-us
text-to-audiominimax/music-3MiniMax Music 3 is a high-performance music generation model for creating complete songs up to five minutes long
video-to-videofal-ai/seedvr/upscale/videoUpscale your videos using SeedVR2 with temporal consistency!
image-to-videofal-ai/pixverse/v6/image-to-videoPixverse's latest V6 Model
image-to-imagefal-ai/flux-2/flash/editImage-to-image editing with FLUX.2 [dev] from Black Forest Labs. Precise modifications using natural language descriptions and hex color control—in a flash.
video-to-videofal-ai/ffmpeg-api/merge-audio-videoMerge videos with standalone audio files or audio from video files.
text-to-imagefal-ai/flux-1/schnellFastest inference in the world for the 12 billion parameter FLUX.1 [schnell] text-to-image model.
image-to-imagefal-ai/image-preprocessors/depth-anything/v2Depth Anything v2 preprocessor.
image-to-videofal-ai/kling-video/o3/standard/reference-to-videoTransform images, elements, and text into consistent, high-quality video scenes, ensuring stable character identity, object details, and environments.
text-to-audiogoogle/lyria-3.5Lyria 3.5 is Google DeepMind's latest music generation model, and you can generate almost any type of music with it
image-to-videogoogle/gemini-omni-flash/image-to-videoAnimates a still image into video with audio. Extends a single frame into coherent motion, grounded in Gemini's physical understanding of how scenes and subjects behave.
image-to-videofal-ai/wan/v2.2-a14b/image-to-video/turboWan-2.2 Turbo image-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts.
image-to-videofal-ai/veo3.1/fast/first-last-frame-to-videoGenerate videos from a first/last frame using Google's Veo 3.1 Fast
video-to-videofal-ai/mmaudio-v2MMAudio generates synchronized audio given video and/or text inputs. It can be combined with video models to get videos with audio.
image-to-3dfal-ai/trellisGenerate 3D models from your images using Trellis. A native 3D generative model enabling versatile and high-quality 3D asset creation.
image-to-imagefal-ai/gemini-3.1-flash-image-preview/editGemini 3.1 Flash Image (a.k.a. Nano Banana 2) is Google's new state-of-the-art fast image generation and editing model
video-to-videofal-ai/kling-video/o3/standard/video-to-video/editEdit videos using Kling O3 from Kling Team!
text-to-audiobytedance/seed-audio-1.0Seed Audio 1.0 is a new audio model from Bytedance that can generate high-quality, natural sounding audio using text, reference audios or an image.
video-to-videofal-ai/wan/v2.2-14b/animate/replaceWan-Animate Replace is a model that can integrate animated characters into reference videos, replacing the original character while preserving the scene’s lighting and color tone for seamless environmental integration.
text-to-imagefal-ai/flux-2-maxFLUX.2 [max] delivers state-of-the-art image generation and advanced image editing with exceptional realism, precision, and consistency.