text-to-audioMiniMax (Hailuo AI) Music v1.5
fal-ai/minimax-music/v1.5Generate music from text prompts using the MiniMax model, which leverages advanced AI techniques to create high-quality, diverse musical compositions.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
text-to-audiofal-ai/minimax-music/v1.5Generate music from text prompts using the MiniMax model, which leverages advanced AI techniques to create high-quality, diverse musical compositions.
image-to-imagefal-ai/glm-image/image-to-imageCreate high-quality images with accurate text rendering and rich knowledge details—supports editing, style transfer, and maintaining consistent characters across multiple images.
text-to-videofal-ai/pixverse/c1/text-to-videoGenerate film-grade videos from text prompts with native audio, up to 1080p and 15 seconds, using PixVerse C1.
image-to-imagefal-ai/flux-2/klein/4b/edit/loraImage-to-image editing with FLUX.2 [klein] 4B from Black Forest Labs and custom LoRA. Precise modifications using natural language descriptions and hex color control.
text-to-imagerecraft/v4/style/text-to-vectorGenerates vector images that hold a consistent style, from either a saved style ID or reference images attached directly.
image-to-image
text-to-imagefal-ai/recraft/v4.1/utility/text-to-imageRecraft V4.1 Utility is a faster, lighter variant of V4.1 made for high-volume creative workflows. Ideal for ideation, A/B exploration, and content pipelines, it keeps Recraft's design sensibility while optimizing for throughput and cost.
text-to-imagefal-ai/ernie-imageHigh-quality text-to-image model by Baidu. Supports English, Chinese, and Japanese prompts with built-in prompt expansion.
image-to-imagefal-ai/image-apps-v2/hair-changeChange hairstyles and hair colors in photos realistically.
image-to-imagefal-ai/image-editing/hair-changeExperiment with different hairstyles, from bald to any style you can imagine, while maintaining natural lighting and realistic results.
image-to-imagebria/fibo-edit/relightPrecise, controllable photo re-lighting with structured text inputs. Apply natural lighting styles, soften harsh shadows, and transform scene illumination - production-ready and trained exclusively on licensed data.
text-to-imagenvidia/cosmos-3-super/text-to-imageCosmos3 is a collection of Omnimodal world models capable of generating dynamic, high-quality video, image, audio, and action commands from combinations of text, image, video, and action trajectory inputs.
text-to-imagefal-ai/sana/sprintSana Sprint is a text-to-image model capable of generating 4K images with exceptional speed.
text-to-imagefal-ai/flux-control-lora-cannyFLUX Control LoRA Canny is a high-performance endpoint that uses a control image to transfer structure to the generated image, using a Canny edge map.
video-to-videofal-ai/ltx-2.3-quality/inpaintInpaint high-quality video using LTX-2.3
image-to-videofal-ai/ltx-2.3-quality/image-to-videoGenerate high-quality video with audio from images using LTX-2.3
image-to-videofal-ai/hunyuan-video-image-to-videoImage to Video for the high-quality Hunyuan Video I2V model.
video-to-videofal-ai/wan-vace-14bVACE is a video generation model that uses a source image, mask, and video to create prompted videos with controllable sources.
text-to-speechfal-ai/dia-ttsDia directly generates realistic dialogue from transcripts. Audio conditioning enables emotion control. Produces natural nonverbals like laughter and throat clearing.
text-to-imagefal-ai/luma-photonGenerate images from your prompts using Luma Photon. Photon is the most creative, personalizable, and intelligent visual models for creatives, bringing a step-function change in the cost of high-quality image generation.
visionfal-ai/florence-2-large/detailed-captionFlorence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks
text-to-imagefal-ai/playground-v25State-of-the-art open-source model in aesthetic quality
visionfal-ai/moondream2Moondream2 is a highly efficient open-source vision language model that combines powerful image understanding capabilities with a remarkably small footprint.
image-to-image
audio-to-textnvidia/nemotron-3-nano-omni/audioAudio reasoning variant of NVIDIA's Nemotron 3 Nano Omni. 30B A3B hybrid Transformer-Mamba MoE - accepts audio plus a prompt and returns text.
image-to-imagefal-ai/qwen-image-edit-plus-loraLoRA endpoint for the Qwen Image Edit Plus model.
video-to-video
image-to-videomirage-api/avatar-x/reference-to-videoThe Avatar X API offers access to Mirage's most advanced generation model yet, delivering industry-leading identity preservation and expressivity in AI video