text-to-imagePhota Text to Image
fal-ai/photaPhota's model empowers developers, photographers, and creators with personalized photograph generation and editing.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
text-to-imagefal-ai/photaPhota's model empowers developers, photographers, and creators with personalized photograph generation and editing.
llmfal-ai/bytedance/seed/v2/miniSeed 2.0 Mini is a high-performance multimodal model optimized for low latency and high concurrency. It supports text, image, and video input with 256K context and configurable thinking/reasoning modes.
text-to-imagerundiffusion-fal/juggernaut-flux/lightningJuggernaut Lightning Flux by RunDiffusion provides blazing-fast, high-quality images rendered at five times the speed of Flux. Perfect for mood boards and mass ideation, this model excels in both realism and prompt adherence.
video-to-videofal-ai/ltx-2.3-quality/render-to-realTransform your 3D video render into realistic using first frame with Ltx 2.3
image-to-imagefal-ai/lora/image-to-imageRun Any Stable Diffusion model with customizable LoRA weights.
text-to-imagefal-ai/wan-25-preview/text-to-imageWan 2.5 text-to-image model.
text-to-videofal-ai/kling-video/o3/4k/text-to-videoKling's Native 4K is a video generation model that directly outputs professional-grade 4K video in one step, eliminating the need for post-production upscaling
text-to-videofal-ai/wan/v2.2-5b/text-to-video/fast-wanWan 2.2's 5B FastVideo model produces up to 5 seconds of video 720p at 24FPS with fluid motion and powerful prompt understanding
image-to-imagefal-ai/bria/genfillBria GenFill enables high-quality object addition or visual transformation. Trained exclusively on licensed data for safe and risk-free commercial use. Access the model's source code and weights: https://bria.ai/contact-us
text-to-imagefal-ai/wan/v2.2-a14b/text-to-imageWan 2.2's 14B model generates high-resolution, photorealistic images with powerful prompt understanding and fine-grained visual detail
image-to-imagefal-ai/minimax/image-01/subject-referenceGenerate images from text and a reference image using MiniMax Image-01 for consistent character appearance.
trainingfal-ai/qwen-image-2512-trainerQwen Image 2512 LoRA training
video-to-videoclarityai/crystal-video-upscalerDo high precision video upscaling that respects the original video perfectly using Crystal Upscaler's new video upscaling method!
text-to-videofal-ai/ltx-video-13b-distilledGenerate videos from prompts using LTX Video-0.9.7 13B Distilled and custom LoRA
visionfal-ai/moondream3-preview/captionMoondream 3 is a vision language model that brings frontier-level visual reasoning with native object detection, pointing, and OCR capabilities to real-world applications requiring fast, inexpensive inference at scale.
image-to-imagefal-ai/firered-image-edit-v1.1FireRed Image Edit v1.1 is an updated version of FireRed Image Edit, with improved image editing capabilities.
video-to-videodecart/lucy2-vton/realtimeRealtime Try On experience with Decart Lucy 2.1 VTON
image-to-imagefal-ai/flux-control-lora-depth/image-to-imageFLUX Control LoRA Depth is a high-performance endpoint that uses a control image using a depth map to transfer structure to the generated image and another initial image to guide color.
image-to-imagefal-ai/qwen-image-max/editImage editing endpoint for Qwen-Image-Max. Qwen Image Max improves upon the Qwen Image Plus series by enhancing the realism and naturalness of images.
image-to-imagetopaz/denoise/imageProfessional photo denoising powered by Topaz Labs. Normal, Strong and Extreme presets clean noise at source resolution; Denoise Max adds generative detail recovery. Best for high-ISO and night photography.
image-to-videofal-ai/kling-video/o1/standard/reference-to-videoTransform images, elements, and text into consistent, high-quality video scenes, ensuring stable character identity, object details, and environments.
image-to-imagefal-ai/live-portrait/imageTransfer expression from a video to a portrait.
text-to-speechfal-ai/orpheus-ttsOrpheus TTS is a state-of-the-art, Llama-based Speech-LLM designed for high-quality, empathetic text-to-speech generation. This model has been finetuned to deliver human-level speech synthesis, achieving exceptional clarity, expressiveness, and real-time performances.
text-to-3dfal-ai/hyper3d/rodin/v2.5/text-to-3d/fastRodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images. Do fast prototyping using the fast model.
image-to-imagefal-ai/qwen-image-edit-2509Endpoint for Qwen's Image Editing Plus model also known as Qwen-Image-Edit-2509. Has superior text editing capabilities and multi-image support.
video-to-videofal-ai/ltx-2.3-quality/reference-video-to-videoGenerate high-quality video with audio from reference video, text and images using LTX-2.3
image-to-jsonbria/ad-delayerTurn any flat ad image into fully editable layers —background, product and logo cutouts, live text with typography, and vector shapes. Commercial-safe, structured JSON output
image-to-videofal-ai/longcat-video/distilled/image-to-video/720pGenerate long videos in 720p/30fps from images using LongCat Video Distilled