text-to-imageFLUX.2 [klein] 9B Base LoRA
fal-ai/flux-2/klein/9b/base/loraText-to-image generation with LoRA support for FLUX.2 [klein] 9B Base from Black Forest Labs. Custom style adaptation and fine-tuned model variations.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
text-to-imagefal-ai/flux-2/klein/9b/base/loraText-to-image generation with LoRA support for FLUX.2 [klein] 9B Base from Black Forest Labs. Custom style adaptation and fine-tuned model variations.
text-to-imagefal-ai/hidream-i1-fullHiDream-I1 full is a new open-source image generative foundation model with 17B parameters that achieves state-of-the-art image generation quality within seconds.
text-to-videofal-ai/wan-t2vWan-2.1 is a text-to-video model that generates high-quality videos with high visual quality and motion diversity from text prompts
video-to-videofal-ai/heygen/v2/translate/precisionHeygen Translate Model with Extreme Precision
image-to-imagefal-ai/wan-25-preview/image-to-imageWan 2.5 image-to-image model.
image-to-imagefal-ai/flux-2-lora-gallery/apartment-stagingVirtually furnishes an empty apartment
text-to-imagefal-ai/stable-diffusion-v3-mediumStable Diffusion 3 Medium (Text to Image) is a Multimodal Diffusion Transformer (MMDiT) model that improves image quality, typography, prompt understanding, and efficiency.
image-to-videofal-ai/ltx-2.3-22b/distilled/image-to-videoGenerate video with audio from images using LTX-2.3 Distilled
image-to-imagebria/extract-objectBria Extract Object uses text prompts to isolate a selected object from an image and return it as an RGBA PNG with a transparent background. Ideal for product, ecommerce, advertising, and creative editing workflows. Bria's Extract Object API leads in product shot extraction, outperforming SAM 3.1 where it counts most for commercial use.
text-to-imagefal-ai/krea-2/turbo/styleGenerate high-fidelity images from text with Krea 2 using a style reference image. Apply a reference image to guide the visual style into new generations, with aspect ratio, creativity, and seed controls.
image-to-imagefal-ai/image-editing/object-removalRemove unwanted objects or people from your photos while seamlessly blending the background.
image-to-imagefal-ai/hunyuan_worldHunyuan World 1.0 turns a single image into a panorama or a 3D world. It creates realistic scenes from the image, allowing you to explore and view it from different angles.
smoretalk-ai/rembg-enhanceRembg-enhance is optimized for 2D vector images, 3D graphics, and photos by leveraging matting technology.
text-to-speechfal-ai/kling-video/v1/ttsGenerate speech from text prompts and different voices using the Kling TTS model, which leverages advanced AI techniques to create high-quality text-to-speech.
audio-to-videolightricks/ltx-2.5/audio-to-video/fastLTX-2.5 is Lightricks' open-source audio-video model. This endpoint generates video timed to a supplied audio clip in a speed-optimized mode — useful for music-driven content, dialogue-led shorts, and ads keyed to a track.
image-to-imagebria/upscale/creativeProfessional-grade creative upscaler that doubles resolution up to 10MP, regenerating sharper textures, refined details, and cleaner faces. Trained exclusively on licensed data for risk-free commercial use.
image-to-imagetopaz/sharpen/imageProfessional photo sharpening powered by Topaz Labs. Models tuned per blur type (lens, motion, portrait, wildlife), plus Super Focus for generative recovery of severely blurred shots. Best for out-of-focus and motion-blurred photos.
visionfal-ai/moondream-nextMoonDreamNext is a multimodal vision-language model for captioning, gaze detection, bbox detection, point detection, and more.
text-to-speechfal-ai/minimax/preview/speech-2.5-hdGenerate speech from text prompts and different voices using the MiniMax Speech-02 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.
image-to-videofal-ai/ltx-video-13b-distilled/image-to-videoGenerate videos from prompts and images using LTX Video-0.9.7 13B Distilled and custom LoRA
image-to-videofal-ai/wan-flf2vWan-2.1 flf2v generates dynamic videos by intelligently bridging a given first frame to a desired end frame through smooth, coherent motion sequences.
text-to-videofal-ai/minimax/video-01Generate video clips from your prompts using MiniMax model
text-to-speechfal-ai/mayaMaya1 is a state-of-the-art speech model by Maya Research for expressive voice generation, built to capture real human emotion and precise voice design.
image-to-imagefal-ai/image-apps-v2/age-modifyModify a face to look younger or older while keeping identity realistic.
image-to-imagefal-ai/qwen-image/image-to-imageQwen-Image (Image-to-Image) transforms and edits input images with high fidelity, enabling precise style transfer, enhancement, and creative modification.
audio-to-audiofal-ai/stable-audio-3/medium/audio-to-audioStable Audio 3 Medium audio-to-audio is a 1.4 billion parameter latent diffusion model that transforms an input audio clip into new stereo variations up to 6 minutes guided by a text prompt.
visionfal-ai/moondream3-preview/pointMoondream 3 is a vision language model that brings frontier-level visual reasoning with native object detection, pointing, and OCR capabilities to real-world applications requiring fast, inexpensive inference at scale.
text-to-imagefal-ai/recraft-20bRecraft 20b is a new and affordable text-to-image model.