Kuaishou (Kling AI)
Kling 3.0
Kuaishou's flagship Kling video model, with multi-shot storyboarding, synchronized audio-video generation, element references, and native 4K output.
- Text to video
- First frame to video
- First and last frame
Models
Modes, reference limits, durations and where to connect each model — read from the same profiles Nomi uses when it calls the model, so this page updates when the models do.
Kuaishou (Kling AI)
Kuaishou's flagship Kling video model, with multi-shot storyboarding, synchronized audio-video generation, element references, and native 4K output.
MiniMax
MiniMax's multimodal video model — text, images, video and audio can all be references, with native stereo sound and up to 2K output. Also known as Hailuo 3.
ByteDance Seed
ByteDance’s video model that generates multi-shot clips of up to 30 seconds with sound, taking up to 30 reference images, 10 videos and 10 audio clips at once.
Alibaba Cloud (Tongyi Wan)
Alibaba's Wan 3.0, with native 30-second duration, up to 20 reference assets, document and webpage parsing, and built-in audio.
Google DeepMind
Google's video-focused entry in its any-to-any multimodal family, built for conversational video creation and editing, and now the model that replaces Veo inside the Gemini app.
ShengShu Technology
ShengShu Technology's video model that generates up to 16 seconds of native audio-synced video, with reference-to-video to keep characters and scenes consistent.
Kuaishou (Kling AI)
The faster variant in Kuaishou's Kling 3.0 series, for text- and image-to-video generation with no watermark by default.
xAI
xAI's video model that turns text or a starting image into clips up to 15 seconds, now with native 1080p and up to seven image or voice references.
MiniMax and fal.ai
The speed-optimized variant of MiniMax H3, jointly released by MiniMax and fal.ai, for faster text-, image-, and reference-to-video generation.
Runway
Runway’s current flagship video model, with both text-to-video and image-to-video support and more realistic physics — liquids, collisions and momentum — than Gen-4.
Runway
Runway’s faster, lighter video model — image-to-video only, generating a 10-second clip in about 30 seconds from a source image and a motion prompt.
Alibaba Cloud (Bailian / Model Studio)
HappyHorse 1.1 from Alibaba Cloud Bailian, one model that auto-routes between text-to-video, image-to-video, and character reference.
Google DeepMind
Google DeepMind’s video model with native synchronized sound and dialogue, able to chain reference-image-driven clips into scenes over a minute long.
MiniMax
MiniMax's Hailuo 2.3, for text- and image-to-video with 15 built-in camera-move commands and automatic prompt optimization — the generation before H3.
Agnes AI
Agnes AI's video model with three modes behind one endpoint — text-only, first/last-frame keyframing, and multi-reference from images, audio or video — plus a faster 2.5 Flash variant.
Higgsfield AI
Higgsfield’s image-to-video model built around camera control — pick a preset and a single photo becomes a moving shot, with 50+ presets to choose from.
OpenAI
OpenAI's image model for generation and editing in one, with transparent backgrounds and resolution up to 4K.
Google DeepMind
Google's image model with Pro-level generation and editing at Flash speed, taking up to 14 reference images.
ByteDance Seed
ByteDance's flagship image model for dense layouts and pixel-level retouching, with click/lasso/recolor editing and native rendering in a dozen-plus languages.
OpenAI
OpenAI's fast tier of GPT Image 2.5, built for everyday generation, batch work and quick drafts.
OpenAI
OpenAI's most capable GPT Image 2.5 tier, built for editing precision on production and ad work.
Black Forest Labs
Black Forest Labs' flagship image model, fusing up to 10 reference images at up to 4MP output, with a strong edge in typography and brand-consistent visuals.
Google DeepMind
Google's Gemini 3 Pro Image, also known as Nano Banana Pro — its best model for complex, multi-turn editing, up to 4K.
xAI
xAI's image model that treats editing as a first-class feature — targeted region edits, background removal, up to 5 reference images, and a claimed
Alibaba Qwen
Alibaba Qwen's newest image model, built for usefulness over looks — it takes prompts up to roughly 4.5k tokens and renders text as small as 10px clearly.
Meta Superintelligence Labs
Meta's image model, served through Runway's API — it plans the full layout before drawing, blends up to 10 reference images, and revises its own output.
Runway
Runway's own reference-based image model — up to 3 reference images carry a character or location straight into a new shot, and it also works from text alone.
Google DeepMind
Google's fastest, most efficient Gemini image model, generating in as little as 4 seconds with up to 10 reference images.
ByteDance Seed
ByteDance's lighter image model that reasons before drawing and can search the web, generating up to 15 images per request for office and visualization work.
Alibaba Tongyi Lab
Alibaba Tongyi Lab's 6B open-weight speed model that generates in 8 steps on consumer GPUs, with accurate bilingual text and the top open-source score on Artificial Analysis.
Google DeepMind
Google's speed-focused Imagen 4 tier for pure text-to-image generation, built for rapid, high-volume work.
Google DeepMind
Google's highest-quality Imagen 4 tier, built for maximum detail and strict prompt adherence.
Agnes AI
Agnes AI's image model tuned for dense, complex compositions, with tiered resolution up to 4K and support for text-to-image, editing and multi-image composition.
Alibaba Tongyi Lab
The full 6B foundation model behind Z-Image-Turbo, trading generation speed for more diversity and fine-tunability — Tongyi Lab's base for LoRA and downstream work.
Alibaba Qwen
Alibaba Qwen's fully open-source image model line, first released August 2025 and updated continuously, known for strong Chinese text rendering with downloadable weights.
Black Forest Labs
Black Forest Labs' open-weight image model with 4-step distilled inference for sub-second generation; BFL says the 9B build matches or beats models five times its size.
Black Forest Labs and Krea
A photorealism-focused image model from Black Forest Labs and Krea that fights the oversaturated "AI look" and matched FLUX.1 Pro on human preference tests.
Higgsfield
Higgsfield's in-house fashion-portrait model with taste built into the model itself, 20+ curated presets, and a trainable, reusable character identity.
Higgsfield
Higgsfield's Soul-family model built specifically for film-still visuals — natural grain and mood lighting, often used as a video keyframe.
macOS · Windows · AGPL-3.0 · No account