MODEL DIRECTORY
Find the right model
for every idea.
Explore the image and video models available in the Imagine Scenes
workspace.
IMAGE GENERATION MODELS
Image models
GPT Image 2.5 FlareA fast, high-fidelity option for prompt-led images and reference-driven edits, tuned for natural detail, reliable composition, and short creative iteration cycles.
GPT Image 2.5 SunburstAn expressive high-detail variant for images where richer lighting, texture, and visual impact matter more than the quickest possible turnaround.
GPT Image-2A dependable general-purpose image model for following detailed prompts, rendering clean text, and producing polished product, marketing, and editorial visuals.
Nano Banana 2Google's balanced all-round image model, combining fast generation with 4K output, reliable text rendering, multiple references, and consistent subjects across iterations.
Nano Banana ProGoogle's premium image model for complex instructions, precise creative control, advanced localization, brand consistency, and high-resolution professional asset production.
Nano Banana 2 LiteGoogle's fastest and most cost-efficient Nano Banana option, suited to rapid 1K concepts and high-volume edits rather than complex multi-reference workflows.
VIDEO GENERATION MODELS
Video models
Seedance 2.5Designed for longer scenes up to 30 seconds, combining text, images, and video references to give more control over motion, continuity, and scene direction.
Seedance 2.0A quality-focused video option for polished movement and reference-guided scenes, with more emphasis on visual control than the faster Seedance draft tier.
Seedance 2.0 FastA faster Seedance tier for exploring prompts, motion, and camera direction before committing to a more quality-focused or longer-form video generation.
Seedance 2.0 MiniA lower-cost Seedance option for quick drafts that still accepts rich references, helping you test concepts before investing in a premium render.
MiniMax H3A 2K-focused video model for detailed, cinematic clips with rich reference input, making it a strong fit when visual fidelity and guided motion matter.
Wan 3.0A flexible video model for text-led generation, frame and reference-guided creation, and video editing, with outputs up to 1080p and 30-second scenes.
Gemini Omni 1.1 FlashA fast, multi-reference Gemini video workflow with first and last frames plus editing support, useful for quickly shaping short scenes from supplied visual direction.
Gemini OmniA reference-led Gemini video option for guiding a scene with images and existing media while retaining flexible creation and editing inputs.
Veo 3.1 QualityGoogle's fidelity-first Veo tier for final-cut quality: cinematic text or image-to-video, first and last-frame control, reference images, and synchronized audio.
Veo 3.1 FastGoogle's faster Veo tier for standard production workflows, retaining cinematic controls and generated audio while prioritizing quicker creative feedback.
Veo 3.1 LiteGoogle's cost-conscious Veo tier for high-volume concepts, with short 16:9 or 9:16 clips and generated audio but fewer reference-image capabilities.
Kling 3.0A Kling video model for native-audio scenes and precise first/last-frame control, suited to directed shots where framing and transitions need close attention.
Kling 3.0 OmniA high-control Kling option supporting multiple image references, start and end frames, and video transformation, with output options reaching 4K.
Kling 3.0 TurboA faster Kling short-form option for prompt or first-frame video, useful when you need rapid motion studies at 720p or 1080p.