Tell a story for up to 30 seconds
Choose any whole-second length from 2 to 30, or let smart duration pick a length from your prompt. A longer shot gives actions, camera moves, and dialogue room to finish.
Plan a Wan 3.0 video with a clear prompt, useful reference media, and the right output settings. Create a draft with the AI video models available in Nana Banana today.
No preview yet
Adjust the settings on the left and generate a video to get started.

Current models: Veo 3.1 Fast and Veo 3.1. Wan 3.0 will appear in the model picker after Nana Banana integration is complete.
Official Wan 3.0 examples
Alibaba published these videos as Wan 3.0 examples. Nana Banana serves web-optimized copies from its CDN; they are official examples, not results generated here.
Official sample · 720P · 30s
Two image references keep a tabby cat and a wolf consistent through fast action, changing camera angles, fur, clothing, and spatial movement.
Official sample · 720P · 24s
Three reference images guide two characters and a dojo, including character movement, Japanese dialogue, and synchronized sound.
Official sample · 720P · 25s
A brand document provides the story, ingredients, and setting. One creative direction turns that material into a complete café commercial.
Wan 3.0 capabilities
Wan 3.0 combines prompts, media references, documents, and native audio in one video workflow.
Choose any whole-second length from 2 to 30, or let smart duration pick a length from your prompt. A longer shot gives actions, camera moves, and dialogue room to finish.
Wan 3.0 can read text, images, video, audio, office files, and a public web page. A slide deck or report can become the starting point for a video.
Audio is on by default. The model can build voices, music, and scene sound together with the visuals, which helps speech and action feel connected.
You can guide a result with up to 10 images, 5 short videos, and 5 short audio clips. These references help hold onto a person, product, place, style, or voice.
Use a first frame to decide how a clip begins, or add both a first and last frame to guide where the motion should end.
Alibaba says Wan 3.0 can change visuals, plot, or dialogue in an existing result. That makes it easier to fix one part without rebuilding the full idea.
Official specifications
Use these published limits to choose the right duration, reference media, aspect ratio, and output quality.
Prompt workflow
Describe the scene in order, add only the references that matter, and choose settings for the final destination.
Say who or what appears, what happens, how the camera moves, what the light looks like, and what should be heard. Put events in time order.
Use an image for appearance, a short video for motion, or audio for a voice. Name each reference in the prompt so its job is easy to understand.
Choose a duration that fits the action, a horizontal or vertical ratio for the destination, and a lower resolution for drafts before a 1080P final.
Model access
Wan 3.0 is available on Alibaba Cloud Model Studio. It is not yet available in Nana Banana's model picker.
Wan 3.0 is available in Model Studio under the model ID wan3.0-video. Your API key, endpoint, and model must use the same region.
Wan 3.0 is not yet connected to Nana Banana. It will appear in the model picker after the connection, credit cost, and generation flow are ready and tested.
Use Veo 3.1 Fast or Veo 3.1 to test your prompt, framing, and timing with the generator above.
Wan 3.0 facts last checked: August 29, 2026
Official resources
Review the Alibaba Cloud model release, product guide, and API reference used for this page.
Answers about the model, access, pricing, inputs, and output.
Wan 3.0 is Alibaba's newest all-in-one AI video model. It makes videos from text or references and can generate up to 30 seconds of visuals and sound in one task.
Not yet. The generator currently uses Veo 3.1 Fast or Veo 3.1. Wan 3.0 will appear in the model picker after Nana Banana completes and tests the integration.
The official API is billed by output second. At the last review, Alibaba listed $0.05 per second for 480P, $0.10 for 720P, and $0.20 for 1080P. Nana Banana provides signup credits for the models currently available here.
Without a video input, choose any whole-second duration from 2 to 30. Smart duration can choose for you. When a video is used as input, input time plus output time must stay within 30 seconds.
Yes. Audio is enabled by default, so Wan 3.0 can create dialogue, voices, music, and scene sound with the picture. You can also turn audio off.
The API accepts text, images, short videos, short audio clips, a document, or a public web page. It also supports first-frame and first-plus-last-frame image control.
The official API lists 480P, 720P, and 1080P output, with horizontal, square, vertical, and adaptive aspect ratios.
Alibaba's official sources describe Wan 3.0 as a cloud model and API, not a downloadable open-weight release. Older Wan versions have open repositories, but that does not make Wan 3.0 open source.
Write a prompt, choose your settings, and create with the AI video models available in Nana Banana today.