Alibaba's Wan3.0 Debuts With 30-Second AI Video Generation

Alibaba's video generation model Wan3.0 entered open beta across the platform on August 24. A single generation run can now reach 30 seconds, and for the first time the model accepts office documents — doc, xls, ppt, pdf and md files — as input, turning existing materials directly into video. Users can try it on Alibaba Cloud Bailian, Wan Jing Yi Ke, the official Wanxiang website and the Qwen Chat desktop client, with a gradual rollout underway inside the Qwen app.

The jump in duration matters most at the editing bench. A clip under ten seconds barely holds one shot, leaving little room for continuous camera movement or long takes; stretching to 30 seconds finally gives a complete short scene somewhere to live. Consistency sits near the top of the capability list Alibaba published for this release — keeping characters, props, sound, spatial relationships and visual style stable across the full clip is exactly what longer duration makes hardest.

Pricing decides whether it enters production pipelines

API pricing runs by the second: 480P, 720P and 1080P cost 0.3, 0.6 and 1.2 yuan per second respectively. From August 24 to September 23, all three tiers are discounted 30%, dropping to 0.21, 0.42 and 0.84 yuan per second.

Converted into finished output: a 30-second 1080P video costs 36 yuan at standard price, 25.2 yuan during the promotion; the same length at 480P costs 9 yuan standard, 6.3 yuan promotional.

Scaling that up to a production-level estimate: a five-minute episode of an AI-generated short drama, shot entirely at 1080P with every clip usable on the first pass, costs roughly 360 yuan in pure generation cost. In practice, discarded takes are unavoidable; assuming a three-to-one yield, real cost lands just over 1,000 yuan. That is nowhere near live-action production costs, and it is also a different order of magnitude from earlier per-minute pricing schemes that ran to tens of dollars. For clients in short drama, advertising and tourism promotion who need large volumes of finished video, the cost has moved from “worth a try” to “something you can actually budget for.”

Spreading 30 seconds across shot count shows where the real difficulty sits. A 30-second finished clip cut at a normal pace typically runs four to six shots, and the model has to keep the same face, the same prop, the same setting consistent across every cut. A ten-second-class model can hide inconsistency behind its short runtime; thirty seconds removes that cover, and any lapse in consistency becomes immediately visible. Alibaba grouping characters, props, sound, spatial relationships and style together into one capability list is, broadly, a response to exactly this problem.

Wan3.0 is not the first to reach the 30-second mark for a single clip — ByteDance's Seedance 2.5 already extended duration to 30 seconds earlier. Several domestic players pushing the same metric to the same point at roughly the same time suggests long-shot coherence, rather than resolution, has become the main battleground of this round of competition.

Moving the entry point from a prompt to a slide deck

Support for doc, xls, ppt, pdf and md input is the detail most worth a closer look here. Making a product-introduction video used to require someone to read through the source material, break it into a shot list, then write it up as a prompt — that manual translation step often took longer than the generation itself. Feeding documents straight into the model removes that step entirely. The trade-off is less control: the model decides on its own which page becomes which shot and which paragraph becomes narration, which makes it harder to pinpoint what to fix on a redo — a workflow that works better on highly standardized source material.

The enterprise use cases line up neatly with this design: training decks, product manuals, financial reports and internal announcements already exist as PPT and PDF files. This path also explains why Wan3.0 was placed inside the Qwen Chat desktop client and Bailian in the first place — it is built for office workflows, not casual creation on social platforms.

On output quality, Alibaba highlights detail in rendering real people, with micro-expressions, skin and eye rendering singled out as areas of focus for this release, alongside adjustments to color grading and style blending. Industry feedback so far clusters around three words: stable, realistic, and polished. Users specifically point to better continuity in long-sequence narrative and fewer rounds of parameter tweaking needed to get there.

These impressions are still anecdotal, and independent benchmark comparisons have yet to catch up. The real answer will come over the month after the promotional pricing ends — how many commercial finished videos actually get produced with it.

Sources: Alibaba Cloud announcement, Quantum Bit (QbitAI), IT Home, CocoLoop; the 30-second single-clip length, doc/xls/ppt/pdf/md input support, the list of launch platforms, and the standard/promotional per-second pricing for the 480P/720P/1080P tiers (0.3/0.6/1.2 yuan) are all cross-checked against the above sources; the production-cost estimates are this site's own rough calculation.