Wan 3.0 leads text-to-video and it leads video editing. MiniMax H3 Max leads image-to-video, at the lowest measured price on the board. Seedance 2.5, five weeks after release, still has no independent score at all.
That is the state of play on 5 September 2026, and it is close to the opposite of what most of the comparisons currently circulating will tell you.
The verdict
The three-way comparisons doing the rounds describe the first week of August. In that world, Wan 3.0 was a listing with a price and no evidence, absent from every leaderboard. MiniMax H3 was the only model with independent scores, and it led the editing board. Both of those facts have expired, and the second one expired by being overtaken by the first.
If you are budgeting AI video capacity for a Q4 campaign on the strength of an August read, you are pricing the wrong models. The correct move now is not consolidation onto a single engine. It is a three-way split: Wan 3.0 for text-driven and edit-driven work, H3 Max for volume iteration, and Seedance 2.5 for any sequence where a face or a product has to survive a long take. Each of the three wins a different job outright, and none of them wins all three.
What changed in thirty days
Alibaba stopped being quiet. Wan 3.0 opened in public beta on 6 August 2026 and formally launched on 24 August, one day after Alibaba closed a HK$80 billion share placement, roughly $10.2 billion, earmarked for AI. The model then entered the Artificial Analysis Video Arena and took the top of two boards.
MiniMax shipped the weights it had promised. They landed on 3 August 2026, with more conditions attached than the announcement implied. Details below, because those conditions determine whether your company can use them at all.
And a variant almost nobody was tracking took a board of its own. H3 Max, an accelerated version of H3 listed on the arena as post-trained by fal, now leads image-to-video with audio while costing a fifth of what the first-placed text-to-video model costs per minute.
Seedance 2.5, meanwhile, did not move. It launched on 31 July 2026 with the loudest specification sheet of the three and it still does not appear on any Artificial Analysis board. The only ByteDance entries are Seedance 2.0 720p and Seedance 1.5 pro. That is not a criticism of the model. It is a statement about what you can and cannot verify before you spend.
The board, as it stands today
All standings below come from the Artificial Analysis Video Arena, which ranks models on blind human preference votes, checked 5 September 2026. Their price column is normalised: the cost of one minute of 1080p video on the model creator's own API at default settings.
| Model | Maker | Max single clip | Published list price | Arena standing today |
|---|---|---|---|---|
| Wan 3.0 | Alibaba Tongyi Lab | 30 seconds | $3 / $6 / $12 per minute at 480p / 720p / 1080p | 1st in text-to-video (1238), 1st in editing (1190), 5th in image-to-video (1175) |
| MiniMax H3 Max | MiniMax, post-trained by fal | 15 seconds | $2.40 per minute, normalised | 1st in image-to-video (1201), 3rd in text-to-video (1235) |
| MiniMax H3 | MiniMax | 15 seconds | $0.13 per second at 2K, about $0.08 at 768p | 4th in text-to-video (1227), 2nd in editing (1130), 3rd in image-to-video (1187) |
| Seedance 2.5 | ByteDance Seed | 30 seconds | Billed by token: $10.70 per million without video input | Not yet ranked |
For context on the rest of the text-to-video board with audio: Gemini Omni Flash is level with Wan 3.0 at 1238 on 18,064 votes, Dreamina Seedance 2.0 720p sits fifth at 1222 on 24,790 votes, Wan 2.7 at 1157, Kling 3.0 1080p Pro at 1108 and Veo 3.1 at 1092. On the editing board, Wan 3.0 leads at 1190, MiniMax H3 follows at 1130, Gemini Omni Flash at 1125, Wan 2.7 at 1077 and Seedance 2.0 720p at 1040.
What each model actually is
Wan 3.0
Alibaba Tongyi Lab's model, API identifier wan3.0-video. Thirty seconds in a single pass at 30fps, in 480p, 720p or 1080p, with audio generated by default and switchable off at no change in price. Alibaba's published list rate is $0.05, $0.10 and $0.20 per output second across the three tiers. A 30 percent discount ran from 24 August to 23 September 2026. It is not open weights; Alibaba's open line stops at Wan 2.2. Documented regions are Singapore, Beijing and US East, and the model, endpoint and key must all match the same region or the call fails.
The feature almost every comparison has missed is Omni-Reference. Alongside text, image, audio and video, Wan 3.0 now accepts documents and web pages: doc, xls, ppt, pdf, txt, key, pages, numbers and md, one file or link per request, up to 100MB and 50 pages. Hand it a deck and it builds a thirty-second film from what is inside. Alibaba names training video from manuals, product films from slide decks and video versions of quarterly reports as the target work. For anyone whose briefs arrive as a brand book and a client deck, that is a more consequential change than three Elo points.
MiniMax H3
Released 31 July 2026, also called Hailuo 3.0 or Hailuo 03. Same model. A 33-billion-parameter dense omni-modal system that reads text, images, video and audio as one context and returns four to fifteen seconds at 24fps, up to 2K, with native 32kHz stereo audio generated in the same pass. Dialogue, effects and room tone arrive with the picture, not after it. API identifier MiniMax-H3. List price $0.13 per second at 2K, with a 768p tier around $0.08 per second that has been intermittently gated. Reference images cost $0.04 each, first five free.
The open-weights story needs care. MiniMax published weights on 3 August 2026, but only H3-Base: the FL2VA and Ref2VA checkpoints, which generate at 768p locally. Context-IR and the 2K regeneration stage stay hosted and closed. The licence is the MiniMax Community License, not Apache or MIT. Commercial use is granted to organisations under $20 million in revenue, and the territory clause excludes the United States, the European Union, the United Kingdom and South Korea from local deployment. A studio registered in Sofia, Paris or London falls inside that exclusion; a studio registered in Riyadh or Dubai does not appear on the list. Those are the licence's own words, not a legal opinion. Have counsel read the actual text before you build anything on a self-hosted checkpoint.
Seedance 2.5
Announced at ByteDance's Volcano Engine FORCE conference on 23 June 2026 and released 31 July. Thirty seconds of native single-pass video plus multi-round extension, and the largest reference budget of the three by a wide margin: 50 assets in one input, split as 30 images, 10 videos and 10 audio. It also has the most production-literate feature in the set, timestamp-level regional editing, which lets you change one detail inside a take, a character's hair colour for instance, without regenerating the whole thing. The performance you liked survives the note. Anyone who has lost a good take to a small client revision will understand why that matters.
Billing is by token rather than by second, which makes cross-model comparison awkward. ByteDance's published rate is $10.70 per million tokens without video input and $6.40 per million with video input, and its own worked examples put a five-second 16:9 clip at $0.514 at 480p and $1.156 at 720p. International model identifier dreamina-seedance-2-5-260628; China-region identifier doubao-seedance-2.5.
On resolution, be sceptical of what you read. The first-party API documented two output tiers at launch, 480p and 720p. The widely repeated claim that Seedance 2.5 delivers native 4K belongs to a separate 4K upgrade granted to Seedance 2.0 at the same keynote, and has been attached to the wrong model across most of the coverage. Some listings added a 1080p tier encoded in H.265 in late August. Until ByteDance publishes a native specification for 2.5, treat any 4K badge as an upscale and check the delivered file rather than the label on the button.
What the leaderboard does not measure
Elo measures which of two short clips a stranger preferred. It does not measure whether a face holds for thirty seconds, whether a brand red comes back as the right red, whether Arabic type renders as language rather than as decoration, or whether a client signs off. Every one of those decides whether a spot ships.
Vote counts matter too. Wan 3.0's text-to-video position rests on 5,554 votes against Seedance 2.0's 24,790. The newer number is real, and it is less settled. Wan 3.0 and Gemini Omni Flash are level at 1238 and sit inside each other's confidence intervals, so first place there is a coin toss rather than a result.
And the boards disagree with each other. Wan 3.0 leads text-to-video and drops to fifth in image-to-video, below the model it beat on the other board. A single overall ranking of video models does not exist. Match the model to the shot.
How to split the work
Treatments and pitch films go to Wan 3.0. The document input maps onto how regional agency work actually arrives, as a brand book, a deck or a spec sheet, and thirty seconds in one pass covers a full broadcast spot with no seams to hide.
Boards and animatics go to H3 Max. Pre-visualisation is a volume problem: you burn a hundred variations to keep one. At the cheapest measured rate on the board, that stops being a budget conversation.
Anything with a recurring face, a hero product or a talent lock goes to Seedance 2.5. Fifty references and regional editing are aimed squarely at consistency drift, which is what actually kills campaign work in this region, where a client will approve a face in shot two and expect it unchanged in shot nine.
Edit and revision passes go back to Wan 3.0, which now leads that board outright.
The number to hold onto
Three Elo points separate first place from third on the text-to-video board. Nine dollars and sixty cents a minute separate their prices.
Wan 3.0 sits at 1238 and $12.00 per minute. H3 Max sits at 1235 and $2.40. On blind preference, that gap is close to invisible. On a campaign that renders four hundred minutes of drafts before anything is graded, it is the whole line item.
Pay the premium where the frame is final. Do not pay it while you are still deciding what the frame is.
This model is available to run at vusion.art.