Up to 30 seconds, one take
4, 6, 8, 10, 15, 20 or 30 seconds. No stitching, no drift between segments. The next longest model here reaches fifteen; Google's stop at eight or ten.
Model · Alibaba Wan 3.0
Alibaba's video model, and the one thing it does that Google's models here don't: length. A single generation gives you 4 to 30 seconds with sound already in it, at 480p, 720p or 1080p, horizontal or vertical. You pay by the second — $0.05, $0.10 or $0.20 depending on resolution — so a short draft costs $0.20 and the longest, sharpest clip on the platform costs $6.00. Crypto, no subscription.
Wan 3.0 is Alibaba's video generation model, released on Alibaba Cloud Model Studio in August 2026, and it renders clips of 4 to 30 seconds with audio generated in the same pass as the picture. Here it reads a prompt, an optional starting frame and closing frame, or up to 10 reference images, and returns one finished clip at 480p, 720p or 1080p. We charge exactly what Alibaba lists for the model — $0.05, $0.10 and $0.20 per second by resolution — rather than marking it up, which is the same policy we apply to Google's models.
Anyone who is tired of stitching. Most video models cap out around eight or ten seconds, so a half-minute sequence means three or four generations, matching them by hand, and living with the seams where the model changed its mind about the lighting. Wan renders the whole thing as one take. For product loops, ambient backgrounds, process shots and anything where the camera just keeps moving, that removes the editing step entirely. Vertical works the same way — 9:16 at any of the three resolutions, which is what Shorts and Reels want.
4, 6, 8, 10, 15, 20 or 30 seconds. No stitching, no drift between segments. The next longest model here reaches fifteen; Google's stop at eight or ten.
Audio is generated with the picture and is on by default. Switching it off is a toggle, and the price stays the same either way — unlike Veo, where native audio is a paid extra.
Give it a starting image to animate, and optionally a closing image to land on. Same picture in both slots gives you a clean loop.
References keep a subject or a style consistent across runs. They are an alternative to frames, not an addition — the model takes one or the other.
Fix the seed and the same prompt lands close to the same shot. Useful when you are tuning one word at a time and want to see what that word did.
480p for drafts at $0.05 a second, 720p for publishing at $0.10, 1080p at $0.20. Draft cheap, render the keeper once.
There is no conversational editing here. With Gemini Omni Flash you can say "make the light warmer" and the model regenerates the same scene; on our side Wan has no memory of the clip it just made, so a change means a new prompt and a new charge. It also renders one clip per run — no four-variant grids to pick from. There is no negative prompt, so anything you do not want has to be written out of the scene rather than banned from it. Aspect ratios are 16:9 and 9:16 only, and 1080p is the ceiling. And starting frames and reference images are mutually exclusive: pick frames or pick references, not both.
Alibaba is candid about where the model is still growing: its own launch notes say audio texture and on-screen text rendering are not yet where they want them, and that matches what we see — the sound is convincing as ambience, less so as a finished mix. Long takes carry their own risk. Our twenty-second pottery clip above held together better than I expected, but thirty seconds is a lot of frames to keep consistent, so budget a retry for anything past fifteen. Two more things the model can do upstream that we have not wired in: it accepts documents (PPT, PDF, spreadsheets) as input, and it can extend or edit a finished generation. If you want audio you can steer sentence by sentence, Gemini Omni Flash is the better tool; if you need 4K or Google's motion quality, Veo 3.1 is still the top of the range.
Subject, what it is doing, the camera, the light. "Single continuous shot" is worth saying out loud. Then add a short "Audio/sound:" line for what should be heard — Wan takes the hint.
Four seconds at 480p costs $0.20. Get the composition right there, then re-run the winner at 720p or 1080p for the length you actually need.
Your balance is charged when the job starts, and refunded automatically if it fails or the filter rejects it. Top up with crypto; there is no monthly fee waiting for you.
| Модель | 4 s | 10 s | 30 s | Sound |
|---|---|---|---|---|
| Wan 3.0 · 480p$0.05/s — drafts and social | $0.20 | $0.50 | $1.50 | Included, free |
| Wan 3.0 · 720p$0.10/s — the usual choice | $0.40 | $1.00 | $3.00 | Included, free |
| Wan 3.0 · 1080p$0.20/s — the sharp one | $0.80 | $2.00 | $6.00 | Included, free |
| Veo 3.1 Fast · 720pGoogle's model, 4–8 s only | $0.35 | — | — | Costs extra |
| Gemini Omni Flash · 720pGoogle's model, 4–10 s only | — | $1.00 | — | Always on |
Prices are per finished clip and include the audio track. A dash means the model cannot render that length at all. Deposits of $50 add 5% to your balance and $100 adds 10%; an active promo code adds another 10% on top of the same deposit.
The full table for every model lives on the pricing page. If you want video cheaper still and can live with six seconds, look at MiniMax H3.
How starting frames, closing frames and loops actually behave, with the prompts that made them.
Читать гайдCamera language, lighting, motion. Written for Veo, but the structure carries over to Wan almost unchanged.
Читать гайдThe model to reach for when you need to edit a clip by talking to it instead of re-prompting.
Читать гайдUp to 30 seconds in a single generation. The lengths you can pick are 4, 6, 8, 10, 15, 20 and 30 seconds, and the clip comes back as one continuous take — it is not four short clips joined together. The next longest option on the platform is MiniMax H3 at fifteen seconds, and Google's Veo and Omni models stop at eight and ten.
You pay by the second: $0.05 at 480p, $0.10 at 720p, $0.20 at 1080p. So four seconds of draft is $0.20, ten seconds at 720p is $1.00, and the maximum — thirty seconds at 1080p — is $6.00. The audio track is included in that price.
Читать полный ответYes, and it is on by default. Sound is rendered together with the picture rather than added afterwards, so footsteps land on the footfall. You can switch it off for a silent clip, and that does not change what you pay. Describe what you want to hear in a short "Audio/sound:" line at the end of the prompt. Worth knowing: Alibaba's own launch notes list audio texture as an area still being refined, so treat the track as usable ambience rather than a finished mix.
Yes. Drop an image into the starting frame slot and the prompt describes what should happen to it. You can add a closing frame to control where the shot lands, and using the same picture in both slots gives you a loop. One catch: frames and reference images cannot be used together.
Veo still wins on motion quality and it is the only one here that reaches 4K, but it stops at eight seconds and native audio costs extra. Wan gives you length, cheaper seconds and free sound. My rule of thumb: hero shot for a client, Veo. Everything that needs to be long, or needs to be cheap, or needs twenty tries — Wan.
Not on Wan. It has no conversational editing and no scene extension, so a change means writing a new prompt and paying again. Gemini Omni Flash is the model that remembers the clip it made and lets you adjust it by describing the change.
Yes. What you generate is yours to use, including in client and commercial work, subject to the content rules in our terms. There is no watermark on the output and no separate licence to buy.
New accounts start with $0.20 on the balance — enough for a 480p draft to see whether Wan is your model. Clips from $0.20, no card, no subscription.