How to Use Wan 3.0: Alibaba’s AI Video Generator

Learn how to use Wan 3.0 to generate, edit, and extend AI videos with multimodal references and native audio.

How to Use Wan 3.0: Alibaba’s AI Video Generator

AI video generators have moved beyond producing isolated silent clips. Newer models can use several reference files at once, generate synchronized audio, and create longer sequences with multiple shots. Alibaba’s Wan 3.0 brings those capabilities into a single video model.

Wan 3.0 can generate videos from text, animate a first frame, interpolate between first and last frames, follow image, video, or audio references, edit existing footage, and extend a clip. It supports videos up to 30 seconds long at 1080p and 30 frames per second, with dialogue, background music, and sound effects generated alongside the visuals.

This guide explains what changed in Wan 3.0, how each generation workflow works, how to write a stronger Wan prompt, and how to turn the resulting clip into a finished video.

Table of Contents

What Is Wan 3.0?

Wan 3.0 is Alibaba Cloud’s all-in-one AI video generation model. Unlike Wan 2.1, which was primarily used for short text-to-video and image-to-video generations, Wan 3.0 combines several production tasks inside the same model.

This makes Wan 3.0 more useful for complete creative workflows than its predecessors. A creator can, for example, supply a product photo, a motion-reference video, and an audio track, then ask Wan to generate a vertical product ad that keeps the product recognizable while borrowing only the movement and pacing from the video. This makes it a strong choice for users looking to add motion to still images or enhance existing video content.

Wan 3.0 Strengths

  • Up to 30 seconds of native video generation
  • Integrated dialogue, music, ambience, and sound effects
  • Multiple reference types within one generation
  • First-frame and first-and-last-frame control
  • Prompt-based video editing and extension
  • 480p, 720p, and 1080p output
  • Both landscape and vertical aspect ratios

Wan 3.0 Limitations

  • Complex 30-second prompts can still lose visual or narrative consistency between shots
  • References can compete if their roles are not clearly assigned
  • Precise text, logos, hands, and fast object interactions may require several iterations
  • Availability and controls vary between Wan’s web interface, Alibaba Cloud regions, and third-party platforms
  • Generation time and credit cost increase with output length, resolution, and model tier

Wan 3.0 Video Generation Specs

Wan 3.0 represents a substantial upgrade from the six-second Wan 2.1 workflow described in earlier versions of this guide.

SpecificationWan 3.0 support
Duration2–30 seconds, plus a smart-duration option
Frame rate30 fps
Resolution480p, 720p, or 1080p
Aspect ratios16:9, 4:3, 1:1, 3:4, 9:16, or adaptive
Output formatMP4
AudioNative dialogue, background music, ambience, and sound effects; audio can be disabled
Reference imagesUp to 10, with a maximum size of 20 MB each
Reference videosUp to 5, totaling no more than 15 seconds; up to 100 MB each
Reference audio filesUp to 5, totaling no more than 15 seconds; up to 15 MB each
Other referencesOne supported document or one public webpage
Frame controlFirst frame or first and last frame
Video toolsPrompt-based editing and forward/backward extension

These specs are very creator-friendly, providing smooth, workable video that is especially helpful when merged together in larger projects.

How to Generate Videos With Wan 3.0

Wan AI is an open-source video generator that can be run locally on your device or accessed through public hosting sites. Public hosting typically has longer processing times, while private hosting requires knowledge of the model’s codebase and API tools.

Likewise, not all hosting sites offer the same features. Some provide more control over aspect ratio, resolution, and prompt type, such as text-to-video or image-to-video. Certain platforms also offer side-by-side access to multiple other AI video generators like Google's VEO 2 and Kling, allowing for a direct comparison between models.

To generate videos using Wan 2.1, visit the official International Experience Page and select AI Videos from the left-hand menu. This will open the video generation interface, where you can choose between text-to-video and image-to-video.

Guide showing how to access the Wan 2.1 image generator
Select the AI Videos option from the left-hand menu to access Wan 2

Use the Text2Video and Image2Video toggles at the top of the prompt menu to switch between these options.

Both methods generate AI-powered videos but function slightly differently.

Side-by-side images of the Wan AI text-to-video and image-to-video interfaces
Generate videos from a text or image prompt using Wan 2.1

Text-to-Video

  • Generate a video from scratch using a written prompt of up to 800 characters.
  • Select an aspect ratio from the available options (16:9, 9:16, 1:1, 4:3, 3:4)
  • Use optional tools:
    • Inspiration Mode: Adds more expressive video features.
    • Sound Effects: Generates sound effects if specified in the prompt. If no sounds are specified, background music is added.
    • Prompt Enhancing: Rewrites the prompt to improve clarity and optimize it for the Wan 2.1 model.

Image-to-Video

  • Upload one or two images to serve as the start and end frames. If only one image is uploaded, it will be used as the start frame.
  • Optionally include a supplemental prompt of up to 800 characters.
  • The generated video will match the aspect ratio of the reference image.
  • Use the same optional tools as text-to-video: inspiration mode, sound effects, and prompt enhancing.

Each video generation, whether from a text or image prompt, costs 10 credits and produces a six-second clip. These credits function as Wan AI’s proprietary currency, though more detailed pricing details will be covered later.

Wan AI is not the fastest video generator, but its speed is expected to improve over time. Enterprise users will likely have a different experience since they can run the model locally or on dedicated servers, avoiding the bandwidth limitations of public access.

Currently, image-to-video generations take approximately 15 minutes to generate, while text-to-video is slightly quicker at around 10 minutes. Generation times vary significantly depending on the complexity of the prompt, concurrent site usage, and optional tools.

Resizing Your Reference Image

Since Wan AI’s image-to-video generator matches the aspect ratio of the uploaded reference image, it is important to size your image correctly before uploading. A free online image resizer can help with this process.

Kapwing makes it easy to resize images with over 15 preset aspect ratios or custom dimensions in just a few clicks. To resize an image, start by uploading it to the editor. Once uploaded, select the background of the image before choosing the Resize Project tool on the right-hand side.

Guide showing how to access the Kapwing Resize Project tool
Use Resize Project tool to resize your image

Then, either enter custom image dimensions or choose from the available preset aspect ratios to resize your image.

Guide showing how to resize a reference image for Wan AI image to video generation
Select from over 15 size presets or enter custom image dimensions

To ensure your image looks exactly how you want it to after resizing, double-click on it to adjust its crop within the frame. This allows you to control the framing while maintaining the new aspect ratio.

Guide showing how to crop a reference image on Kapwing
Crop your image by double clicking it or by selecting the crop tool

Once completed, your image is ready to be exported and used as a reference for a Wan 2.1 video generation.

You can also use a video resizing tool to adjust the aspect ratio of any video generated by Wan AI if you need a different format than the default options: 16:9, 9:16, 1:1, 4:3, and 3:4.

How to Edit Wan 3.0 Videos

Generating short AI video clips can be a great way to fill brief gaps in your content with relevant visuals. Many content creators also combine generated clips with real footage for a more dynamic result.

To edit your Wan AI videos online, open the Kapwing editor in your browser. Start a new project or make small adjustments to existing clips. Use the media sidebar to add videos, text, graphics, images, audio, or AI-powered features like automatic subtitles.

Easily arrange layers in the project timeline by dragging and dropping them into place.

Guide showing the Kapwing online video editing interface
Make quick edits or build a video from scratch using the Kapwing online editor

To streamline your editing process, here are a few key shortcuts:

  • Spacebar: Play/Pause the video
  • Ctrl/Cmd + A: Select all layers in the scene
  • Ctrl/Cmd + Z: Undo | Ctrl/Cmd + Shift + Z (or Y on Windows) – Redo
  • S: Split selected layers at the playhead
  • H: Hide selected layers
  • Backspace/Delete: Remove selected layer or gap
  • Arrow Keys: Move layers (Up/Down to switch tracks, Left/Right to adjust position)
  • Ctrl/Cmd + ] or [:Bring layer forward/backward in order

Beyond filling gaps in longer videos, many creators use AI-generated clips to create short promotional content like trailers or teasers. These are especially useful for podcasters, YouTubers, and social media managers looking to build anticipation for an upcoming release without revealing actual footage.

To create a video in this style using Kapwing, upload and arrange your clips in the editor. Then, add automatic narration by selecting the AI Voice tool from the left-hand sidebar.

Guide showing how to generate a voice over on Kapwing
Open the AI Voice tool to generate a voice over

Enter or paste your script into the prompt box (up to 5,000 characters at a time). Once you've confirmed your script, selected a voice, and optionally chosen a persona, click Add Layer to generate the voice over.

Kapwing will automatically add subtitles that sync with the spoken content, giving your video a polished finish. To adjust the subtitles, use the editing menu on the right to change the font, size, color, border, and animations. If you need to modify the narration, reopen the AI Voice tool and update your script.

Guide showing how to edit automatically generated captions
Adjust the style of your captions by using the editing menu on the right-hand side

When you're finished, export your video by clicking Export in the top-right corner. For best compatibility, export as an MP4 and use the compression slider if you need to reduce file size.

Guide showing how to export a video for social media compatibility.
Export your project as an MP4 for best compatibility across platforms

The final video will be an effective promotional piece that integrates Wan AI-generated content into your production.

0:00
/0:15

Example video trailer using a Wan AI generated video clip

Wan 3.0 Pricing

For enterprise users, Wan AI offers a customized pricing structure based on usage and support needs. To learn more about enterprise pricing, visit the official Wan AI enterprise website.

For individual users, Wan AI operates on a credit system. New users receive 50 credits upon creating an account, which is enough for five video generations. Each video, whether created from a text or image prompt, costs 10 credits.

Image showing the Wan AI credits menu
Check your credit balance by selecting Credits from the generation screen

Currently, there is no option to purchase credits, but users can earn free credits through platform interactions:

  • Daily Check-In: Clicking the Check In button once per day adds 50 credits.
  • Publishing Videos: Sharing a generated video earns 20 credits. To publish, select the three-dot icon in the video window and choose Publish from the menu.
  • Providing Feedback: Rating a generated video with a thumbs-up or thumbs-down gives 5 credits, up to 10 credits per day.
Graphic showing how to earn Wan AI credits for free
Earn free credits by rating your video generations or publishing them to the Wan AI site

At the moment, credits cannot be purchased, but a paid option may be introduced later to increase generation limits.

Frequently Asked Questions

Is Wan 3.0 free?

Wan may offer limited trial credits or free generations, but availability varies by account, region, and current promotion. Ongoing use generally requires a consumer subscription, credits, or paid Alibaba Cloud API usage.

How long can Wan 3.0 videos be?

Wan 3.0 supports native generations from 2 to 30 seconds. When a reference video is included, the input and output duration together cannot exceed 30 seconds in the Alibaba Cloud API.

Does Wan 3.0 generate audio?

Yes. Wan 3.0 can generate synchronized dialogue, background music, ambience, and sound effects with the video. Audio is enabled by default in the API but can be disabled.

Can Wan 3.0 use multiple reference images?

Yes. Wan 3.0 supports up to 10 reference images, as well as video and audio references. Alibaba describes the model as supporting up to 20 multimodal reference materials in total, subject to the limit for each file type.

Can Wan 3.0 edit an existing video?

Yes. You can instruct Wan 3.0 to add, remove, replace, or restyle elements; change lighting or dialogue; and extend footage forward or backward. For exact trimming, captions, layout, and multi-clip assembly, finish the result in a timeline-based video editor.

What is the best aspect ratio for Wan 3.0?

Use 9:16 for TikTok, Instagram Reels, and YouTube Shorts; 16:9 for YouTube and widescreen video; 1:1 for square social posts; or adaptive when you want Wan to preserve the aspect ratio of an input image or video.