How do you generate a video from text with Wan 2.2?

By Updated 4 min read

Wan 2.2 text-to-video generates a clip of about five seconds, at 480p or 720p, from a single sentence that describes the scene, the action and the camera. It is Alibaba's open-weights video model, released in 2025 under the Apache 2.0 licence, and one of the few open models whose motion and lighting compare with commercial services.

What Wan 2.2 text-to-video is

Wan 2.2 comes in three open models. Wan2.2-T2V-A14B is the text-to-video model: a mixture-of-experts diffusion transformer with 27 billion parameters, 14 billion active at each step, where a high-noise expert lays out the motion and a low-noise expert refines textures and light. Wan2.2-I2V-A14B animates an existing image, and Wan2.2-TI2V-5B is a smaller hybrid that does both and runs on a consumer card at 720p and 24 frames per second.

Compared with Wan 2.1, the 2.2 models were trained on far more video, with labels for lighting, composition, colour and camera movement. That is why the prompt can ask for "golden hour backlight" or "handheld camera" and get it, something that earlier open models mostly ignored.

How to write a video prompt

A video prompt describes a short scene, not a still. The reliable structure is subject, action, setting, camera, style: "A red fox walks through fresh snow in a pine forest, turns its head towards the camera, low angle tracking shot, soft morning light, cinematic". Each element answers a question the model would otherwise guess.

  • Subject and action: one subject, one clear action with a verb. Five seconds is enough for a gesture, a walk, a wave, not for a story.
  • Setting: the place and the time of day set the light.
  • Camera: static, pan, tilt, tracking, push-in, drone shot. Name it, or the model picks.
  • Style: cinematic, documentary, anime, claymation, 35 mm film. One style word is enough.

Write in plain sentences; keyword lists that work in Stable Diffusion give static, incoherent motion here. Use the negative prompt for the classic defects: blurry, flicker, distorted hands, extra fingers, watermark, subtitles.

Settings and generation time

The defaults are 81 frames at 16 fps, 480p or 720p, 20 to 30 steps and a guidance scale around 5. At 480p a clip takes one to two minutes on a data-centre GPU; 720p takes four times longer for the same number of frames. The seed works as in image generation: fix it to compare two prompts, change it to get another take of the same prompt. The aspect ratio is chosen before generation, usually 16:9 or 9:16; the model cannot change it afterwards.

Because each attempt is slow, test the prompt at 480p, keep the seed of the take you like, and only then re-run at 720p. Most bad results come from prompts with too many actions or contradictory camera instructions, not from the settings.

What text-to-video cannot do yet

Five seconds is the practical limit for a consistent clip; longer generations drift, and the subject changes face or clothes. Text inside the video is unreliable. Hands and fast, complex interactions between several people still fail often. Physics is plausible rather than exact: water, cloth and hair move well, but a ball can bounce wrong. There is no audio. For a precise look, image-to-video is easier: generate the first frame with Stable Diffusion, SDXL or FLUX, then animate it with Wan 2.2.

Running Wan 2.2 T2V locally

The weights are on Hugging Face and run in ComfyUI, Diffusers and the official repository. The 14B T2V model needs about 80 GB of VRAM in full precision, or a 24 GB card with fp8 checkpoints and offloading at several minutes per clip. The 5B model produces a 720p, five-second clip on an RTX 4090 in under ten minutes. Below 16 GB of VRAM, use a cloud GPU or the free demo on this site, which also accepts an image as the first frame.

Questions people also ask

Is Wan 2.2 better than Wan 2.1?

Yes on motion quality, prompt adherence and aesthetic control: the mixture-of-experts models and the richer training labels give smoother motion and reliable lighting and camera instructions. Wan 2.1 remains useful on small GPUs, where the 5B Wan 2.2 model is the natural replacement.

Can Wan 2.2 make a video from an image and a text?

Yes. The I2V model uses an image as the first frame plus a text describing the motion, and the 5B TI2V model accepts text alone or text with an image. This site's Wan 2.2 demo works from an image.

Can I use the videos commercially?

The models are released under the Apache 2.0 licence, which allows commercial use of the outputs. Check the terms of the hosting service you generate on.