What Wan 2.2 text-to-video is
Wan 2.2 comes in three open models. Wan2.2-T2V-A14B is the text-to-video model: a mixture-of-experts diffusion transformer with 27 billion parameters, 14 billion active at each step, where a high-noise expert lays out the motion and a low-noise expert refines textures and light. Wan2.2-I2V-A14B animates an existing image, and Wan2.2-TI2V-5B is a smaller hybrid that does both and runs on a consumer card at 720p and 24 frames per second.
Compared with Wan 2.1, the 2.2 models were trained on far more video, with labels for lighting, composition, colour and camera movement. That is why the prompt can ask for "golden hour backlight" or "handheld camera" and get it, something that earlier open models mostly ignored.
How to write a video prompt
A video prompt describes a short scene, not a still. The reliable structure is subject, action, setting, camera, style: "A red fox walks through fresh snow in a pine forest, turns its head towards the camera, low angle tracking shot, soft morning light, cinematic". Each element answers a question the model would otherwise guess.
- Subject and action: one subject, one clear action with a verb. Five seconds is enough for a gesture, a walk, a wave, not for a story.
- Setting: the place and the time of day set the light.
- Camera: static, pan, tilt, tracking, push-in, drone shot. Name it, or the model picks.
- Style: cinematic, documentary, anime, claymation, 35 mm film. One style word is enough.
Write in plain sentences; keyword lists that work in Stable Diffusion give static, incoherent motion here. Use the negative prompt for the classic defects: blurry, flicker, distorted hands, extra fingers, watermark, subtitles.
Settings and generation time
The defaults are 81 frames at 16 fps, 480p or 720p, 20 to 30 steps and a guidance scale around 5. At 480p a clip takes one to two minutes on a data-centre GPU; 720p takes four times longer for the same number of frames. The seed works as in image generation: fix it to compare two prompts, change it to get another take of the same prompt. The aspect ratio is chosen before generation, usually 16:9 or 9:16; the model cannot change it afterwards.
Because each attempt is slow, test the prompt at 480p, keep the seed of the take you like, and only then re-run at 720p. Most bad results come from prompts with too many actions or contradictory camera instructions, not from the settings.
What text-to-video cannot do yet
Five seconds is the practical limit for a consistent clip; longer generations drift, and the subject changes face or clothes. Text inside the video is unreliable. Hands and fast, complex interactions between several people still fail often. Physics is plausible rather than exact: water, cloth and hair move well, but a ball can bounce wrong. There is no audio. For a precise look, image-to-video is easier: generate the first frame with Stable Diffusion, SDXL or FLUX, then animate it with Wan 2.2.
Running Wan 2.2 T2V locally
The weights are on Hugging Face and run in ComfyUI, Diffusers and the official repository. The 14B T2V model needs about 80 GB of VRAM in full precision, or a 24 GB card with fp8 checkpoints and offloading at several minutes per clip. The 5B model produces a 720p, five-second clip on an RTX 4090 in under ten minutes. Below 16 GB of VRAM, use a cloud GPU or the free demo on this site, which also accepts an image as the first frame.