I run this site's tools every day, but I'd never actually sat down and used Wan 2.2 the way a regular visitor would — pick a prompt, hit generate, see what comes out, try again. So one week I set myself a small challenge: make 10 different videos, log every prompt and every result honestly, including the ones that came out badly. This post is that log, cleaned up.
Quick context: I ran everything through the studio on this site (which sits on top of Wan 2.2), on the default settings unless I say otherwise. Nothing here is sponsored or cherry-picked — video #4 and #7 genuinely didn't work, and I've kept them in.
The Setup
Ten prompts, one per day, roughly 5 seconds each, no re-rolling more than 3 times per prompt (to keep it honest — endless re-rolling would make every tool look perfect eventually). I varied the subject matter on purpose: people, animals, objects, indoor and outdoor scenes, slow shots and fast motion.
| # | Prompt (Shortened) | Result | Attempts |
|---|---|---|---|
| 1 | Woman walking through a rainy street market, neon signs | ✅ Good on 1st try | 1 |
| 2 | Golden retriever running across a beach at sunset | ✅ Good, slight paw-blur on landing | 2 |
| 3 | Close-up of coffee being poured into a cup, steam rising | ✅ Excellent, best result of the week | 1 |
| 4 | Two people shaking hands in an office, wide shot | ❌ Fingers merged/warped | 3, kept the failure |
| 5 | Drone shot flying over a mountain range at dawn | ✅ Good, very usable | 1 |
| 6 | Cat knocking a glass off a table, slow motion | ✅ Good after adjusting motion wording | 2 |
| 7 | Crowded dance floor, fast cuts, multiple people dancing | ❌ Faces smeared with too much motion | 3, still imperfect |
| 8 | Single candle flame close-up, dark background | ✅ Excellent, very clean | 1 |
| 9 | Skateboarder doing a trick down stairs | ⚠️ Usable but board clips through foot once | 3 |
| 10 | Elderly man reading a newspaper by a window, soft light | ✅ Good, very natural lighting | 1 |
What Actually Failed, And Why
Two clips didn't work, and both failures followed the same pattern.
Video #4 (handshake) and #7 (dance floor) both involved multiple people interacting closely, with hands or faces overlapping. In both cases, the model handled a single subject fine but struggled the moment two sets of hands or faces needed to occupy the same space at the same time. Fingers fused together in the handshake clip, and faces blurred into each other on the dance floor.
This matches what I'd seen mentioned around: current text-to-video models are noticeably weaker at close multi-person contact than at solo subjects, wide landscapes, or single-object close-ups. If your prompt needs two people to touch, expect to re-roll more, or reframe the shot so the contact point is off-camera or partially obscured.
What Consistently Worked
- Single-subject shots — one person, one animal, or one object, doing one clear action — came out clean almost every time.
- Close-ups with simple backgrounds (the coffee pour, the candle flame) were the two best results of the whole week, by a clear margin. Less going on in frame seems to mean fewer places for the model to make mistakes.
- Naming the camera move explicitly ("drone shot flying over," "close-up of") gave noticeably more stable results than leaving the camera behavior for the model to guess at.
If I Did This Again
I'd keep prompts to one clear subject doing one clear thing, save multi-person shots for when I have patience to re-roll a few times, and plan for close-ups over wide group scenes whenever the shot allows it. None of this is a knock on the tool — it's just where the honest edges of the technology are right now, and knowing them ahead of time saves a lot of wasted generations.
Request A Custom AI Video
Tell us what you're trying to create and we'll point you to the right tool — or help you set it up.