Wan vs Kling vs Hailuo: China's video models, compared
By the Infer teamUpdated
Three Chinese video models, three different budgets. Hailuo 02 Pro is the cheapest to run on Infer at $0.08/sec. Kling 3.0 1080p Pro costs more per second ($0.10/sec), but it buys the highest quality ceiling of the three: native 4K capture and the top Elo score (1248, Infer's 2026-04-30 leaderboard snapshot). Wan 2.2 T2V-A14B is the most expensive to host on Infer at $0.13/sec, but it's the only one of the three with open weights, meaning you can run it on your own GPU for nothing but electricity. If you need a one-line answer: pick Hailuo for volume, Kling for the shot a client will actually watch on a big screen, and Wan if you'd rather own the model than rent it.
Infer hosts all three, so we have no favorite here. This is what the per-second rate, the leaderboard, and the license terms say when you put them side by side.
What 30 seconds of ad footage actually costs
None of these three models generates a 30-second clip in one pass; you're stitching multiple generations regardless of which you pick. Here's the raw per-second math for 30 seconds of finished footage, before any retries:
| Hailuo 02 Pro | $0.08/sec | $2.40 | 3 clips (10s each) |
| Kling 3.0 1080p Pro | $0.10/sec | $3.00 | 3–6 clips (5–10s each) |
| Wan 2.2 T2V-A14B | $0.13/sec | $3.90 | 6 clips (5s each) |
Wan is the outlier on that table, and the number is misleading if you stop reading there. Wan 2.2 T2V-A14B ships under an Apache 2.0 license, so it's open weights, and you can run it on your own hardware instead of paying Infer's per-second rate at all. A 24GB RTX 4090 handles Wan's smaller 1.3B variant comfortably at 480p, but 720p is tight on that card's VRAM, and Spheron's Wan deployment guide points to H100/H200 hardware for 720p production runs. Self-host at 480p, and the $3.90 above becomes a one-time hardware cost amortized over every render you ever make; self-host at 720p and the hardware bill climbs with it. The tradeoff either way is you're managing your own inference stack instead of hitting an API.
Infer's own listing for Hailuo 02 Pro claims it's "~50% cheaper than Kling and Veo." Check that against Infer's actual per-second cards, though, and it doesn't hold up literally: $0.08 vs. Kling's $0.10 is 20% cheaper, not 50%. That 50% figure most likely describes MiniMax's broader standard-vs-pro pricing tiers, not the specific rate Infer charges per second, so treat it as marketing positioning rather than a number you can rebuild yourself.
Spec table
| Developer | Alibaba | Kuaishou | MiniMax |
| Modality | Text-to-video | Text/image-to-video | Text/image-to-video |
| Max resolution | 1080p (typical clips) | Native 4K capture; outputs 1080p | 1080p Pro tier (fal.ai spec sheet) |
| Max duration | 5s typical clip | 5–10s typical | 6 or 10s per generation (fal.ai) |
| Native audio | Not mentioned on Infer's page | No (this version) | Unconfirmed, Infer page erroring |
| Price on Infer | $0.13/sec | $0.10/sec | $0.08/sec |
| Elo (Infer leaderboard, 2026-04-30) | #1 open-weights leaderboard (Q4 2025, per Infer copy) | #3 overall, 1248 | Not stated (page 500-errored) |
| License / self-hosting | Apache 2.0, self-hostable | Closed, API only | Closed, API only |
| Try it | Run Wan 2.2 on Infer → | Run Kling 3.0 Pro on Infer → | Try Hailuo 02 Pro in the Infer playground → |
A data caveat worth stating plainly: Hailuo 02 Pro's tryinfer.com model page was returning an HTTP 500 error at the time of writing, on every fetch attempt. The resolution and duration figures above come from MiniMax's own pricing docs (via platform.minimax.io) and fal.ai's hosted spec sheet for the same model, not from Infer's page body, which was unreachable. Treat Hailuo's row as best-available outside evidence until Infer's page is back up. Separately, Wan's spec table lists 1080p as its typical output, but Infer's own per-model notes describe the 14B variant as "optimized for 720p generation." The two claims sit on the same page without reconciling, so test at both resolutions before committing to one in a pipeline.
Scenario breakdown
Here's what the documented specs predict for three scenarios, not a claim any of the three has been run against.
A 20-second sneaker ad: tracking shot following a runner through a rain-slicked street, ending on a hero shot of the shoe. This is squarely Kling's advertised strength. Camera control is its "Top 1" claim on Infer, and a tracking-to-hero-shot cut is the kind of move that claim is built for. Neither Wan's nor Hailuo's Infer listing advertises camera control as a named feature, so on paper this scenario favors Kling by default.
A busy night market with English-language storefront signage in frame. Legible on-screen text is a known weak spot across video diffusion models generally, and it's the fastest way a quality gap between an open-weights model and a closed frontier one would show up. Wan's own model page carries a direct, sourced claim here: "~85% quality compared to Hailuo 02 Pro" (Infer's comparative line on Wan's page), which is Wan's own documented admission that it trails on exactly this kind of detail work.
A 10-second talking-presenter clip reading a product script direct to camera. None of the three has confirmed native audio on Infer. Kling's page states no native audio for this version, Wan's page doesn't mention audio at all, and Hailuo's audio support is unconfirmed since its Infer page is currently erroring (see the note below). Whichever you pick, plan on dubbing dialogue in separately; none of them is a one-pass answer for synchronized speech.
Compare these models on Infer →
Where Wan 2.2 T2V-A14B wins
Wan is the only model of the three you can own outright. Apache 2.0 licensing means no per-generation billing if you're willing to run your own GPU, a real advantage for teams doing high-volume prototyping, fine-tuning a base model for a specific style, or working under data-residency rules that rule out sending footage to a third-party API at all. It's also the model to reach for if you need to iterate hundreds of times on a single scene without watching a meter; on Infer's hosted rate that iteration would cost $0.13/sec every single time.
The honest flaw: Wan trails on quality by Infer's own comparative claim, "~85% quality compared to Hailuo 02 Pro," and its typical clip length (5 seconds) is the shortest of the three, meaning more stitching work for the same finished runtime.
Where Kling 3.0 1080p Pro wins
Kling is the quality ceiling. Native 4K capture (even though Infer outputs at 1080p), the highest Elo of the three (1248 vs. Hailuo's unstated score and Wan's open-weights-category-only ranking), and camera control Infer rates "Top 1" in its own catalog. If the output is going in front of a client on a screen bigger than a phone, or the shot depends on a specific camera move, Kling is built for exactly that brief.
The honest flaw: no native audio in this version, and at $0.10/sec it's pricier than Hailuo for the same runtime. You're paying for the resolution and camera-control ceiling, not for a bargain.
Where Hailuo 02 Pro wins
Hailuo is the value pick, full stop. At $0.08/sec it's the cheapest of the three on Infer, and MiniMax's own paygo pricing confirms the same $0.08/sec rate for the 1080p Pro tier directly, with a $0.045/sec Standard tier underneath it for even lower-stakes work. That makes it the right choice for volume social content: a UGC-style batch of 50 six-second variants runs about $24 total, cheap enough to actually A/B test at scale instead of picking one prompt and hoping.
The honest flaw: because Infer's own hailuo-02-pro page is currently erroring, its exact resolution, duration, and audio specs on Infer specifically are unverified. What's documented above comes from MiniMax's and fal.ai's listings for the same model, not Infer's page body. Re-check Infer's page directly before you commit a production pipeline to it.
The verdict
Choose Hailuo 02 Pro if the job is volume: social clips, ad variant testing, anything where you're generating dozens of takes and cost per clip matters more than any single frame being perfect. Choose Kling 3.0 Pro if the output goes in front of a client or on a screen where resolution and camera work actually show. It's the most expensive per second of the two hosted options here, but it's built for exactly that brief. Choose Wan 2.2 T2V-A14B if you'd rather own the model than rent it: self-hosting under Apache 2.0 turns per-second billing into a one-time GPU cost, at roughly 85% of Hailuo's quality by Infer's own comparison. And if the brief needs native synchronized audio, none of these three delivers it reliably today. Reach for Veo 3.1 Fast or Seedance 2.0 Pro instead.
For the full price ladder across every model Infer hosts, see the cheapest AI video APIs in 2026 and the AI video pricing index. For where these three land against the rest of the field, see the best AI video models in 2026, ranked and the state of AI video, July 2026. More head-to-heads live at the compare hub.
Frequently asked questions
Is Wan really free to self-host?
Wan 2.2 T2V-A14B ships under Apache 2.0, so there's no license fee and no per-second charge if you run it yourself. You still pay for compute: Spheron's deployment guide says a 24GB RTX 4090 handles Wan's smaller 1.3B variant comfortably at 480p, but 720p is tight on that card's VRAM, and the guide points to H100/H200 hardware for 720p production runs. Free of billing, not free of hardware.
Which Chinese video model is best for English prompts?
All three — Wan 2.2, Kling 3.0 Pro, and Hailuo 02 Pro — are built for English and Chinese prompting, and Infer's own catalog copy documents no English-specific limitation for any of them. Kling's camera-control language ('dolly in,' 'orbit left') is the most literal to prompt in English, per its documented 'Top 1' camera-control claim.
Is Hailuo 02 Pro good enough for client work?
For social-length, volume output, yes — at $0.08/sec it's the cheapest of the three on Infer and MiniMax's own Pro tier caps out at 1080p. For a client deliverable meant for a big screen, Kling 3.0 Pro's native 4K capture and top camera-control rating make it the safer pick.
What's the difference between Wan 2.2 and Wan 2.6?
Wan 2.2 T2V-A14B is the open-weights model Infer hosts today — Apache 2.0, self-hostable, priced at $0.13/sec if you use Infer's API. Wan 2.6 is Alibaba's newer, closed, API-only generation (GA since December 2025, native lip-synced audio, up to 15-second clips per Alibaba Cloud's announcement) — its weights were never published, per Wan27.org's tracking of the release.
Do any of these three generate audio natively?
Not on Infer, not yet. Kling 3.0 Pro has no native audio in this version, Wan 2.2 T2V-A14B doesn't mention audio on its Infer page, and Hailuo 02 Pro's audio support is unconfirmed (its Infer page is currently erroring — see the note below). If native audio is the requirement, Veo 3.1 Fast or Seedance 2.0 Pro are the models to reach for instead.
How many clips do I need to stitch together for a 30-second ad?
At least three with any of these models. Kling 3.0 Pro tops out around 10 seconds per clip, Hailuo 02 Pro caps at 10 seconds too, and Wan 2.2 T2V-A14B's typical clip length on Infer is 5 seconds — so a 30-second spot means 3–6 separate generations stitched in post, regardless of which model you pick.
Sources
- tryinfer.com/models/kling-3-0-pro
- tryinfer.com/models/wan-2-2-t2v-a14b
- tryinfer.com/models/hailuo-02-pro
- tryinfer.com/leaderboards
- fal.ai/models/fal-ai/minimax/hailuo-02/pro/text-to-video
- platform.minimax.io/docs/guides/pricing-paygo
- www.alibabacloud.com/en/press-room/alibaba-unveils-wan2-6-series-enabling-everyone
- artificialanalysis.ai/video/leaderboard/text-to-video
- wan27.org/blog/wan-2-6-open-source-guide
- www.spheron.network/blog/deploy-wan-2-1-ai-video-generation-gpu-setup/