Category research · Text and image to video

Best AI video generators, by what the evidence supports

We researched 2 of these tools in depth against 12 criteria. None clears the coverage threshold we require before naming a category leader — so this recommends by use case, and tells you how confident anyone can be.

Read this before the list

We have researched 2 of these 8 tools in depth. None of them currently clears the 75% evidence coverage our method requires before we name a category leader, so this page recommends by use case and says how confident anyone can be — it does not crown a winner.

The short answer

If you want…Best supported optionEvidence
Cinematic look developmentKlingModerately supported
Atmospheric B-roll and directed shotsRunway Gen-4.5Limited evidence
Native audio with the videoKling — Runway Gen-4.5 has noneLimited evidence
Lowest cost per 5 secondsKling Standard, $0.40 vs Runway $1.44Published pricing
Precise physics or on-screen textNothing here is well supportedNot established

Kling AI

Moderately supported 71% coverage VIDEO 3.0 Standard T2V; base-mode family proxies

Best supported for: Cinematic shots and look development.

What the evidence supports: Lighting and cinematic motion. 8 of 12 criteria have enough accepted evidence to score.

Where it warns: Fast anatomy; multi-shot retakes.

Full research profile, all criteria and 10 sources →

Runway Gen-4.5

Not established 57% coverage Gen-4.5 T2V 720p; I2V findings excluded; some route-unknown proxies

Best supported for: Atmospheric B-roll and directed shots.

What the evidence supports: Visual composition and scene control. 6 of 12 criteria have enough accepted evidence to score.

Where it warns: Complex actions; no native audio.

Overall verdict withheld. Overall verdict withheld. Coverage fell to 57% once image-to-video evidence was excluded from text-to-video scoring.

Full research profile, all criteria and 14 sources →

Not yet researched

In this category and on the list, but we have not published research on them. They are not ranked here because we would be guessing.

  • Higgsfield — Camera-move control and stylised cinematic shots, with access to several models in one place.
  • PixVerse — Fast short-form output for TikTok, Reels and YouTube Shorts on a low monthly budget.
  • Luma Dream Machine — Naturalistic motion and image-to-video from a single still. We earn nothing
  • Vidu — Reference-to-video — feeding it a character or object and keeping it consistent. We earn nothing
  • Google Flow — Scene-by-scene filmmaking with Veo, if you are already inside Google's AI subscription. We earn nothing
  • Sora 2 — Prompt-faithful scenes with native audio, if you already pay for ChatGPT. We earn nothing

What this page deliberately does not do

  • Name an overall winner. No tool here clears our coverage threshold.
  • Compare against avatar platforms. Synthesia and HeyGen do a different job — see all research.
  • Publish a score we cannot support. Where evidence is thin, the page says so.
Affiliate disclosure. Some links below earn us a commission at no extra cost to you. It does not change what we publish — see which tools we earn nothing from.