Veo 3.1
Google's Veo 3.1 text-to-video with synchronized audio, in full and fast tiers.
Google$0.20 / secondtext-to-videogoogleaudio-sync
Tier
Est. cost$0.80$1 ≈ 5 seconds of video.
Example output
$0.20/ second
$0.20 per second. $1 ≈ 5 seconds of video.
Pay only for successful generations. No idle, no minimums, no per-seat.
API
Wire it up.
Endpoint
POST https://api.tryinfer.com/v1/inference/veo-3.1/text-to-videorequest
curl https://api.tryinfer.com/v1/inference/veo-3.1/text-to-video \
-H "Authorization: Bearer $INFER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input": {
"prompt": "slow cinematic dolly across a quiet city street at dawn"
}
}'
Parameters
| Name | Type | Required | Default | Description |
|---|---|---|---|---|
| seed | integer | — | — | — |
| prompt | string | Yes | — | — |
| aspect_ratio | string | — | 16:9 | 16:99:16 |
| duration_seconds | integer | — | — | 468 |
Authenticate with a bearer token. Get an API key →
Further reading
From the Infer content directory.
Best ofBest AI video models for TikTok, Reels and ShortsSeedance 2.0 Pro wins talking-head hooks, Veo 3.1 Fast wins sound-on ambient clips, and Hailuo 02 Pro is cheapest for daily trend-response posting in 2026.GuidesHow to prompt Veo 3.1, audio includedVeo 3.1 generates dialogue, SFX, and ambient sound in the same pass as video, so audio is its own prompt clause. Google's syntax and eight worked templates.GuidesFal vs Replicate vs Infer: which platform in 2026Fal wins on latency, Replicate on catalog depth, Infer on multimodal price. This is Infer's own blog comparing itself, so here's the honest breakdown.