Google Veo
Google Veo (Veo 3.1): The Foundation Cinematic Clip Model — DeepMind T2V/I2V with Native Spatial Audio, Ingredients-to-Video (≤4 Ref), True 4K, Scene Extension to 60s, SynthID Watermark, Priced by Output Second
- Name: Google Veo (Veo 3.1, formerly Veo 3 / Veo 2)
- Official Website: https://deepmind.google/models/veo/ (Consumer‑facing entry points: Gemini app / Google Flow / Google Vids; Developer‑side: Gemini API + Vertex AI)
- Short Description: Google DeepMind generative-video foundation model, T2V/I2V/Ingredients-to-Video, 8 s native clip (scene-extension chain to 60 s+), 720p/1080p native + 4K via high-quality upscale, native spatial audio (dialogue + SFX + ambient co-generated, lip-synced), up to 4 reference images for character/product/style hold, prompt-obedient camera/lighting keywords (dolly, Rembrandt light, ARRI Alexa look). Distributed inside Gemini app, Google Flow, YouTube Shorts (Fast variant), Canva, Gemini API, Vertex AI. Priced per output second (0.60 4K Standard w/ audio). Not a clip editor (Runway), not social-effects (Pika/PixVerse), not avatar (HeyGen), not repurpose (Lumen5/Pictory/Opus Clip).
1. One-Line Positioning
Veo = the prompt-obedient cinematic foundation model with the only real spatial audio co-generation, where the default idiom is detailed text or still → 8 s 1080p clip with synced dialogue/ambient in one pass, extended by Ingredients-to-Video consistency + Flow for scene-extension storytelling; billed by output second (not source minute like Opus, not per-clip like PixVerse/Pika, not credit-pack like Kling), consumer floor is Google AI Pro 249.99 or API.
ai-productivity-data/pipeline_stage=none. 2026 trap: 8 s 1080p Standard w/ audio = 4.80/clip; Free only via Flow 50 cr/day (~12 clips) or AI Studio prototype; SynthID watermark baked into every frame (provenance, not removable on consumer tiers).
2. Core Features (Table)
Layer Capability Notes Generation T2V, I2V, Ingredients-to-Video (≤4 ref imgs: character/env/style/product) 8 s native; scene extension chains to 60 s+ using last 1 s as anchor Model line (2026) Veo 3.1 (default, Jan 2026: 4K + spatial audio + 3-ref→4-ref), Veo 3.1 Fast (cheaper, Shorts-used), Veo 3.1 Lite (<$0.07/clip 8 s), legacy Veo 3 / Veo 2 No multi-model router; tier = speed/quality not separate architecture Resolution 720p / 1080p native, 4K 3840×2160 via high-quality upscale (not native 4K sensor, but DeepMind-grade SR) Kling 3.0 only other 4K-native in this pool Audio Native spatial audio — dialogue (lip-synced), SFX, ambient co-trained in one pass; car L→R pans stereo field Only Veo does true spatialization in 2026; Kling Omni/Veo-rivals do mono native Consistency Ingredients-to-Video up to 4 anchors, improved face/garment in 3.1, Flow manages cross-shot character lock Behind Runway Gen-4.5/Aleph on surgical control, ahead of Kling on prompt-obedience Camera/Style Prompt-embedded: dolly, crane, vertigo, bullet-time, “Rembrandt light”, “shot on ARRI Alexa”, lens blur No node UI (vs Runway Director Mode), no region Motion Brush Format Native 9:16 vertical + 16:9, 24/30/60 fps Vertical default for Shorts/Reels Editing None standalone — pairs with Google Flow (shot seq, scene extend, reference board) and Google Vids Generation-first, not NLE Provenance SynthID watermark in pixels + metadata on every output Non-removable on consumer/API tiers Access Gemini app, Flow, Vids, YouTube Shorts (Fast), Canva “Create a Video Clip”, Whisk, Gemini API, Vertex AI No standalone binary; all cloud
3. Pricing (2026-08, USD)
Consumer subscriptions (Google AI stack):
Tier $/mo Veo allowance Notes Free $0 Flow 50 cr/day (~12 × 8 s Fast clips), Vids ~10/mo, AI Studio prototype calls 8 s cap, watermark, non-commercial-ambiguous (personal account OK, brand use needs Pro) AI Plus $7.99 limited Flow/Vids gen Veo 3.1 Fast only AI Pro $19.99 1,000 cr/mo (~50 Veo 3 Fast or ~10 Quality clips) Entry for serious creators AI Ultra $249.99 12,500 cr/mo (~625 Fast / ~125 Quality), priority Veo 3.1 Standard, 4K, Flow max Studio/power-user floor API (Gemini API + Vertex AI, per output second, 2025-09 降价后生效):
Variant 720p 1080p 4K Veo 3.1 Lite (video+audio) $0.05/s $0.08/s — Veo 3.1 Fast (video+audio) $0.15/s $0.15/s — Veo 3.1 Standard (video only) $0.20/s $0.20/s $0.40/s Veo 3.1 Standard (video+audio) $0.40/s $0.40/s $0.60/s → 8 s 1080p Standard w/ audio = 4.80/clip; 8 s Fast w/ audio = $1.20. Enterprise volume discounts on Vertex.
pricing_type = freemium;starting_price = $0;free_tier = "Flow 50 cr/day (~12 8s Fast clips) + Vids ~10/mo + AI Studio prototype, 8s cap, SynthID watermark, personal-use only";note = "billed-by-output-second-not-credits-except-consumer-subs; ai-pro-19-99-1000cr-ai-ultra-249-99-12500cr; api-0-05-lite-0-15-fast-0-40-std-1080p-0-60-4k-std-audio; synthid-nonremovable; flow-free-50-day; sora-2-api-eol-2026-09-24-veo-is-de-facto-migration-target"
4. Access Type (v2)
access_type: web-app, api
access_display: 🌐 Gemini app (gemini.google.com) · 🎬 Google Flow (flow.google.com, dedicated Veo director surface) · 📊 Google Vids (workspace) · ▶️ YouTube Shorts / YouTube Create (Fast variant) · 🎨 Canva “Create a Video Clip” (embed) · 🧪 Whisk / AI Studio (prototype) · 🔌 Gemini API (ai.google.dev, key auth, per-second bill) + Vertex AI (cloud console, enterprise SLA, Beijing-not-available, global endpoints) · ❌ no desktop client, ❌ no iOS/Android native Veo app (Gemini app wraps it), ❌ no MCP server first-partyDifferences within the same category: Runway, PixVerse and Opus Clip feature MCP support; Pika comes with an MCP plugin but no REST API key; Kling offers dual‑endpoints plus prepaid billing; Veo is the only base model fully tied to the Google account ecosystem, with no standalone client and no MCP. It operates via dual Gemini/Vertex APIs with per‑second metered billing.
5. Reviews (2026, quality-high / cost-friction)
Source Read DIY AI 2026 bench 9.1/10 cinematic, prompt adherence 9.5, audio 10/10, cost 4/10 Promptyze / Rangy “Most obedient model” — detailed prompt = exact scene; physics slightly floaty vs Kling; audio unmatched G2/Capterra sparse (not a standalone SKU, reviewed inside Gemini/Flow) Reddit r/AIVideoCreators “Veo 3.1 + Flow is the Sora replacement” vs “$3.20 for 8 s kills batch social” Trustpilot (Google AI subs) billing/cancel complaints inherited from Google One; Veo-specific praise on audio Praise: best prompt obedience (camera/lighting/style keywords actually execute), only spatial native audio in class, 4K upscale clean enough for broadcast, Ingredients-to-Video consistency good enough for brand spokes-character, Flow scene-extension enables 60 s narrative without stitching, SynthID = enterprise provenance win, Vertex SLA + Google Cloud identity for regulated teams.
Gripes: **3.20), 8 s native cap forces Flow chaining for anything longer, 4K is upscale not native sensor (Kling 3.0 is true native 4K), complex physics (cloth/collision) still behind Kling, long-dialogue lip-sync slightly robotic in profile, Free Flow 50 cr/day evaporates, AI Ultra $249.99 gate for full 3.1 Standard, geo-blocked in CN + some markets, no MCP/no standalone app frustrates devs wanting local piping.
6. Best For / Not For
✅ Best For ❌ Not For Dialogue-driven ads / explainer heroes (native audio = no post Foley) High-volume social batch (Pika/PixVerse/Kling cheaper per clip) Broadcast-leaning 4K deliverables (upscale accepted) Frame-accurate editorial (Runway/Kapwing/Premiere) Brand spokes-character via Ingredients-to-Video (≤4 ref) Avatar translate-dub product (HeyGen) Filmmakers storyboarding in Flow (scene extend to 60 s) Text/URL/blog → social video (Lumen5/Pictory/InVideo) Enterprise on Vertex AI needing SLA + SynthID + Cloud IAM Creators needing $8–30/mo predictable clip quota Sora refugees wanting closest parity (long clips + audio + prompt faith) Offline/air-gapped gen (100% Google cloud) Nature/env shots (fluid dynamics + spatial ambient unbeatable) Keyframe-exact camera paths (Luma Ray3.2 / Runway Director)
7. Competitors (Table)
Tool Lane vs Veo 3.1 Kling 3.0 / 3.0 Omni Motion physics + true native 4K + 15 s single-pass Kling wins physics/4K-native/price-per-sec; Veo wins audio-spatial/prompt-obedience/Flow/Vertex Runway Gen-4.5 Editorial control + Aleph + Act-Two Runway wins surgical control; Veo wins realism/audio/long-form-chain Sora 2 (dying) Closest former peer, API EOL 2026-09-24 Veo is migration target; Sora led complex-prompt cinema, Veo led audio PixVerse / Pika Social effects Veo not in their price/lane Luma Ray3.2 Cinematic I2V + HDR keyframes Luma wins keyframe EXR; Veo wins audio + prompt obedience Seedance 2.0 ByteDance multi-shot, 12 @tag refs Seedance more ref-flexible; Veo audio-spatial unique HeyGen/Synthesia/Lumen5/Pictory/Opus Clip Different lanes Not generative foundation rivals Closest: Kling 3.0 (physics/4K/price) and Runway Gen-4.5 (pro workflow); Veo moat = spatial native audio + prompt obedience + Flow + Vertex/SynthID + Google distribution.
