VEO 3.1 is an AI video generator that builds clips from text or photos and is known for realistic motion with audio in the frame. Describe a scene in plain words or upload images, and the model returns a short video where movement feels natural and sound can match the action. Three quality tiers let you trade speed, cost, and polish.

The model runs in two directions. Text-to-video: you write out the scene, the subjects, the camera move, and VEO 3.1 films your script. Image-to-video: you upload one or two pictures and a frozen frame becomes a living scene — leaves sway, traffic flows, a portrait turns its head.
Audio is the standout feature. Unlike most video models, VEO 3.1 can generate clips with a soundtrack baked in: street ambience, effects, even lines of dialogue that fit the scene. Occasionally a clip comes back silent — that is a quirk of the model, not a bug, and a rerun usually fixes it.
Three generation tiers are available. Lite is the budget option for sketches and idea-testing, when you want many cheap attempts. Fast is the workhorse — solid output with quick turnaround, enough for most everyday jobs. Quality delivers the best image for final versions.
Pricing depends on the tier and the clip parameters, drawn from one token balance shared across the whole site. Treat it like draft and final renders in editing: explore on Lite or Fast, then redo the winning prompt on Quality once you know it works.

Upload an image and it becomes the opening frame; your text then directs what happens next — branches start moving, the camera pushes in, someone glances over their shoulder. This is how landscapes, interiors, product shots, and old family photos get brought to life.
With two images the model treats the first as the starting frame and the second as the ending frame, then builds a coherent transition between them. That is a ready-made recipe for transformation clips: product before and after, day turning to night, a sketch becoming a finished object.
Settings include 16:9 landscape, 9:16 portrait, and an Auto option where the model picks the ratio itself. Every finished video is saved to your generation history and can be downloaded whenever you need it.
A clip you like can be extended so the story keeps going without starting over. You can also pull the last frame out of any result and use it as the starting image for the next generation — a simple way to stitch short pieces into longer, continuous sequences.
VEO 3.1 enforces strict internal standards: it will not generate children, animal cruelty, or adult content. A request may also be declined for less obvious reasons when a scene fails the model's internal review — rewording the description is the way through.
Refusals and failures are free: tokens are only deducted for videos that were actually created. If a run does not go through, simplify the scene, drop the contested elements, and try again, or switch to a different quality tier.
Yes — VEO 3.1 can produce clips with ambience, sound effects, and even dialogue that matches the scene. Some runs come back silent; if that happens, just generate again.
Lite is the cheapest tier for drafts and experiments. Fast balances speed and quality for everyday work. Quality renders the best-looking image for final cuts. The price of each clip follows the tier you choose.
Yes. Pick the 9:16 portrait format in the settings — made for Reels, Shorts, and stories. Landscape 16:9 is there too, plus an Auto mode that lets the model choose the ratio.