← All articles

Gemini Omni (aka Google Omni and Veo Omni): videos with characters, voice and 4K

Gemini Omni - multimodal video generation

Briefly about the main thing (BLUF)

Gemini Omni, Google Omni and Veo Omni are three names of one video model: she keeps one hero in all videos, voices them and gives out up to 4K. Let’s figure out what’s behind the word “Omni,” how much generation costs, and where to start.

This model has three names, and you've probably come across all three: Gemini Omni, Google Omni And Veo Omni. Some call it by the family Gemini, some by the video line Veo, some simply “omni” - we are talking about the same neural network for generating video. She works for us in "Video" section, and you can run it without a subscription, with payment for a specific generation.

What is behind the word "Omni"

A regular video model accepts text or one picture. Here the input is mixed: a text description plus up to seven images, video or audio as additional context. You need a video in the spirit of a specific frame and with the mood of a specific melody - bring everything at once, the model will take into account each source.

This changes the very formulation of the problem. Instead of the paragraph “imagine a house like this, the light like this,” you show a photograph of the house, and only describe the action in words. The text is responsible for the plot, the attachments are responsible for the appearance.

Characters: one hero in all videos

The main problem with conventional generators is that the hero changes from video to video. In Gemini Omni there is a separate character tab for this: a hero is created once, he is saved to your library and then connected to any generation with one click. The appearance remains consistent from video to video.

This is how the series is assembled: a virtual presenter of a column, a brand mascot, a character for children’s stories. There is no need to re-upload references every time and hope that the model will “remember” the face.

Mixed input: text, image and sound in one generation

Voice: ready-made presets and your own

The video can immediately come out with sound. The settings have a set of ready-made voices - male and female, from soft high to low with hoarseness. You choose a voice, write a line in the description - and the character speaks right in the frame. If the presets are not suitable, your own voice is created in the audio tab based on the base one: it also ends up in your personal library. This is a way for the brand to have one recognizable voice behind all of its videos.

Up to 4K, duration and frame format

This is one of the few models on the site with 4K output: 720p, 1080p and 4K are available. Duration - 4, 6, 8 or 10 seconds, horizontal frame 16:9 or vertical 9:16 for social networks. The practical route is simple: run drafts at 720p, and a good option is to restart at 1080p or 4K.

How much does it cost

The price depends on the duration, quality and whether you are submitting a video as an input, and is always shown in the form before launch - you see the amount before you click the button. There is no subscription: you pay for a specific video from the general balance, you can top it up with a Russian card. A four-second 720p draft costs significantly less than a ten-second 4K, so testing an idea is cheaper than shooting the finale right away.

Where to start

  1. Open model page and make the first video from one text - this way you will feel how she reads the descriptions.
  2. Add a reference picture: let the words represent the action, and the image represent the appearance of the scene.
  3. Include a voice and a line if someone is speaking in the frame.
  4. Get a permanent character when you need a series of videos with one hero.

When describing a scene, keep order: who is in the frame, what they are doing, where it is happening, how the camera behaves. Submit your response in a separate sentence. This structure almost always produces a coherent video on the first or second try.

You are looking for a model by name Google Omni or Veo Omni - this is the same: Gemini Omni available here. Nearby, in "Video" section, there are other video models if you want to compare.

Mixed input: text, image and sound in a single generation
混合输入:在一次生成中使用文本、图像和声音

Frequently asked questions (FAQ)

Question: How to get the best result from a neural network?
Answer: Use detailed prompts (descriptions) in English, set the style and details of the scene.

Question: Can these materials be used for commercial purposes?
Answer: Yes, the generated content is entirely yours.