• Picasso AI LogoLogo Picasso IA
  • Home
  • Image
  • Video
  • Video Edit
  • Lipsync
  • Enhance
  • Music
  • Voice
  • Transcribe
  • Chat
  • 3D
  • Upscale
  • Remove BG
  • Effects
  • AI ToolkitNEW
  • Generations
  • Billing
  • Support
  • Account
Seedance 2.5 IS HERE ยท Nano Banana 2 & GPT Image 2.5 UNLIMITED UNTIL September 15Upgrade
  1. Collection
  2. Text to Video
  3. Gemini Omni 1.1

Generate Video with Audio Using Gemini Omni 1.1

Gemini Omni 1.1 turns a text description, a still photo, or an existing video clip into a new video that comes with its own soundtrack already built in. Instead of writing a scene and then hunting for music or sound effects separately, you describe the action, the mood, and the audio you want, and the model produces a video where the visuals and sound match from the first frame. This solves the usual back-and-forth of stitching silent AI clips together with audio in a separate editor. It handles three different jobs well. Give it only a prompt and it generates a video from scratch in resolutions up to 4K, choosing between a 16:9 widescreen frame or a 9:16 vertical one for social feeds. Give it a starting image and an ending image and it fills in a smooth transition between the two, useful for product reveals or before-and-after shots. Give it a finished video plus a written instruction like 'make the sky stormy' and it edits that specific detail while leaving the rest of the footage untouched. For a marketer or content creator, that means going from an idea to a shareable clip in one pass, without switching between separate tools for animation, voice, and sound design. Reference images can also lock in a particular character or object so it stays consistent across shots. Open Gemini Omni 1.1 on Picasso IA, describe the video you want, and download the result once it renders.

Official

Google

3.1k runs

Gemini Omni 1.1

2026-08-26

Commercial Use

Generate Video with Audio Using Gemini Omni 1.1

Table of contents

  • Overview
  • How It Works
  • Frequently Asked Questions
  • Credit Cost
  • Features
  • Use Cases
  • Examples
Get Nano Banana Pro

Overview

Gemini Omni 1.1 turns a text description, a starting image, or an existing clip into a finished video with synchronized audio built in, all inside Picasso IA. It works well for anyone who needs a short video fast, whether that means animating a product photo, patching one detail in a clip without re-shooting it, or building a scene from scratch with a few descriptive sentences. You describe the shot, camera movement, lighting, mood, and sound, and the model renders a clip that matches, saving the back-and-forth of traditional editing software. A marketer prepping a social ad, a teacher illustrating a concept, or a filmmaker testing a shot list can all get a usable draft in one pass.

How It Works

  • Write a prompt describing the scene, camera movement, lighting, mood, and any audio you want in the clip
  • Optionally upload a starting image to animate, or add an ending frame so the model builds a smooth transition between the two
  • Upload an existing video instead and describe the single change you want, like adjusting the weather or lighting, while everything else stays untouched
  • Choose a resolution from a quick 360p draft up to 4k, and pick a 16:9 or 9:16 aspect ratio for the final output
  • Add reference images to keep a specific character, product, or style consistent across the generated video

Frequently Asked Questions

Do I need programming skills or technical knowledge to use this? No, just open Gemini Omni 1.1 on Picasso IA, adjust the settings you want, and hit generate.

Is it free to try? You can test Gemini Omni 1.1 on Picasso IA with the credits included in your account, so there is no separate purchase before your first clip.

How long does it take to get results? A 360p draft renders in under a minute, while higher resolutions like 1080p or 4k take a bit longer depending on the length and detail of the scene.

What output formats are supported? Every clip is delivered as a standard video file with audio already mixed in, ready to download and drop straight into your editing timeline or social post.

Can I customize the output quality or style? Yes, pick a resolution from 360p to 4k, choose a 16:9 or 9:16 frame, and add reference images to lock in a specific character or visual style.

How many times can I run the model? You can run it as many times as you have credits for, so testing a few prompt variations before settling on a final clip is normal practice.

Where can I use the outputs? The videos are yours to use in ads, social content, presentations, or any other project, with no watermark attached.

Credit Cost

The credit cost for this model varies based on the settings you choose. Below are the costs per configuration:

ConfigurationCredits
360p10per video
720p30per video
1080p46per video
4k90per video

Features

Everything this model can do for you

Built-in audio

Every generated video includes matching sound, so you skip separate audio editing.

Up to 4K resolution

Render sharp video at 360p, 720p, 1080p, or full 4K depending on your needs.

Image-to-video animation

Turn a single photo into a moving video with natural motion.

Frame interpolation

Supply a start and end image and get a smooth transition between them.

Video editing mode

Change one detail in an existing clip with a text instruction, keeping the rest intact.

Reference-guided consistency

Lock in a character, object, or style across a video using reference images.

Flexible aspect ratios

Switch between widescreen 16:9 and vertical 9:16 for any platform.

Draft mode

Preview a scene fast in 360p before rendering a higher resolution version.

Use Cases

Generate a 4K video from a text prompt describing the scene, camera movement, lighting, and audio

Animate a single still photo into a moving video clip with matching sound

Create a smooth transition video by supplying a starting photo and an ending photo

Edit an existing video with a text instruction, like changing the weather or background, while keeping the rest of the footage unchanged

Produce a vertical 9:16 video for social media stories directly from a written description

Keep a specific character or product consistent across multiple video clips using reference images

Render a fast 360p draft to preview a scene before generating the final high-resolution version

Add background audio without speech to a video by specifying 'no dialogue' in the prompt

Examples

720p
16:9
47.5s

A cinematic drone shot through misty pine mountains at sunrise, gentle wind and birdsong. No dialogue.

Switch Category

Effects

Text To Image

Text To Video

Large Language Models

Text To Speech

Super Resolution

Lipsync

AI Music Generation

Video Editing

Speech To Text

AI Enhance Videos

Remove Backgrounds