movie_editGemini Video Agent in Infoseek

Gemini Video Agent

Turn creative direction into polished video with Gemini Omni Flash. Create a scene, bring a reference image to life, refine the result in conversation, or extend the story with natural language.

  • check_circle Generate from a prompt with cinematic motion, timing, and sound
  • check_circle Use image and video references to keep subjects and movement on brief
  • check_circle Edit and extend videos through simple follow-up instructions
  • check_circle Share your videos in one click
  • check_circle Low cost usage based pricing. No subscription required.

Prompt

A wide cinematic shot of a misty mountain range at sunrise. The camera glides slowly above the clouds; soft wind and distant birds create a calm, spacious soundtrack.

Make the brief more visual

Start with a scene, a frame, a style, or a continuation. Gemini Omni gives your agent a conversational path from first idea to usable clip.

Text to cinematic video

Describe subject, camera, light, atmosphere, movement, and sound in one natural-language request.

Move between two frames

Give the agent a starting image and an ending image, then describe the transition you want between them.

Keep the story moving

Extend a generated scene while preserving its characters, motion, camera language, and audio continuity.

Extend a coastal drive

The paired samples show the same orange classic car continuing along a coastal road with the rear tracking view, ocean, and changing light preserved.

schoolGemini Omni prompt guide

Prompt tutorials

These patterns combine a clear subject with camera movement, timing, lighting, and audio—the controls that make Omni outputs easier to direct. Use the media-role tags with source declarations to bind reference images and videos to the brief.

1. Describe a scene

The sample is a single aerial glide over misty evergreen mountains at sunrise, so name the landscape, camera movement, and quiet atmosphere directly.

Prompt pattern

“One continuous aerial shot over misty evergreen mountains at sunrise. Glide slowly above the treetops and low clouds, keep the distant ridgelines soft and layered, and add gentle wind with distant birds. No people, dialogue, or scene cuts.”

2. Bring a reference image to life

Upload the landscape as Image1. Bind it with [# References <IMAGE_REF_0>@Image1], then refer to the image as <IMAGE_REF_0> in the prompt.

Prompt pattern

“[# References <IMAGE_REF_0>@Image1] Animate the misty evergreen mountain landscape in <IMAGE_REF_0> into a 10-second cinematic aerial shot. Preserve the mountain shapes and fog banks; add a slow forward glide, warm sunrise light, and soft wind with distant birds. Use Image1 as a reference, not as a literal first frame.”

3. Direct a reference video

The player shows a sample result. In your own request, attach a separate short clip as Video1. The source declaration maps it to <VIDEO_REF_0>, which you use to describe the subject or motion to borrow.

Prompt pattern

“[# References <VIDEO_REF_0>@Video1] Generate a new cinematic shot of the silver sports car shown in <VIDEO_REF_0> driving along a coastal road at golden hour from a rear tracking camera, with the ocean on the right. Keep the car’s shape and reflective finish consistent. Use Video1 as a reference; do not use it as the source video for editing.”

4. Guide a frame-to-frame transition

Provide the two lake images as Image1 and Image2, then bind them as the first and last frames while directing the camera movement between them.

Prompt pattern

“[# Sources <FIRST_FRAME>@Image1 <LAST_FRAME>@Image2] One continuous shot: move low across the bright alpine lake in Image1, arc toward the sunlit peaks, and transition smoothly to the same lake in Image2 at pink-blue dusk. Keep the mountain silhouettes, shoreline, and reflection consistent. No jump cuts.”

5. Make a focused edit

Name the orange classic car and coastal highway so a small visual change stays anchored to the footage.

Prompt pattern

“Keep the orange classic car driving away on the coastal highway with the ocean on the right. Shift the daylight toward late golden hour and turn on the rear lights. Preserve the road, camera angle, speed, and composition. Keep everything else the same.”

6. Direct audio and timing

The visible sequence is a classic car cruising beside the ocean, which gives the sound and timing instructions a concrete scene to follow.

Prompt pattern

“As the orange classic car follows the coastal road, keep the engine low and steady with a soft ocean-surf bed. Let the warm road noise and surf continue as the car moves toward the bend; keep the sunset light and camera movement continuous.”

7. Extend video

The extended sample keeps the same orange classic car, rear tracking view, coastal road, and ocean as daylight fades toward sunset.

Prompt pattern

“Continue this coastal-road video for the next 10 seconds. Keep the orange classic car, rear tracking camera, ocean, road curve, and sunset continuity. Let the car follow the bend as the sky darkens slightly; end with the car still moving forward in one continuous shot.”

helpAI video generator FAQs

AI Video Generator FAQs

Straight answers about using Google Omni Flash and AI video to animate images, create scenes, and refine videos.

What is the easiest way to use Google Omni Flash to generate video?

The easiest way is to describe the scene in plain language. Gemini Video Agent sends your brief to Google Omni Flash and returns a reviewable video; you can start with text and add reference media when you need more control.

What is the easiest way to use AI video to animate images?

Upload an image, then describe how it should move, including the camera path, subject motion, lighting, and sound. The image provides the visual starting point while your prompt directs the animation.

What is the easiest AI video generator for beginners?

For beginners, an approachable AI video generator accepts natural-language direction and follow-up requests. Gemini Video Agent lets you start with a simple idea, then refine the result conversationally.

What AI video generator can I use without a subscription?

Gemini Video Agent uses usage-based account credits, so you can generate videos without a recurring subscription. Availability and cost depend on your account and generation settings.

Can I pay for AI video generation as I use it?

Yes. Generation is usage-based: add credits to your account and use them when you create or extend videos. There is no subscription requirement.

Can I create an AI video from a text prompt?

Yes. Describe the subject, setting, camera movement, style, timing, and audio in a text prompt. Google Omni Flash can turn that brief into a video with generated audio by default.

Can I turn a photo into an AI video?

Yes. Upload a photo and describe how to animate it, such as adding a slow push-in, product movement, or environmental motion, while specifying which details to preserve.

Can I use my own images or videos as references?

Yes. Provide your own image or video as reference media. Images can guide a subject or visual style, while videos can guide movement and appearance.

Can AI video generators edit videos through chat?

Yes. Ask for focused changes in plain language, such as changing the lighting or background, and tell the agent to keep everything else the same. Follow-up turns can continue from the previous result.

Can I make vertical videos for social media?

Yes. Ask for vertical 9:16 videos for TikTok, Instagram Reels, or YouTube Shorts. Landscape 16:9 video is also supported.

developer_boardMCP for AI video products and agents

Gemini Video Agent MCP

Add Gemini Omni video creation to your AI product through one agent boundary. Your app sends the user’s creative request and optional file context; the agent returns user-facing text plus structured video workflow results.

  • check_circle One public tool: run_agent
  • check_circle Natural-language requests with optional reference media
  • check_circle Structured results for rendering, persistence, and follow-up turns
Login to access the MCP server URL
Gemini Video Agent icon

Agent request

Creative video pipeline

Prompt

Create a 9:16 product story from this image. Add a slow camera push-in, soft studio light, and a subtle ambient soundtrack.

Status

reviewable

Output

video/mp4

Structured result

response_text, tool_result, agent_runtime, widget_bridge

FAQ

What teams need to know before connecting Gemini Video Agent.

What does the agent create?

It supports prompt-based generation, image and video reference workflows, focused conversational edits, and stateful video extensions through the Gemini Omni integration.

What can I send with a request?

Send the required prompt, plus optional file_urls or reference_files when the request depends on attached images or videos.

What are the current media limits?

The current Infoseek integration accepts up to two image references and one video reference. Uploaded video references are normalized and limited to ten seconds; uploaded audio references are not supported.

Can users keep editing a result?

Yes. The agent can continue from the account-scoped generated video state. Extensions add one to ten seconds at a time, up to forty seconds total, and only append to the end of a clip.

Which formats and audio modes are supported?

The integration supports 16:9 and 9:16 video. Gemini generates audio by default and controls the final audio and watermark behavior.

Can I call the downstream video tools directly?

No. Gemini Video Agent intentionally exposes one public agent boundary. The supported Infoseek MCP route handles the downstream workflow behind that contract.

Tools

Gemini Video Agent keeps its public contract small: provide the user’s request and optional file context, then render the structured result.

Single public tool

run_agent

smart_toy

Runs the Gemini Video Agent for the user’s creative request and preserves the video workflow result for rendering and follow-up turns.

Required input

  • prompt

Optional inputs

  • file_urls[]
  • reference_files[]

Useful output

  • response_text
  • tool_result
  • agent_runtime
  • widget_bridge

Response JSON

The agent returns a user-facing response alongside structured runtime and video results so your product can render the answer, persist state, and continue the conversation without guessing from prose.

response_text

Immediate user-facing answer.

tool_result

Structured video job and artifact result.

agent_runtime

Traceable agent execution context.

widget_bridge

Display metadata for an interactive result.

{
  "response_text": "Your Gemini video is ready to review.",
  "agent_runtime": { "app_mode": "skills_based_agent", "response_id": "resp_..." },
  "tool_result": { "status": "succeeded", "job": { "status": "succeeded", "operation": "video_generation", "artifacts": [{ "kind": "video", "mime_type": "video/mp4", "preview_url": "https://infoseek.ai/tmp/..." }] } },
  "widget_bridge": { "bridge_state_id": "bridge-...", "resource_uri": "ui://widgets/..." }
}

Quickstart + Code

Discover the one public tool, pass the creative brief, and preserve the structured result for the next turn.

One-call workflow

  1. 1. Discover run_agent with tools/list.
  2. 2. Send a detailed video request with optional file context.
  3. 3. Render response_text and persist tool_result for follow-up turns.
const result = await mcp.callTool("run_agent", {
  prompt: "Create a 9:16 product story from the attached image. Use a slow camera push-in, soft studio light, and subtle ambient sound.",
  file_urls: ["https://files.example.com/product-reference.png"]
});

const answer = result.structuredContent?.response_text || "";
const videoResult = result.structuredContent?.tool_result;

Access

The agent wrapper is designed for account-scoped use so user media, generated artifacts, and follow-up state stay associated with the authenticated Infoseek account.

shield_lock

OAuth-protected access

Use Infoseek authentication when your AI app needs to work with user-owned reference files, generated video, and private workflow state.

account_balance_wallet

Usage-based generation

Video generation uses account credits. Available resolution tiers depend on the account plan; access is checked again when a user proceeds with generation.

Why these outputs matter

Agent continuity

Stable structured results let your product continue from the last real video state instead of reconstructing it from a transcript.

Media-aware prompts

Optional file context lets the agent preserve the relationship between a user’s brief and the image or video they actually attached.

UI-ready artifacts

Response text, video results, runtime context, and bridge metadata give your product a direct path from agent request to reviewable output.

Use cases

campaign

Creative studios

Turn campaign briefs and product references into reviewable storyboards and short-form clips.

edit_video

AI video editors

Let users request focused changes and keep the rest of a good take intact.

shopping_bag

Product storytelling

Animate product imagery, add camera direction, and create reusable launch assets.

smart_toy

Agent platforms

Add multimodal video creation to an assistant without exposing provider-specific workflow stages.