Text to cinematic video
Describe subject, camera, light, atmosphere, movement, and sound in one natural-language request.
Turn creative direction into polished video with Gemini Omni Flash. Create a scene, bring a reference image to life, refine the result in conversation, or extend the story with natural language.
Prompt
A wide cinematic shot of a misty mountain range at sunrise. The camera glides slowly above the clouds; soft wind and distant birds create a calm, spacious soundtrack.
Start with a scene, a frame, a style, or a continuation. Gemini Omni gives your agent a conversational path from first idea to usable clip.
Describe subject, camera, light, atmosphere, movement, and sound in one natural-language request.
Give the agent a starting image and an ending image, then describe the transition you want between them.
Extend a generated scene while preserving its characters, motion, camera language, and audio continuity.
The paired samples show the same orange classic car continuing along a coastal road with the rear tracking view, ocean, and changing light preserved.
These patterns combine a clear subject with camera movement, timing, lighting, and audio—the controls that make Omni outputs easier to direct. Use the media-role tags with source declarations to bind reference images and videos to the brief.
The sample is a single aerial glide over misty evergreen mountains at sunrise, so name the landscape, camera movement, and quiet atmosphere directly.
Prompt pattern
“One continuous aerial shot over misty evergreen mountains at sunrise. Glide slowly above the treetops and low clouds, keep the distant ridgelines soft and layered, and add gentle wind with distant birds. No people, dialogue, or scene cuts.”
Upload the landscape as Image1. Bind it with [# References <IMAGE_REF_0>@Image1], then refer to the image as <IMAGE_REF_0> in the prompt.
Prompt pattern
“[# References <IMAGE_REF_0>@Image1] Animate the misty evergreen mountain landscape in <IMAGE_REF_0> into a 10-second cinematic aerial shot. Preserve the mountain shapes and fog banks; add a slow forward glide, warm sunrise light, and soft wind with distant birds. Use Image1 as a reference, not as a literal first frame.”
The player shows a sample result. In your own request, attach a separate short clip as Video1. The source declaration maps it to <VIDEO_REF_0>, which you use to describe the subject or motion to borrow.
Prompt pattern
“[# References <VIDEO_REF_0>@Video1] Generate a new cinematic shot of the silver sports car shown in <VIDEO_REF_0> driving along a coastal road at golden hour from a rear tracking camera, with the ocean on the right. Keep the car’s shape and reflective finish consistent. Use Video1 as a reference; do not use it as the source video for editing.”
Provide the two lake images as Image1 and Image2, then bind them as the first and last frames while directing the camera movement between them.
Prompt pattern
“[# Sources <FIRST_FRAME>@Image1 <LAST_FRAME>@Image2] One continuous shot: move low across the bright alpine lake in Image1, arc toward the sunlit peaks, and transition smoothly to the same lake in Image2 at pink-blue dusk. Keep the mountain silhouettes, shoreline, and reflection consistent. No jump cuts.”
Name the orange classic car and coastal highway so a small visual change stays anchored to the footage.
Prompt pattern
“Keep the orange classic car driving away on the coastal highway with the ocean on the right. Shift the daylight toward late golden hour and turn on the rear lights. Preserve the road, camera angle, speed, and composition. Keep everything else the same.”
The visible sequence is a classic car cruising beside the ocean, which gives the sound and timing instructions a concrete scene to follow.
Prompt pattern
“As the orange classic car follows the coastal road, keep the engine low and steady with a soft ocean-surf bed. Let the warm road noise and surf continue as the car moves toward the bend; keep the sunset light and camera movement continuous.”
The extended sample keeps the same orange classic car, rear tracking view, coastal road, and ocean as daylight fades toward sunset.
Prompt pattern
“Continue this coastal-road video for the next 10 seconds. Keep the orange classic car, rear tracking camera, ocean, road curve, and sunset continuity. Let the car follow the bend as the sky darkens slightly; end with the car still moving forward in one continuous shot.”
Straight answers about using Google Omni Flash and AI video to animate images, create scenes, and refine videos.
The easiest way is to describe the scene in plain language. Gemini Video Agent sends your brief to Google Omni Flash and returns a reviewable video; you can start with text and add reference media when you need more control.
Upload an image, then describe how it should move, including the camera path, subject motion, lighting, and sound. The image provides the visual starting point while your prompt directs the animation.
For beginners, an approachable AI video generator accepts natural-language direction and follow-up requests. Gemini Video Agent lets you start with a simple idea, then refine the result conversationally.
Gemini Video Agent uses usage-based account credits, so you can generate videos without a recurring subscription. Availability and cost depend on your account and generation settings.
Yes. Generation is usage-based: add credits to your account and use them when you create or extend videos. There is no subscription requirement.
Yes. Describe the subject, setting, camera movement, style, timing, and audio in a text prompt. Google Omni Flash can turn that brief into a video with generated audio by default.
Yes. Upload a photo and describe how to animate it, such as adding a slow push-in, product movement, or environmental motion, while specifying which details to preserve.
Yes. Provide your own image or video as reference media. Images can guide a subject or visual style, while videos can guide movement and appearance.
Yes. Ask for focused changes in plain language, such as changing the lighting or background, and tell the agent to keep everything else the same. Follow-up turns can continue from the previous result.
Yes. Ask for vertical 9:16 videos for TikTok, Instagram Reels, or YouTube Shorts. Landscape 16:9 video is also supported.
Add Gemini Omni video creation to your AI product through one agent boundary. Your app sends the user’s creative request and optional file context; the agent returns user-facing text plus structured video workflow results.
run_agentAgent request
Prompt
Create a 9:16 product story from this image. Add a slow camera push-in, soft studio light, and a subtle ambient soundtrack.
Status
reviewable
Output
video/mp4
Structured result
response_text, tool_result, agent_runtime, widget_bridge
What teams need to know before connecting Gemini Video Agent.
It supports prompt-based generation, image and video reference workflows, focused conversational edits, and stateful video extensions through the Gemini Omni integration.
Send the required prompt, plus optional file_urls or reference_files when the request depends on attached images or videos.
The current Infoseek integration accepts up to two image references and one video reference. Uploaded video references are normalized and limited to ten seconds; uploaded audio references are not supported.
Yes. The agent can continue from the account-scoped generated video state. Extensions add one to ten seconds at a time, up to forty seconds total, and only append to the end of a clip.
The integration supports 16:9 and 9:16 video. Gemini generates audio by default and controls the final audio and watermark behavior.
No. Gemini Video Agent intentionally exposes one public agent boundary. The supported Infoseek MCP route handles the downstream workflow behind that contract.
Gemini Video Agent keeps its public contract small: provide the user’s request and optional file context, then render the structured result.
Single public tool
Runs the Gemini Video Agent for the user’s creative request and preserves the video workflow result for rendering and follow-up turns.
Required input
promptOptional inputs
file_urls[]reference_files[]Useful output
response_texttool_resultagent_runtimewidget_bridgeThe agent returns a user-facing response alongside structured runtime and video results so your product can render the answer, persist state, and continue the conversation without guessing from prose.
response_text
Immediate user-facing answer.
tool_result
Structured video job and artifact result.
agent_runtime
Traceable agent execution context.
widget_bridge
Display metadata for an interactive result.
{ "response_text": "Your Gemini video is ready to review.", "agent_runtime": { "app_mode": "skills_based_agent", "response_id": "resp_..." }, "tool_result": { "status": "succeeded", "job": { "status": "succeeded", "operation": "video_generation", "artifacts": [{ "kind": "video", "mime_type": "video/mp4", "preview_url": "https://infoseek.ai/tmp/..." }] } }, "widget_bridge": { "bridge_state_id": "bridge-...", "resource_uri": "ui://widgets/..." } }
Discover the one public tool, pass the creative brief, and preserve the structured result for the next turn.
run_agent with tools/list.response_text and persist tool_result for follow-up turns.const result = await mcp.callTool("run_agent", {
prompt: "Create a 9:16 product story from the attached image. Use a slow camera push-in, soft studio light, and subtle ambient sound.",
file_urls: ["https://files.example.com/product-reference.png"]
});
const answer = result.structuredContent?.response_text || "";
const videoResult = result.structuredContent?.tool_result;The agent wrapper is designed for account-scoped use so user media, generated artifacts, and follow-up state stay associated with the authenticated Infoseek account.
Use Infoseek authentication when your AI app needs to work with user-owned reference files, generated video, and private workflow state.
Video generation uses account credits. Available resolution tiers depend on the account plan; access is checked again when a user proceeds with generation.
Stable structured results let your product continue from the last real video state instead of reconstructing it from a transcript.
Optional file context lets the agent preserve the relationship between a user’s brief and the image or video they actually attached.
Response text, video results, runtime context, and bridge metadata give your product a direct path from agent request to reviewable output.
Turn campaign briefs and product references into reviewable storyboards and short-form clips.
Let users request focused changes and keep the rest of a good take intact.
Animate product imagery, add camera direction, and create reusable launch assets.
Add multimodal video creation to an assistant without exposing provider-specific workflow stages.