movie_edit AI Video Maker Agent in Infoseek

AI Video Maker Agent

Turn a prompt, image, video, or audio reference into polished short-form video using the industry leading Seedance 2. Use it for social posts, product stories, cinematic concepts, and AI-assisted creative production.

  • check_circle Our agent UX helps optimize your prompt and set the proper configuration, making it much easier to generate a high-quality video
  • check_circle Create videos from natural-language creative direction
  • check_circle Use images, video, or audio as generation references
  • check_circle Download or share completed video artifacts
  • Usage based pricing. Cost depends on video resolution and duration. Credit your account to get started.

Prompt

"Cinematic butterfly lighting. Medium shot of a handsome pianist wearing a simple, elegant black suit. His hands move smoothly across the keys. Cut to close-up: his face immersed in music, eyes gently closed, focused and serene expression. Audio: beautiful piano music and subtle ambient audience noise from a concert hall."

Explore Your Creative Side

Create an ultra-realistic immersive experience.

Prompt

360-degree panoramic camera selfie. The camera rotates counterclockwise, capturing the dessert shop interior. Then show a woman posing in different scenes, wearing different outfits and using different props. 【Note: Use of real-person images references for video generation not allowed. Provide non indentifiable person images with blurred faces to generate similar people.】

Drag and drop images, and video as references

Prompt

A car sketch on paper. The camera pushes in. The sketch lines rise off the paper, gaining dimensionality and color, transforming into a photorealistic 3D car driving on a road.

Cinematic output quality

Prompt

Cyberpunk style, game CGI, dark scene. A run-down city corner. A young assassin fights enemies. The camera rapidly pulls back, revealing the assassin surrounded by multiple enemies. The assassin wields a lightsaber and fights fluidly, defeating enemies one by one. Enemies collapse. The entire sequence is fast yet smooth. At the end, the assassin looks up toward the camera. The video blends instantaneous "teleportation" effects with visually smooth transitions.

Immersive Audio-visual Experience

Prompt

Ultra-luxury perfume commercial. Music: dreamy electronic soundscape with steady drum beats. Macro shot: a transparent rectangular glass bottle surrounded by violently swirling purple liquid. Liquid churns with bubbles and splashes, accompanied by crisp water sounds. Dissolve transition to surface ripples of purple liquid. Close-up of delicate iris flowers floating underwater. Cut to: perfume bottle tilted on textured light surface, refracting dreamy halos. Then side still-life close-up with sharp focus on the "Iris" label in the upper-left of frame, minimal clean background. Camera tilts up to upper half of bottle rotating slowly left. Style shift: a Latina female model elegantly lifts the perfume bottle to her neck and shoulder in a pure white high-end studio background, gazing into camera. Music: rhythmic modern electronic throughout.

school Video Maker Tutorials

Tutorials

Multi Angle Product Video

Three camera reference images labeled Image 1, Image 2, and Image 3

Use multi-angle subject references and multi-image references in a precise order.

Prompt

Extract the camera from Image 1, Image 2, and Image 3, replace the background with white. The camera sits on a white table, the lens focuses on the camera in close-up, then slowly rotates around the camera as the main subject, clearly showcasing the front, side, and back.

Edit Reference Video

Before and after reference showing a squirrel changed into a pig

Upload a reference video and make a focused edit with a simple prompt.

Prompt

Change the squirrel to a pig

Using Video and Image References

Video Maker setup showing a fight-scene output, action reference video, and two character images

Combine an action reference video with character images while keeping the image order explicit.

Prompt

Reference the character actions and camera language from Video 1, generate a fight scene between Image 2 and Image 1. Image 2 is the character on the left, Image 1 is the character on the right.

Using text in Video

Illustrated coffee shop scene with a couple sitting at a table and a text-filled menu board

Keep generated text light and reliable: reserve space for the wording, then add the final copy in an editor.

Prompt

No text, no captions, no logos. Leave clean negative space in the upper third for a title.

Keep in mind

Use AI-generated text only for decorative, nonessential details—never critical prices, URLs, names, product claims, or subtitles. If you must generate text inside the video generator, keep it very short (1–5 words), large and high-contrast, static, in one location, on screen for the entire shot, and quoted exactly in the prompt.

Example

At the end of the shot, display the exact words “BOOK DIRECT” in uppercase white sans-serif letters, centered in the clean sky area, large and highly legible, held completely static for the final three seconds. No other letters, words, captions, or logos.

Extending Video

Anime-style school walkway scene with a girl walking beneath cherry blossoms at sunset

Continue from the previous output with one clearly defined next action while preserving the clip’s camera, lighting, pacing, style, sound, and aspect ratio.

Prompt

Continuing from @Video1, extend the video by 5 seconds. Keep the same characters, environment, lighting, camera style, and sound design. Show [one clearly defined next action]. End with [specific final pose or composition].

Keep in mind

Feed the previous output back as the next video reference. State the added duration explicitly. Describe one next story beat, not an entire new sequence. Preserve the existing visual language: camera movement, lighting, pacing, style, and sound. Decide the aspect ratio in the original clip; extensions inherit it.

lightbulb Pro Tip

Pro Tip! Drag and Drop Your Reference Images

For the most reliable image, video, and audio uploads, drag and drop your files directly into the AI Video Maker Agent widget.

view_quilt Advanced Video Tutorial

Tutorial: Generating High Quality Videos with Accurate Text using Storyboarding

Plan the visuals, timing, and text-safe areas in ChatGPT first, then give that storyboard to AI Video Maker Agent as a visual production brief. This workflow is especially useful when your video needs readable headlines, consistent composition, and accurate scene-to-scene timing.

1

Create the storyboard in ChatGPT

Ask ChatGPT to generate one high-resolution, landscape 16:9 storyboard image for your desired video. Describe the subject, shot timing, visual direction, voiceover, headline, text position, synchronization, camera, audio, and transition for every panel.

Show the complete BMW M3 example prompt
Generate one high-resolution, landscape 16:9 storyboard image for a premium product promotional video. Use real-world photographic imagery in every panel. Do not use illustrations, sketches, cartoons, CGI mockups, concept art, wireframes, or placeholder imagery.

Create exactly four storyboard panels in a clean 2-by-2 layout. Each panel must contain:

1. A large photorealistic live-action image at the top.
2. A clearly separated professional production-information section directly below the image.
3. All settings written explicitly below the image, including visual direction, voiceover, headline, text position, synchronization, camera, audio, and transition.

The storyboard sheet itself must be landscape 16:9. The production information below each image must be large enough to read clearly. Do not place the production notes over the photographic imagery.

Project:

* Product: A red 1988 BMW M3, specifically the classic E30 M3
* Concept: A cinematic drive through the Swiss Alps
* Video format: 16:9 widescreen
* Total duration: 12 seconds
* Number of panels: 4
* Duration per panel: Exactly 3 seconds
* Style: Photorealistic live-action automotive commercial using real Alpine roads, authentic mountain landscapes, natural lighting, accurate vehicle proportions, realistic reflections, and believable driving motion
* Mood: Emotional, aspirational, elegant, powerful, and adventurous

Panel timing and content:

PANEL 1 — 0:00–0:03

Photographic image: A realistic close-up of the red 1988 BMW E30 M3’s rectangular headlights, front grille, hood, and glossy red paint. Position the car in the lower half of the image and leave clean Alpine sky or mountain space above it.

Below the image, display these exact production notes:

VISUAL: Close-up of the red 1988 BMW E30 M3 front end; slow cinematic push-in; golden-hour reflections; authentic E30 proportions.
VOICEOVER: None.
HEADLINE: None.
TEXT POSITION: No headline in this panel.
SYNCHRONIZATION: No voiceover or headline; establish the vehicle from 0:00 to 0:03.
CAMERA: 85mm lens, slow forward push, shallow depth of field.
AUDIO: Soft cinematic music begins; subtle engine idle.
TRANSITION: Gentle cut to a wide driving shot.

PANEL 2 — 0:03–0:06

Photographic image: A realistic wide shot of the red BMW driving along a winding Alpine mountain road with snow-capped peaks, green valleys, guardrails, and natural atmospheric depth. Position the BMW in the lower-right portion of the image. Reserve the upper-left quadrant as clean negative space.

Place the headline only in the upper-left quadrant, entirely over the open sky or distant mountain background. The headline must not overlap or touch the BMW.

Below the image, display these exact production notes:

VISUAL: Wide Alpine driving shot; red 1988 BMW E30 M3 positioned in the lower-right; expansive clean sky and mountain negative space in the upper-left.
VOICEOVER: “The Ultimate Driving Machine”
HEADLINE: “The Ultimate Driving Machine”
TEXT POSITION: Upper-left quadrant, fully outside the BMW silhouette, with at least 10% clearance from the vehicle.
SYNCHRONIZATION: At 0:03, the narrator begins saying “The Ultimate Driving Machine” and the headline appears simultaneously. Keep the headline visible until 0:06.
CAMERA: 35mm lens, smooth aerial tracking shot, realistic vehicle motion.
AUDIO: Confident, warm voiceover synchronized word-for-word with the headline; rising orchestral music.
TRANSITION: Match cut following the BMW’s forward motion.

PANEL 3 — 0:06–0:09

Photographic image: A realistic low-angle tracking shot of the BMW accelerating through a dramatic mountain curve. Position the BMW in the lower-left portion of the image. Reserve the upper-right quadrant as clean mountain or sky space.

Place the headline only in the upper-right quadrant, completely separate from the BMW and road details. Never place the text over the hood, roof, windshield, wheels, doors, or body panels.

Below the image, display these exact production notes:

VISUAL: Low-angle tracking shot of the red E30 M3 accelerating through an Alpine curve; vehicle in the lower-left; clean mountain negative space in the upper-right.
VOICEOVER: “Stir Your Soul”
HEADLINE: “Stir Your Soul”
TEXT POSITION: Upper-right quadrant, entirely outside the BMW silhouette, with at least 10% clearance from the vehicle.
SYNCHRONIZATION: At 0:06, the narrator begins saying “Stir Your Soul” and the headline appears simultaneously. Keep the headline visible until 0:09.
CAMERA: 50mm lens, low side-tracking shot, dynamic but realistic motion.
AUDIO: Clear, emotionally resonant voiceover; engine acceleration rises underneath the music.
TRANSITION: Smooth motion blur into the final hero shot.

PANEL 4 — 0:09–0:12

Photographic image: A realistic hero shot of the BMW arriving at an Alpine overlook during golden-hour sunset. Position the vehicle in the lower-center or lower-right portion of the image. Reserve the upper-left quadrant as clean open sky or distant mountain space.

Place the headline only in the upper-left quadrant, with generous separation from the vehicle. Never place the text on the car.

Below the image, display these exact production notes:

VISUAL: Hero shot of the red 1988 BMW E30 M3 at a breathtaking Alpine overlook during sunset; vehicle positioned in the lower-right; clean glowing sky in the upper-left.
VOICEOVER: “Live Life to the Fullest”
HEADLINE: “Live Life to the Fullest”
TEXT POSITION: Upper-left quadrant, fully outside the BMW silhouette, with at least 10% clearance from the vehicle.
SYNCHRONIZATION: At 0:09, the narrator begins saying “Live Life to the Fullest” and the headline appears simultaneously. Keep the headline visible until 0:12.
CAMERA: 35mm lens, slow cinematic orbit, final hero composition.
AUDIO: Voiceover finishes clearly at 0:12; music reaches an emotional conclusion.
TRANSITION: Fade to black at 0:12.

Critical product-and-text composition rules:

* Never place headline text on top of the product.
* Text must never overlap, touch, obscure, or sit across any part of the BMW.
* Keep every headline outside the complete product silhouette.
* Maintain at least 10% visual clearance between text and the vehicle.
* Deliberately reposition the car or change the camera angle whenever necessary to create negative space.
* Do not place text over busy scenery, road markings, guardrails, or important vehicle details.
* The photographic image must visibly demonstrate the intended text-safe area.
* The production notes must appear below each image, never over the image.

Text and voice accuracy rules:

* Use professional Title Case exactly as follows:
  1. “The Ultimate Driving Machine”
  2. “Stir Your Soul”
  3. “Live Life to the Fullest”
* Preserve every letter, space, and word order.
* The spoken voiceover must match the corresponding headline exactly.
* Do not add any unapproved narration.
* Text must be sharp, stable, correctly spelled, and perfectly legible.
* Do not create warped letters, missing letters, duplicated letters, random symbols, flickering, or changing typography.
* Use elegant white typography with subtle shadow or a transparent dark backing.
* Do not include unintended text, road signs, license-plate lettering, watermarks, fake logos, or random symbols anywhere in the photographic imagery.

Final storyboard-image requirements:

* Exactly four panels.
* Exactly three seconds per panel.
* Landscape 16:9 format.
* Real-world photographic imagery in every panel.
* Large image on top and explicit production settings directly below each image.
* Every panel must explicitly show VISUAL, VOICEOVER, HEADLINE, TEXT POSITION, SYNCHRONIZATION, CAMERA, AUDIO, and TRANSITION.
* Generate only the finished storyboard image, not a written explanation.

Example output: a four-panel BMW E30 M3 commercial storyboard with explicit timing, text-safe areas, voiceover, camera, audio, and transition direction.

2

Save the storyboard to your Desktop

When ChatGPT finishes generating the image, download it and save it somewhere easy to find—your Desktop is ideal for the next step. Keep the original high-resolution file rather than a screenshot or compressed copy.

3

Start the AI Video Maker Agent workflow

In ChatGPT, issue this command exactly:

@AI Video Maker Agent Make a video using my storyboard

The agent will prepare the video-generation widget and its settings for you.

4

Attach the storyboard to the widget

When the AI Video Maker Agent widget appears, drag and drop the storyboard image directly into it. This attachment is what ensures the storyboard is used as the visual reference for generation.

Use the example storyboard shown here. If you saved it to your Desktop, drag the file from there into the widget’s reference area.

Example widget state after the storyboard is attached: the reference appears as @Image1 and the prompt can be updated before proceeding.

Storyboard image ready to drag into the AI Video Maker Agent widget

Drag this storyboard image into AI Video Maker Agent before proceeding.

5

Reference the attached image in your prompt

With the storyboard attached, edit the prepared prompt wherever you want the storyboard to guide the generation. Refer to the attachment using @Image1.

Example

Use @Image1 as the storyboard reference. Preserve the four-panel shot order, the BMW’s placement, and the exact headline wording and timing shown in the storyboard.

6

Choose State of the Art quality

In the widget’s Quality setting, select State of the Art to get the best results and preserve the storyboard’s visual detail and text accuracy.

7

Review settings and proceed

Review all available settings in the AI Video Maker Agent widget, including duration, aspect ratio, resolution, audio, and any watermark options. Confirm that the prompt references @Image1, then click Proceed to generate your video.

8

Enjoy your video

Once generation finishes, play the result below. The example uses the BMW M3 storyboard to produce a cinematic vertical video with the planned sequence and accurate headlines.

smart_toy Claude Setup

Use AI Video Maker Agent with Claude

Connect AI Video Maker Agent to Claude in a few steps. Go to Claude.ai, then open Settings and Connectors.

1. Go to Settings > Connectors > Add

In Connectors, click Add Custom Connector to open the setup form.

2. Enter the connector details

Set Name to AI Video Maker Agent, then set Remote Server URL to the address given in the app store.

3. Set Authentication Always, No Client ID

Set Authentication to Always required and select No client ID. Congrats! You can now login to your Infoseek account and start making video with Claude.

help AI Video Maker Agent FAQ

Frequently Asked Questions

Learn how to make videos with the AI Video Maker Agent in ChatGPT, Claude, or another AI app using the Infoseek MCP.

What's the easiest way to make videos with Seedance?

The easiest way is right here in ChatGPT: use the AI Video Maker Agent and describe what you want in plain English. It uses Seedance 2 / 2.5 for video generation.

For example:

@AI Video Maker Agent Create a cinematic video of a giraffe walking through Times Square at night, then looking into the camera and saying "I think we're lost." 5 seconds. Audio on.

For best results, specify subject + action + setting + style + duration + audio/dialogue. You can also attach an image and ask it to animate that image.

What's the best way to make videos in ChatGPT?

Use the AI Video Maker Agent app for ChatGPT, then write a clear prompt with the subject, action, camera movement, style, aspect ratio, and duration. You can attach reference media when the visual identity needs to stay consistent.

Can I use Seedance as an AI video generator from text?

Yes. The AI Video Maker Agent turns a natural-language prompt into short-form video with Seedance for social content, product stories, cinematic concepts, and creative prototypes.

Can Seedance create a video from an image?

Yes. Add an image reference and describe the motion, framing, subject behavior, and scene you want. The reference is preserved as a handle so the prompt can clearly point to it.

Can I use video or audio references with Seedance?

Yes. The AI Video Maker Agent supports image, video, and audio references to guide the visual subject, motion, pacing, or sound direction while the creative instructions stay in the prompt.

How do I write a good AI Video Maker Agent prompt with Seedance?

Describe the subject, action, environment, camera shot or movement, visual style, lighting, audio direction, aspect ratio, and duration. Concrete shot-by-shot instructions usually produce more controllable results.

How do I edit a video?

Upload the video as a reference, then describe the specific change you want in a focused prompt. Explain what should change and what should stay the same, such as the characters, setting, lighting, camera style, or pacing.

How do I extend a video?

Use the previous output as the next video reference and say exactly how much longer it should be, such as “extend the video by 5 seconds.” Describe one clear next action and preserve the original characters, environment, lighting, camera style, and sound.

How long of a video can I make?

Seedance 2.5 supports videos up to 30 seconds long. Other Seedance models support videos up to 15 seconds long. Use the extend feature to create longer videos.

Can I make AI videos with Seedance in Claude?

Yes. Connect the AI Video Maker Agent MCP to Claude, authenticate with Infoseek, and ask Claude to create a video from a prompt and optional references. Claude can follow the job status and return the completed artifact.

How much does an AI Video Maker Agent video cost?

You do not need a subscription to use the AI Video Maker Agent. It uses usage-based pricing, and the cost depends on factors such as video resolution and duration, so credit your Infoseek account before starting a generation.

developer_board MCP for AI video apps and creative agents

AI Video Maker Agent MCP

Add prompt-to-video, media-reference generation, async status tracking, and shareable video artifacts to your AI product without building the provider workflow yourself.

  • check_circle Create video jobs from prompts plus image, video, or audio references
  • check_circle Poll normalized job state until queued, running, succeeded, or failed
  • check_circle Return safe preview artifacts and public share links for completed videos
Login to access this app
AI Video Maker Agent MCP icon

Video job

Generated creative pipeline

Prompt

Create a 9:16 product launch video using @Image1 as the subject reference.

Status

running

Output

video/mp4

Artifacts

Preview URL, artifact ID, MIME type, and share link when ready.

Tools

The AI Video Maker Agent keeps its public contract simple: send the user’s request to one agent wrapper and optionally include reference files.

Single public tool

run_agent

smart_toy

Runs the AI Video Maker Agent for the user’s request. The agent interprets the creative intent, handles the video workflow, and returns the user-facing answer plus the generated video result.

Required input

  • prompt

Optional inputs

  • file_urls[]
  • reference_files[]

Useful result

  • User-facing response
  • Generated video output

Quickstart + Code

One call starts the AI Video Maker Agent. Send the user’s creative request with optional reference context, and let the wrapper handle the video workflow.

One-call workflow

  1. 1. Write a clear creative request and add optional reference context.
  2. 2. Call run_agent with the user’s prompt.
  3. 3. Render the returned user-facing response and generated video result.
const result = await mcp.callTool("run_agent", {
  prompt: "Create a 9:16 social video of our product with a slow camera push-in, soft studio lighting, and subtle ambient audio."
});

Access

The public AI Video Maker Agent wrapper requires Infoseek end-user OAuth. The Infoseek MCP route is the supported integration surface.

shield_lock

OAuth-Protected Tools

Use Infoseek auth when your AI app needs user-owned video jobs, private artifact handling, and continuity across follow-up turns.

movie_filter

Safe Artifact Surface

The public response omits private provider task IDs, private storage keys, and billing internals while preserving useful preview fields.

Why These Outputs Matter

Agent Continuity

Stable job IDs and sanitized job payloads let agents continue from the last real video state instead of guessing from prose.

Media-Aware Prompts

Reference handles keep image, video, and audio inputs aligned with prompt instructions for better creative control.

UI-Ready Artifacts

Preview URLs and artifact IDs give your app a direct path from generation status to playback and sharing.

Use Cases

campaign

Creative Automation

Let marketing agents generate short social variations from product images, motion references, and approved campaign prompts.

smart_display

AI Video Editors

Build chat-based editing workflows that use video references, image replacements, duration matching, and status-aware polling.

auto_awesome

Generated Media Products

Add video generation to creator tools, ad builders, and internal brand studios with a clean MCP contract.