Image-to-Video AI Models Compared: Cost, Resolution, and Control

OpenRouter ·

Image-to-Video AI Models Compared: Cost, Resolution, and Control

If you already have the image a video should start from, choosing an image-to-video model comes down to what has to happen after that first frame. The clip may need to end on a second exact image, run longer than a few seconds, include generated audio, or stay under a price per second.

The video models we serve handle those requirements differently. This post covers four model lines: Veo 3.1, Seedance, Kling, and Grok Imagine Video. They are not the whole image-to-video catalog. The video models collection has the full list.

You call all of them through the same video generation API. The submit, poll, and download lifecycle is the same for every model. The durations, resolutions, frame roles, audio options, and prices each model accepts are not. This post compares those settings, then shows how to submit an image-to-video job with the TypeScript SDK.

Summary

  • Veo 3.1, Fast, and Lite generate 4, 6, or 8 second clips with first-frame and last-frame control and generated audio. Veo 3.1 and Fast go up to 4K, Lite stops at 1080p.
  • Seedance 2.5 generates 4 to 30 second clips at 480p or 720p, accepts first and last frames, and takes up to 50 image, video, and audio reference assets. The Seedance 2.0 line covers 4 to 15 seconds.
  • Kling v3.0 Standard and Pro generate 3 to 15 second clips at 720p with first-frame and last-frame control. Audio is optional and changes the price.
  • Grok Imagine Video 1.5 generates 1 to 15 second clips at 480p, 720p, or 1080p from a first frame only, with generated audio.
  • Every model in this post uses POST /api/v1/videos. GET /api/v1/videos/models returns each model’s supported settings and pricing SKUs.

How an image can be used in a video request

An image can define a frame in the finished video or act as reference material. We handle those two cases with two different request fields.

Text-to-video, image-to-video, and reference-to-video

With text-to-video, the model builds the scene from your prompt alone. You don’t control the opening frame beyond what the prompt describes.

With image-to-video, the image becomes the first frame, the last frame, or both. If a product video has to open on an exact product shot, this is the mode to use.

With reference-to-video, the model receives source material without fixing it to a frame position. You might provide several images of one character so the model keeps their appearance while composing a new shot.

If the image must appear at a specific point in the clip, use image-to-video. If it only needs to guide the result, use reference-to-video.

The frame_images and input_references fields

We expose the two modes as separate fields on POST /api/v1/videos.

FieldWhat it does
frame_imagesUses an image as the exact first frame, last frame, or both
input_referencesGives the model image, video, or audio material to follow without fixing it to a frame

Each frame_images entry carries a frame_type of first_frame or last_frame. Not every model supports both roles. The supported_frame_images array in the video models catalog tells you which roles a model accepts.

If you send both fields in one request, frame_images takes precedence and we treat the job as image-to-video.

Image-to-video models compared

These tables cover the Veo, Seedance, Kling, and Grok Imagine Video lines. We serve other image-to-video models as well, and the video models collection is the current catalog.

Resolutions, durations, frame roles, and audio settings come from GET /api/v1/videos/models. Prices come from each model page. We checked both on September 11, 2026. Video pricing varies by resolution, audio, and input type, so treat these figures as a comparison and check the model page before estimating production cost.

Output settings

ModelResolutionDurationGenerated audioFrame control
google/veo-3.1720p, 1080p, 4K4, 6, 8 sOptionalFirst and last
google/veo-3.1-fast720p, 1080p, 4K4, 6, 8 sOptionalFirst and last
google/veo-3.1-lite720p, 1080p4, 6, 8 sOptionalFirst and last
bytedance/seedance-2.5480p, 720p4 to 30 sOptionalFirst and last
bytedance/seedance-2.0480p, 720p, 1080p, 4K4 to 15 sOptionalFirst and last
bytedance/seedance-2.0-fast480p, 720p4 to 15 sOptionalFirst and last
bytedance/seedance-2.0-mini480p, 720p4 to 15 sOptionalFirst and last
kwaivgi/kling-v3.0-pro720p3 to 15 sOptionalFirst and last
kwaivgi/kling-v3.0-std720p3 to 15 sOptionalFirst and last
kwaivgi/kling-video-o1720p5 or 10 sOptionalFirst and last
x-ai/grok-imagine-video-1.5480p, 720p, 1080p1 to 15 sYesFirst only
x-ai/grok-imagine-video480p, 720p1 to 15 sNot listedFirst only

Reference inputs and price

ModelReference inputsPrice
google/veo-3.1Not listed$0.20/s to $0.60/s
google/veo-3.1-fastNot listed$0.08/s to $0.30/s
google/veo-3.1-liteNot listed$0.03/s to $0.08/s
bytedance/seedance-2.5Up to 50 image, video, and audio assetsFrom $0.1028/s at 480p
bytedance/seedance-2.0Image, video, and audioFrom $0.06726/s at 480p
bytedance/seedance-2.0-fastImage, video, and audioFrom $0.04035/s at 480p
bytedance/seedance-2.0-miniImage, video, and audioFrom $0.03363/s at 480p
kwaivgi/kling-v3.0-proNot listed$0.112/s without audio, $0.168/s with audio
kwaivgi/kling-v3.0-stdNot listed$0.084/s without audio, $0.126/s with audio
kwaivgi/kling-video-o1Not listed$0.112/s
x-ai/grok-imagine-video-1.5Not listed$0.08/s to $0.25/s plus $0.01 per input image
x-ai/grok-imagine-videoUp to 7 images$0.05/s to $0.07/s plus $0.002 per input image

Three notes on reading the tables.

  • Optional audio means the model exposes the generate_audio request parameter. The Grok Imagine Video models don’t expose it. The Grok Imagine Video 1.5 description says the model generates sound effects, ambience, and dialogue, so the table marks it as Yes. The earlier Grok Imagine Video description doesn’t mention audio, so the table marks it as Not listed.
  • Reference inputs reflects what the model description on its OpenRouter page states. Not listed means the description doesn’t mention reference images, video, or audio. It doesn’t mean we tested and found that references are rejected. The reference-to-video cookbook explains how to confirm support before you build on it.
  • Seedance prices are per video token, not per second. The Seedance 2.0 model pages state that the token count is height times width times duration times 24, divided by 1,024. The per-second figures in the table are the starting rates the model pages display for 480p output. Seedance 2.5 lists $10.70 per million video tokens, Seedance 2.0 lists $7.00 at 480p and 720p with separate 1080p and 4K rates, Fast lists $4.20, and Mini lists $3.50. Requests with video input use a lower per-token rate.

Because most models bill per second, duration is also a cost decision. At Seedance 2.5’s starting rate of $0.1028 per second, a 4 second clip starts at about $0.41 and a 30 second clip starts at about $3.08. The charge is higher at 720p or with a different input configuration.

How to choose

Duration and frame control narrow the list first. Once you know how long the clip must be and whether you need a fixed first frame, last frame, or both, compare the remaining models on audio, resolution, and cost.

Clips of 8 seconds or less with generated audio

Veo 3.1, Veo 3.1 Fast, and Veo 3.1 Lite generate 4, 6, or 8 second clips in 16:9 or 9:16 with first-frame and last-frame control and a generate_audio option. Veo 3.1 and Fast accept 720p, 1080p, and 4K. Lite accepts 720p and 1080p.

Lite starts at $0.03 per second for 720p without audio and Veo 3.1 starts at $0.20 per second without audio, so Lite or Fast cost less while the prompt and motion are still changing. The image-to-video cookbook uses google/veo-3.1-lite for this reason.

The Veo 3.1 Fast and Lite model descriptions list SynthID watermarking, Google’s imperceptible marker for AI-generated content. If your workflow has provenance or watermark requirements, check the model description for the Veo variant you plan to use.

Clips up to 30 seconds and many reference assets

Seedance 2.5 generates 4 to 30 second clips at 480p or 720p, accepts first and last frames, and takes up to 50 image, video, and audio reference assets. It is the only model in this comparison that produces a 20 or 30 second shot in one generation.

The Seedance 2.0 line covers 4 to 15 seconds at lower starting rates. Seedance 2.0 accepts 480p, 720p, 1080p, and 4K. Fast and Mini accept 480p and 720p. All three list image, video, and audio references in their descriptions.

If you don’t need the 30 second window or the 50 asset allowance, the 2.0 line covers the same first-frame and last-frame workflow at a lower rate.

Every Seedance model runs through hosted APIs only. We found no published weights or model license for any of them, so none can run locally or air-gapped.

First and last frame control from 3 to 15 seconds

Kling v3.0 Standard and Pro generate 3 to 15 second clips at 720p in 16:9, 9:16, or 1:1 with first-frame and last-frame control. Audio is optional. Standard costs $0.084 per second without audio and $0.126 with audio. Pro costs $0.112 and $0.168. Kuaishou describes Pro as the higher visual quality tier.

Kling Video O1 generates 5 or 10 second clips at 720p with the same frame roles and costs $0.112 per second. Use it when one of those two durations already fits the shot.

One starting frame and clips up to 15 seconds

Grok Imagine Video 1.5 generates 1 to 15 second clips at 480p, 720p, or 1080p in seven aspect ratios. Its supported_frame_images array lists first_frame only, so you can fix the opening image but not the closing one. The model description says it generates sound effects, ambience, and dialogue.

The earlier Grok Imagine Video stops at 720p and costs less per second. Its description lists reference-to-video with up to seven reference images, which the 1.5 description does not mention.

If the clip must end on a second exact image, use Veo, Seedance, or Kling instead.

Submit an image-to-video job

Once you have chosen a model, check that the source image is reachable through a stable, directly downloadable HTTPS URL and that the model supports the frame role you need.

A URL can work in your browser and still fail when the provider fetches it, for example when it depends on a session cookie, redirects through an HTML page, or sits behind a bot check. Check the URL directly before you spend credits on a generation.

curl -I "$FIRST_FRAME_URL"

You want a 200 status and an image content type such as image/jpeg, image/png, or image/webp.

The example below uses google/veo-3.1-lite to generate a 4 second 720p clip with the supplied image as the first frame and audio turned off. The REST API uses snake_case field names and the TypeScript SDK maps them to camelCase. frame_images becomes frameImages, frame_type becomes frameType, image_url becomes imageUrl, and generate_audio becomes generateAudio. Values such as first_frame and image_url stay the same.

The example also passes { retries: { strategy: "none" } } as the second argument to generate. By default the SDK retries requests that return a 5XX status with exponential backoff for up to one hour. A video submission starts billable work, so if the server accepts the job and then returns a transient error before you receive the job ID, an automatic retry submits a second paid job and the client only sees the second one. With retries off for the submission, a failed submission surfaces as an error and you decide whether to resubmit. The per-call option leaves the default retries in place for the read-only getGeneration polling calls, which are safe to repeat.

import { OpenRouter } from "@openrouter/sdk";

const apiKey = process.env.OPENROUTER_API_KEY;
const firstFrameUrl = process.env.FIRST_FRAME_URL;

if (!apiKey) {
  throw new Error("Set OPENROUTER_API_KEY first.");
}

if (!firstFrameUrl) {
  throw new Error("Set FIRST_FRAME_URL to a directly downloadable image URL.");
}

const openRouter = new OpenRouter({ apiKey });

const job = await openRouter.videoGeneration.generate(
  {
    videoGenerationRequest: {
      model: "google/veo-3.1-lite",
      prompt:
        "The camera slowly pushes in as the subject turns toward warm window light, cinematic, realistic motion",
      duration: 4,
      resolution: "720p",
      aspectRatio: "16:9",
      generateAudio: false,
      frameImages: [
        {
          type: "image_url",
          imageUrl: { url: firstFrameUrl },
          frameType: "first_frame",
        },
      ],
    },
  },
  { retries: { strategy: "none" } },
);

console.log(job.id, job.status);

A successful submission returns a job, not the finished MP4. Store the job ID and poll GET /api/v1/videos/{jobId} until the status is completed, then download the video from the job’s unsigned_urls array, which the SDK exposes as unsignedUrls, with your API key in the Authorization header. Stop polling on failed, cancelled, or expired. Our video generation guide recommends a 30 second interval between polls.

let current = job;

while (!["completed", "failed", "cancelled", "expired"].includes(current.status)) {
  await new Promise((resolve) => setTimeout(resolve, 30_000));
  current = await openRouter.videoGeneration.getGeneration({ jobId: job.id });
}

if (current.status !== "completed") {
  throw new Error(current.error ?? `Video generation ${current.status}.`);
}

If the model supports a supplied last frame, add a second frameImages entry with frameType: "last_frame".

frameImages: [
  {
    type: "image_url",
    imageUrl: { url: firstFrameUrl },
    frameType: "first_frame",
  },
  {
    type: "image_url",
    imageUrl: { url: lastFrameUrl },
    frameType: "last_frame",
  },
],

If your application lets users switch models, query GET /api/v1/videos/models before you build the request. The response includes each model’s supported_durations, supported_resolutions, supported_aspect_ratios, supported_frame_images, generate_audio flag, and pricing_skus. A request with an unsupported duration, resolution, or aspect ratio returns a 400 that lists the supported values.

FAQ

Which image-to-video models are available on OpenRouter?

The video catalog includes Veo 3.1 and its Fast and Lite variants, Seedance 2.5 and the Seedance 2.0 line, Kling v3.0 Standard and Pro, Kling Video O1, and Grok Imagine Video 1.5 and its predecessor, among others. The video models collection and GET /api/v1/videos/models list the current set.

What is the difference between image-to-video and reference-to-video?

Image-to-video uses frame_images. The image becomes the exact first frame, last frame, or both. Reference-to-video uses input_references. The model uses the images, video, or audio as guidance without placing them at a fixed frame. If a request includes both fields, frame_images takes precedence and we treat the request as image-to-video.

How much does image-to-video generation cost?

Most video models are priced per second of output, and the rate can change with resolution, whether audio is generated, and whether the request includes video input. The Seedance models are priced per video token instead. Each model page lists the current rates, and GET /api/v1/videos/models returns the pricing SKUs for each model.

How long does an image-to-video job take?

Video generation is asynchronous. POST /api/v1/videos returns a job immediately, and you poll GET /api/v1/videos/{jobId} until the status is completed. Our documentation recommends a 30-second interval between polls. Turnaround varies by model, duration, and resolution.

What happens if a video generation job fails?

A job can end as failed, cancelled, or expired. Stop polling when the status reaches one of those values and read the error field. A transient error on the poll request itself is not a failed job, so retry the status request for the same job ID instead of submitting a new paid generation.

By subscribing you agree to receive the OpenRouter newsletter: model usage data, product updates, and research reports, about one email a week. Unsubscribe anytime via the link in every email. See our Privacy Policy.