cd /news/artificial-intelligence/which-ai-video-models-actually-take-… · home › topics › artificial-intelligence › article
[ARTICLE · art-140064] src=mer.vin ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Which AI Video Models Actually Take a First and Last Frame

Of the 29 video models exposed by OpenRouter's public model list, 15 accept a last frame for first-and-last-frame (FLF2V) generation, according to a capability and price survey that measured seam, churn, drift and duration floors across the field. The survey found the API field name is not standardised, with four conventions in circulation: flat URLs (`first_frame_url`/`last_frame_url`), a typed media list, start/end (`start_image`/`end_image`), and image-plus-last variants such as Google Veo 3.1's `image`/`lastFrame`. The survey also notes that a 4-second duration floor makes a 2-second clip cost double, and that video inpainting with a mask plus source video is available only in VACE.

by read13 min views28 publishedSep 19, 2026
Which AI Video Models Actually Take a First and Last Frame
Image: Mer (auto-discovered)

Turning a still image into a short moving clip is now a commodity API, but the models differ far more than their marketing suggests. Some accept a first and last frame, most do not. Some hold a composition still, and pay for it in motion. Almost none will let you say only this part moves. This is a capability and price survey, with measurements.

%%{init: {"theme": "base", "themeVariables": {"background": "transparent", "lineColor": "#000000"}}}%%
graph TD
    A[Still image] --> B{What controldoes the model take?}
    B -->|first frame only| C[Image-to-videodrifts freely]
    B -->|first + last frame| D[Keyframe / FLF2Vcomposition anchored]
    B -->|mask + source video| E[Video inpaintingVACE only]
    C --> F[Loop it yourself]
    D --> G[Loops if last = first]
    E --> H[Region-limited motion]

    classDef hook fill:#189AB4,color:#fff
    classDef agent fill:#8B0000,color:#fff
    classDef decision fill:#444,color:#fff

    class B decision
    class D agent
    class E agent
    class C hook

What is image-to-video generation? #

You hand a model one picture. It returns a few seconds of video that starts from that picture. That is the whole idea, and everything below is about the controls wrapped around it.

  • Image-to-video (i2v) — one input frame. The model invents where the shot goes. Most motion, least control.
  • First-and-last-frame (FLF2V, keyframe) — two input frames. The model must arrive at the second one. Pass thesame image twice and the clip returns to where it began.
  • Video inpainting — a source video plus a mask. Only the masked region is regenerated.
  • Motion brush — paint the region that may move. Common in web apps, almost never in APIs.
Term Means Why it matters
Seam Difference between the last frame and the first A large seam means the clip cannot be repeated
Churn Mean pixel change between consecutive frames How much is actually moving
Drift Distance of frame 0 from the input image High drift means the model re-rendered rather than animated
Duration floor Shortest clip the API will bill A 4s floor makes a 2s clip cost double

An absolute seam number means nothing on its own. Compare it against an ordinary frame-to-frame change: when the two match, the repeat is invisible.

Which models accept a first and last frame #

This is the single most useful capability and it is unevenly distributed. Of 29 video models exposed by OpenRouter’s public model list, 15 accept a last frame. The field name is not standardised — there are four conventions in circulation.

Convention Fields Used by
Flat URLs first_frame_url /last_frame_url Alibaba wan2.2-kf2v-flash
Typed media list media: [{type: first_frame}, {type: last_frame}] Wan 2.7
Start / end start_image /end_image Kling, Luma, Vidu
Image + last image /last_frame_image Seedance, PixVerse, LTX, p-video
Image + lastFrame image /lastFrame Google Veo 3.1
Model Durations Resolutions
google/veo-3.1-lite 4, 6, 8 720p, 1080p
google/veo-3.1-fast 4, 6, 8 720p, 1080p, 4K
google/veo-3.1 4, 6, 8 720p, 1080p, 4K
bytedance/seedance-1-5-pro 4–12 480p, 720p, 1080p
bytedance/seedance-2.0 4–15 480p → 4K
bytedance/seedance-2.5 4–30 480p, 720p
kwaivgi/kling-v3.0-std 3–15 720p
kwaivgi/kling-v3.0-pro 3–15 720p
kwaivgi/kling-video-o1 5, 10 720p
minimax/hailuo-3 5–15 2K
minimax/hailuo-3-max 5–15 480p, 768p
alibaba/wan-2.7 2–10 720p, 1080p
alibaba/wan2.2-kf2v-flash 5 fixed 480p → 1080p
black-forest-labs/flux-3-video 5–20 720p, 1080p
lightricks/ltx-2.5-fast 2–20 720p → 4K
prunaai/p-video 1–20 720p, 1080p
wan-video/wan-2.2-i2v-fast 81–121 frames @16fps 480p, 720p

Confirmed first-frame only, despite names that suggest otherwise: alibaba/wan-3, runwayml/gen-4.5, minimax/hailuo-2.3, kwaivgi/kling-v2.6, wan-video/wan-2.5-i2v-fast. Passing two frames to a model that takes one does not error — it ignores both images and bills you for a clip built from the prompt alone.

The keyframe trade: stability costs motion #

Passing the same image as both frames gives a mathematically exact loop. It also suppresses motion, because the model has half the clip to travel out and half to come back. Measured on one still, same prompt, per region:

Region wan2.2-kf2v-flash (first+last) wan2.6-i2v-flash (first only)
Tree / leaves 2.67 34.14
Water 2.91 34.48
Cloth 2.94 26.91
Sky / hills 1.85 17.06

Roughly ten times less movement. The keyframe clip loops perfectly because almost nothing happens in it. Google’s Veo 3.1 Lite behaves the same way: given one image as both frames it holds the composition superbly — whole-frame drift 2.5 against 29.3 for an unconstrained model — while grass churned 0.23 and foliage 0.38.

Strategy Seam Tree Water Cloth Sky
No loop (raw i2v) 25.62 31.01 19.55 30.65 13.50
Model keyframe loop 2.89 1.88 1.94 2.01 1.49
Crossfade, done locally 2.69 39.67 11.59 19.22 13.55
Ping-pong, done locally 1.33 31.00 19.54 30.64 13.53

Paying a keyframe model for a loop buys stillness. Generating normal motion and closing the loop afterwards costs nothing and keeps the movement.

A further measurement: a clip the model plans as 2 seconds packs more motion per second than the first 2 seconds of a 5-second shot. Excursion 31.6 against 23.3 for the same footage trimmed down, with a cleaner seam (1.05 vs 2.31).

Region and mask control: what is actually sold #

“Only this part moves” is the most requested control and the least available. Web apps ship it; APIs mostly do not.

Two traps here. Kling’s motion-control models on Replicate and fal are not motion brush — they take image_url + video_url + character_orientation and do motion transfer from a reference video. And Vertex AI’s shared type documentation still describes a mask field with maskMode: insert/remove/outpaint; it is only valid alongside a video input, and the only model that ever shipped it was Veo 2, retired in June 2026.

The one genuinely mask-native family is Wan VACE:

One caveat decides whether this is worth it: in VACE, even the retained region passes through the VAE encode/decode round trip, so “untouched” pixels come back altered. Masking upstream reduces how hard the model fights you; it does not hand back the original pixels.

Native loop support #

Exactly one family exposes a real loop flag. Luma’s Ray models take loop: true, described as “the last frame matching the first”.

So the loop flag and the end-frame control cannot be combined, and the shortest looping Luma clip is 5 seconds. Everyone else expects you to write “seamless loop” in the prompt and hope.

Lightricks publishes a Cinemagraph LoRA for LTX-2.3 and LTX-2.5, trained to produce continuous-loop cinemagraphs where one designated element moves and backgrounds, people and static objects stay still. Training resolution 512×704, 25 frames, LoRA strength 1.0–1.2. It runs in ComfyUI against the LTX base; there is no hosted endpoint for it.

Price per second by provider #

Silent output, 1080p unless noted. Prices read from vendor pricing pages and provider rate cards.

Two billing models hide in that table. Seedance bills by video token: tokens = width × height × fps × seconds ÷ 1024, so cost scales with resolution rather than sitting at a flat rate. And wan-video/wan-2.2-i2v-fast bills flat per run — $0.05 at 480p regardless of length, which works out around $0.0066 per second at its maximum 121 frames.

The duration-minimum trap #

Headline price per second is misleading for short clips, because most models enforce a floor. Cost of one 2-second clip:

Veo is the sharpest example: durationSeconds accepts only 4, 6 or 8, and 1080p forces 8. A 2-second clip bills as eight. Very few models honour a genuinely short request.

Four models on the same still #

Same input image, 2 seconds each, same intent. Drift is distance of frame 0 from the input; wrap/step under 2 means the clip can be repeated without a visible jump.

  • p-video is the cheapest model honouring a 1-second minimum, but on this input it was near-static — churn 0.32, water 0.39 — and its wrap cost nine times an ordinary frame step, so it does not loop either.
  • LTX 2.5-fast re-rendered the scene: drift 34.3 where every other model sat at 6. Faces came back visibly different. Fine when the input is a suggestion, useless when it is the subject.
  • wan-3 produced by far the most motion, and moves people hard — faces churned 9.00. It hasno last-frame input, so any loop must be made afterwards.
  • None of the four looped on its own. Every clip needed a local crossfade or ping-pong to be repeatable.

Repeating a 2-second clip out to 20 seconds is the practical test. With a properly closed loop, the change at each repeat boundary falls below an ordinary frame step — 0.86–2.57 against a 6.55 step in one case, 0.71–0.80 against 0.82 in another. Both are invisible in playback.

Prompting does not increase motion #

A widely repeated tip is to name the moving elements. It was tested directly: a generic wind prompt against a prompt naming every specific object in the frame, on three different images, 2 seconds each with first = last frame.

Noise in both directions. The limiting factor is not wording — it is anchoring both ends of a short clip. Want more motion: generate longer, or drop the last frame.

Erasing people: inpainting models #

A common preprocessing step is removing figures from a still before animating it. The split that matters is erasers (extrapolate surrounding texture, prompt-free) versus fillers (compose new content from a prompt).

Note the polarity inversion between two endpoints at the same vendor. Always test a mask before trusting it.

Models with no mask parameter at all — and therefore unusable as an erase step, whatever their quality: Gemini 3 Pro Image / Nano Banana Pro ($0.134), Gemini 3.1 Flash Image ($0.045–0.151), Qwen Image 3.0 (docs state plainly that mask is not supported), wan2.7-image-pro ($0.069–0.075, offers bbox_list instead).

Option Weights GPU on macOS Verdict
Big-LaMa ( simple-lama-inpainting ) 206 MB Yes, MPS Best free. ~5s at 1536×864
IOPaint + lama 206 MB No — silently forced to CPU Same weights, slower
IOPaint + migan 27 MB Yes 512² internal, weakest on large holes
IOPaint + mat 251 MB No Built for large holes
IOPaint + zits 391 MB No Best structure priors, slowest
Core ML LaMa 217 MB ANE + GPU 800×800 fixed input
SDXL inpainting fp16 6.95 GB Yes Heavy
FLUX.1-Fill-dev 23.8 GB Yes Impractical locally
cv2.inpaint TELEA / NS 0 n/a Unusable — diffuses colour inward, smears a figure

Two install traps. simple-lama-inpainting dates from 2023 and will downgrade numpy and Pillow unless installed with --no-deps. And IOPaint hard-blocks MPS for every good eraser in MPS_UNSUPPORT_MODELS, silently downgrading to CPU — a guard that predates working MPS FFT support.

pip install --no-deps simple-lama-inpainting

python3 - <<'PY'
import torch
from PIL import Image
from simple_lama_inpainting import SimpleLama

device = torch.device("mps" if torch.backends.mps.is_available() else "cpu")
lama = SimpleLama(device=device)

plate = Image.open("plate.jpg").convert("RGB")
hole  = Image.open("mask.png").convert("L")    # WHITE = remove
lama(plate, hole).save("clean.png")
PY

None of the free local models will invent structure. They are Places2-trained and text-free: they extend surrounding texture across the hole. Expect plausible continuation of hills, sky and foliage, and no new stone well. Standard mitigations: dilate the mask 8–15 px (more around hair), erase one figure per pass largest first, and crop-and-stitch so the untouched area never passes through a VAE.

Person segmentation: options and speed #

Measured on an M2 Max at 820×1024:

MediaPipe’s multiclass selfie segmenter is the only free option that returns hair, body-skin, face-skin and clothes as separate labels, which is what makes “hold the faces, release the hair” expressible at all. Its Metal GPU delegate works from Python on macOS and is 6.8× the CPU path — not documented anywhere obvious. Meanwhile rembg reports CPUExecutionProvider even when CoreML is installed and available, which is why BiRefNet-lite takes nearly six seconds a frame.

SAM 2 and SAM 3 are broken on MPS today — open issues cite Placeholder storage has not been allocated on MPS device and a Triton dependency that is CUDA-only. The working Mac path is Apple’s Core ML SAM 2.1 build through Swift, not Python.

Running video models locally on Apple Silicon #

Published benchmarks on Apple hardware are scarce. The clearest set, from an M1 Max 64GB running ComfyUI:

33 frames is about 1.4 seconds at 24fps. Only the distilled LTX build is in a usable range; anything Wan-14B-shaped is economically dead locally. Three Mac-specific failures are worth knowing before starting: the official LTX-2 FP8 checkpoint fails on Metal with Undefined type Float8_e4m3fn (use GGUF), the two-stage sampler pipeline produces NaN during VAE decode (use a plain KSampler), and Wan 2.2 needs wan_2.1_vae.safetensors or it throws a 16-vs-48 channel mismatch.

Google’s line: Veo 3.1 and Gemini Omni #

There is no Veo 4. The current line is Veo 3.1 in three tiers, plus Gemini Omni Flash, which replaced Veo in the consumer Gemini surface.

  • Veo keyframes use instances[0].image andinstances[0].lastFrame ;lastFrame is only valid withimage .
  • The Gemini API cannot mute audio. Audio is always on for every Veo 3.1 variant there. Vertex hasparameters.generateAudio: false and a cheaper silent SKU — $0.03/s vs $0.05/s at 720p for Lite.
  • Durations are 4, 6 or 8 only, 24fps, 16:9 or 9:16, and generated videos are deleted from the server after 2 days .
  • Veo 3.1 on Vertex is us-central1 only . In the EU, UK, Switzerland and MENA,personGeneration accepts onlyallow_adult .
  • Gemini Omni takes two images in order as first and last frame — no named field. Its editing is conversational (“keep everything else the same”), never geometric, and it is billed by token at roughly $0.10/s at 720p.
  • conditioningScale is exposed as a passthrough on all three Veo tiers — an explicit dial for how tightly output adheres to the conditioning frames.

API traps worth knowing #

Two ffmpeg traps come up constantly when compositing generated clips. overlay defaults to format=yuv420, which chroma-subsamples exactly the alpha edge you care about — pass format=auto. And feeding a grey mask into maskedmerge in a YUV pipeline drains colour, because gray → yuv444p sets both chroma planes to 128. Measured residual: 1.057 for maskedmerge against 0.004 for alphamerge + overlay.

ffmpeg -y -i moving.mp4 -loop 1 -framerate 30 -i cutout.png \
  -filter_complex "\
[0:v]format=yuv444p,setpts=PTS-STARTPTS[bg];\
[1:v]format=rgba,scale=1920:1080:flags=lanczos,setsar=1[fg];\
[bg][fg]overlay=format=yuv444:shortest=1,format=yuv420p[v]" \
  -map "[v]" -an -r 30 -c:v libx264 -crf 18 -pix_fmt yuv420p out.mp4

A third, subtler one: -loop 1 on a still input defaults to 25fps. Composite that against a 30fps clip and ffmpeg silently resamples, producing a perfectly valid 25fps file with one frame in six duplicated. It judders, and nothing in the output reveals the cause — the only way to catch it is to compare frame rates against the source.

What to take away #

The recurring lesson across every measurement: the controls that sound most valuable are the ones least likely to exist, model names routinely overstate capability, and the cheap parts of the pipeline — looping, compositing, masking — are better done locally than bought.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openrouter 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/which-ai-video-model…] indexed:0 read:13min 2026-09-19 · —