Martini added a Gaussian-splat camera motion controller and a persistent media library, letting filmmakers previs shots before AI video generation.
What is Martini’s new camera motion feature? #
Martini, a canvas-based AI video platform, added a camera motion controller that lets users choreograph camera movement on a still image before running an image-to-video generation. Instead of hoping a text prompt produces the pan, dolly, or tilt you want, you set the camera path directly, save it, and feed that motion data into the generation itself. It builds on Martini’s existing “step into set” feature, which turns a flat image into a navigable environment.
The tool opens a scene into what looks like a Gaussian splat style 3D space. From there, users can adjust lens length, move through the environment using WASD keys and mouse look, and drop keyframes at points along a timeline. Multiple keyframes can be strung together to create more complex moves, including deliberately chaotic ones. Once a move feels right, it gets saved as a “camera move” asset that sits alongside the start frame, reference elements, audio, and prompt text used to drive the final generation, which currently runs on Kling 2.0.
TL;DR #
- Martini’s camera motion controller lets users plot camera paths on a still image inside a Gaussian-splat-style 3D view before generating video, rather than relying on prompt language to describe camera movement. - Users set keyframes along a timeline, adjust lens length, and can build multi-point moves, then save the result as a reusable camera move asset attached to the generation. - The feature works alongside Martini’s element and audio system, so a saved camera move can be combined with reference characters, locations, and even audio clips (including outputs from models like Seed 1.0) in the same generation. - Martini also introduced a persistent media library, letting characters, locations, and other assets carry over between canvases and projects instead of being locked to a single session. - The library integrates with MCP-based tools like Claude, so a user can ask an AI assistant to pull a specific character asset from their library and place it directly onto a new canvas. - Martini has said a phone-based motion controller is coming, which would let users physically move a phone to drive the virtual camera instead of using keyboard and mouse. - The current generation backend for these workflows is Kling 2.0, with the camera move and reference elements feeding into that model to produce the final video.
Other agents start typing. Remy starts asking. #
Scoping, trade-offs, edge cases — the real work. Before a line of code.
How does the camera motion tool actually work? #
The workflow starts with an existing image, for example a character standing in front of a building. From the image’s menu, a user selects “step into set” and then “camera motion.” This opens a 3D-style reconstruction of the scene that behaves like a Gaussian splat environment: navigable in three dimensions, even though it originated from a single 2D image.
From there, the controls are straightforward. Lens length can be adjusted at the top of the interface. WASD keys move the camera through the space, while the mouse aims it. A timeline sits at the bottom of the screen. Scrubbing to a point on the timeline and hitting “add key” drops a keyframe capturing the camera’s current position and orientation. Keyframes can be dragged along the timeline to adjust timing, and a motion intensity setting controls how pronounced the movement feels between keyframes. Multiple keyframes can be chained to build more elaborate paths, including intentionally jarring or unconventional moves. Once the move looks right, “save camera move” locks it in as an asset. That asset then appears alongside the start frame in Martini’s generation panel, where it can be combined with “elements” (reference characters or objects), audio files, and a text prompt before the final render, currently produced through Kling 2.0.
Why does pre-planning camera movement matter for AI video? #
Most image-to-video and text-to-video tools handle camera movement through prompt language: “slow zoom in,” “pan left,” “dolly out.” That approach is inconsistent. Models interpret camera language loosely, and small wording changes can produce wildly different results. For anyone trying to work the way a filmmaker actually works, plotting a shot before committing resources to it, that unpredictability is a real obstacle.
A direct camera motion controller flips the order of operations. Instead of generating first and hoping the camera behaves, users can block out the move first, using a real spatial interface, then generate once the path is confirmed. That mirrors how previsualization works in traditional filmmaking, where directors and cinematographers plan shots on virtual sets before a single frame of final footage gets shot.
This matters more as scenes get complex. A simple push-in is easy enough to describe in a prompt. A shot with multiple keyframes, changing focal length, and a specific arc around a character is much harder to describe in words and much easier to simply place with a mouse.
What is the persistent media library and why does it matter? #
Martini’s library feature addresses a structural limitation of canvas-based tools: assets tend to live inside the canvas or project where they were created. A character, location, or prop generated in one session hasn’t historically been easy to reuse in a different project without re-up or regenerating it.
Remy doesn't build the plumbing. It inherits it. #
Other agents wire up auth, databases, models, and integrations from scratch every time you ask them to build something.
Remy ships with all of it from MindStudio — so every cycle goes into the app you actually want.
The library changes that by letting users save assets, a character like “flamethrower girl,” for instance, into a persistent collection that’s accessible from any canvas or project going forward. That turns one-off generations into a reusable asset pipeline, which matters a lot for anyone building a series of shots, a short film, or ongoing content that depends on a consistent cast or set of locations.
The library also connects to MCP (Model Context Protocol) integrations. A user working inside an MCP-enabled assistant like Claude can issue a plain-language instruction, along the lines of pulling a specific character image from the library and placing it onto a new project’s canvas, and have that action carried out automatically. On its own this is a small convenience. For anyone managing a large or growing library of characters and locations across multiple projects, it removes a repetitive manual step.
Is Martini’s camera motion tool worth using right now? #
For anyone doing AI-assisted filmmaking or previs work, the camera motion tool addresses a real gap: precise, repeatable camera control that doesn’t depend on guessing what phrasing a model will interpret correctly. Being able to set a move once, save it, and reuse or adjust it is a meaningfully different workflow than prompt-only camera direction. That said, the tool is still developing. A planned phone-based motion controller, which would let users physically move a device to drive the virtual camera, isn’t available yet. And the quality of the final output still depends on the underlying generation model, currently Kling 2.0, so results will vary by scene complexity and how well the model executes the planned move. The library feature is more of a workflow convenience than a generation capability, but for anyone managing recurring characters or locations across multiple projects, it removes friction that canvas-based tools have generally struggled with.
Frequently Asked Questions #
What generation model does Martini’s camera motion feature use?
Based on the demonstrated workflow, camera moves created in Martini feed into Kling 2.0 for the final image-to-video generation.
Do I need a 3D model or depth map to use the camera motion tool?
No. The tool reconstructs a navigable, Gaussian-splat-style version of the scene directly from a single existing image, so no separate 3D asset or depth map is required to plot a camera path.
Can I combine camera motion with reference characters and audio?
Yes. Saved camera moves sit alongside reference “elements” like characters or locations, optional audio files, and a text prompt, all of which feed into the same generation.
What is the difference between Martini’s library and a regular project folder?
A project folder typically keeps assets tied to one canvas or project. Martini’s library is meant to persist across canvases and projects, so a character or location saved once can be reused anywhere without regenerating or re-up it.
Is a phone-based camera controller available yet?
Not yet. It has been described as an upcoming capability that would let users physically move a phone to control the virtual camera, in addition to the current keyboard and mouse controls.