Run Kandinsky 6.0 Locally: RTX 4090 Guide to the Open Video Model Kandinsky 6.0, an open-source text-to-video model released under the MIT license on Hugging Face under the kandinskylab organization, runs on consumer GPUs including the RTX 4090 in its Lite SD configuration, while the Pro variant targets full HD output. The model supports text-to-video and image-to-video with built-in audio generation, and Fal.ai hosts it at roughly $1.40 per 5 seconds of image-to-video output. Kandinsky 6.0 arrived the same week as Tencent's Prism video model, which requires an 80GB card at 720p with no consumer-friendly build yet. Run Kandinsky 6.0 Locally: RTX 4090 Guide to the Open Video Model Kandinsky 6.0 is an open-source text-to-video model that runs on consumer GPUs. Here's how to install it locally or use it via Fal.ai. What is Kandinsky 6.0? Kandinsky 6.0 is an open-source video generation model that, unlike most recent releases in the space, is actually built to run on consumer graphics cards. It ships in multiple configurations, from a lightweight “Lite SD” variant that works on something like an RTX 4090 up to a “Pro” full HD version for people with more serious hardware. It generates video from text or from a starting image, and it can also produce audio alongside the visuals. For anyone who has been priced out of the current wave of 80GB-VRAM video models, Kandinsky 6.0 is notable mainly because you can actually run it at home. TL;DR - Kandinsky 6.0 is a new open-source video generation model that supports both text-to-video and image-to-video, with audio generation built in. - It comes in multiple sizes , including a Lite SD version aimed at consumer GPUs like the RTX 4090 and a Pro version for full HD output. - The model is distributed through Hugging Face under the kandinskylab organization using the Diffusers library, with a dedicated pipeline class for text/image-to-video. - It’s released under the MIT license , which is permissive compared to many recent video model releases. - For people who don’t want to manage local weights and VRAM, Fal.ai hosts Kandinsky 6.0 , with image-to-video generation priced at roughly $1.40 per 5 seconds of output. - Kandinsky 6.0 arrived the same week as Tencent’s Prism video model, but Prism needs an 80GB card at 720p and has no consumer-friendly build yet, which makes Kandinsky’s hardware accessibility the bigger practical story for most builders. - The model’s documentation and technical details are published alongside an arXiv paper, so the architecture and training details are open for inspection, not just the weights. Other agents ship a demo. Remy ships an app. Real backend. Real database. Real auth. Real plumbing. Remy has it all. Where does Kandinsky 6.0 come from? Kandinsky has a long history as a Russian-developed open image and video generation line, and version 6.0 continues that lineage into video. The release lands on Hugging Face under the “kandinskylab” organization, with the Pro variant packaged as a Diffusers-compatible checkpoint Kandinsky-6.0-Pro-5s-Diffusers . That naming convention tells you the default clip length the Pro model targets out of the box is around 5 seconds. The repository bundles everything the Diffusers pipeline needs: a transformer backbone, a VAE for video, a separate audio VAE and vocoder for sound generation, plus text encoder and tokenizer components. That audio pipeline is part of what makes Kandinsky 6.0 interesting. Instead of generating a silent clip that you then score separately, the model is built to produce audio and video together. It’s released under the MIT license, which is about as permissive as licensing gets in this space. Many competing open video models carry more restrictive research-only or non-commercial clauses, so the licensing terms alone make Kandinsky 6.0 worth a second look if you’re building something you intend to ship. What hardware do you need to run Kandinsky 6.0 locally? This is the headline feature. Kandinsky 6.0 was explicitly built with a tiered approach to hardware: the Lite SD configuration is designed to run on consumer-class GPUs, with the RTX 4090 called out as a target card. The Pro, full HD configuration is heavier and intended for users with more VRAM and compute headroom. That tiering matters because it’s the opposite of how most frontier open video models have been shipping lately. For comparison, Tencent’s Prism model released around the same time generates at native 2K with audio, but requires an 80GB card just to run at 720p, with no quantized or community-optimized version available yet. That puts Prism out of reach for almost anyone without access to datacenter-class hardware. Kandinsky 6.0 takes the opposite approach: trade some resolution and fidelity for the ability to actually run on hardware people already own. If you’re planning to install it locally, expect the same general requirements as other Diffusers-based video pipelines: a recent version of PyTorch, the Diffusers library, enough system RAM to load and offload model components, and patience, since even “consumer friendly” video generation is slower than image generation. Generation time will scale with resolution and clip length, so the Lite SD tier on a 4090 will be meaningfully faster than attempting the Pro tier on the same card. How do you actually set it up? Because the model is distributed as a standard Diffusers repository, the installation path looks similar to other open diffusion models: 1. Install Python, PyTorch with CUDA support, and the diffusers library a recent version, since Kandinsky 6.0 uses a dedicated pipeline class . 2. Pull the model weights from Hugging Face under the kandinskylab organization, choosing the Lite SD checkpoint if you’re on a single consumer GPU like a 4090, or the Pro checkpoint if you have more VRAM to spare. 3. Load the pipeline using the Kandinsky6TI2VAPipeline class referenced in the model’s tags, which handles both text-to-video and image-to-video generation. 4. Pass in a text prompt or a starting image plus prompt and let the pipeline generate the video and, if enabled, the accompanying audio track through the bundled audio VAE and vocoder. Remy doesn't build the plumbing. It inherits it. Other agents wire up auth, databases, models, and integrations from scratch every time you ask them to build something. Remy ships with all of it from MindStudio — so every cycle goes into the app you actually want. Because the repository ships the text encoder, tokenizer, VAE, audio VAE, and vocoder all together, there isn’t a lot of hunting around for separate components, which is a common pain point with other open video model releases that split weights across multiple repos. Is the hosted Fal.ai option worth it instead? If you don’t want to deal with local setup, driver versions, or VRAM limits, Fal.ai has already put Kandinsky 6.0 up as a hosted endpoint. Image-to-video generation through Fal is priced at around $1.40 per 5 seconds of output. That’s a reasonable way to test the model’s quality and behavior before committing to a local install, especially if you’re not sure your GPU can handle even the Lite SD tier comfortably. The tradeoff is the usual one between hosted and local: Fal costs money per generation and depends on an external service staying available and priced the way it is today, while a local install is a one-time setup cost your GPU, your time troubleshooting dependencies in exchange for unlimited generations afterward. For anyone doing high-volume experimentation or building a product around video generation, the local route on a 4090-class card will pay for itself quickly. For a one-off test or an occasional creative project, the Fal.ai pricing is simple enough to not think twice about. How does it compare to other open video models right now? Kandinsky 6.0 isn’t claiming to be the best-looking open video model available. The more relevant comparison is accessibility. Most of the recent open releases chasing top-tier visual quality, like Tencent’s Prism, demand enterprise GPU memory just to run at reduced resolution, with no community quantization yet to bring them down to consumer reach. Kandinsky 6.0 flips that priority: it ships multiple size tiers from day one specifically so people with gaming GPUs can use it without waiting for someone else to shrink it down. That makes it one of the more practical choices right now for developers who want to experiment with open video generation, including audio, without renting cloud GPU time or waiting for a quantized fork to show up weeks later. Frequently Asked Questions What GPU do I need to run Kandinsky 6.0? The Lite SD version is built to run on consumer GPUs such as an RTX 4090. The Pro, full HD version needs more VRAM and compute, so it’s better suited to higher-end or multi-GPU setups. Does Kandinsky 6.0 generate audio as well as video? Yes. The model includes a dedicated audio VAE and vocoder alongside its video components, so it can produce synchronized audio as part of generation rather than requiring a separate sound pipeline. Can I use Kandinsky 6.0 for commercial projects? The model is released under the MIT license, which is permissive and generally allows commercial use, unlike many research-only licenses attached to other open video models. How much does the hosted version cost through Fal.ai? Fal.ai’s image-to-video endpoint for Kandinsky 6.0 is priced at roughly $1.40 per 5 seconds of generated video. How does Kandinsky 6.0 compare to Tencent’s Prism model? Seven tools to build an app. Or just Remy. Editor, preview, AI agents, deploy — all in one tab. Nothing to install. Prism targets higher native resolution 2K with audio but requires an 80GB GPU even at reduced settings and has no consumer-optimized build yet. Kandinsky 6.0 trades some of that top-end quality for the ability to run on hardware people already own, like an RTX 4090.