I video has moved beyond the days of unrealistic physics and excess fingers. As these models get better at mimicking the world, their use cases may be boundless.
On Wednesday at the Runway AI Summit, the AI video firm showed off a number of use cases for its video generation technology, extending far beyond creative fields. The company debuted Runway Ads for generating and optimizing paid ad campaigns, as well as Praxis-1, an open-weight world action model that adapts to new robotic embodiment with light fine-tuning.
Though robotics, advertising and Hollywood seem worlds apart, Cristóbal Valenzuela, co-CEO of Runway, told The Deep View that Runway is made to bridge that gap. In an interview, we sat down with Valenzuela to discuss the most promising use cases for video and world models, his thoughts on scaling laws, creatives getting over the AI hurdle, and why the US may be "LLM-pilled." This interview was edited for brevity and clarity.
Nat Rubio-Licht: Runway has a lot of varied bets for its technology, from creative to enterprise to robotics. Are you equally bullish on all of these?
Cristóbal Valenzuela: I'm particularly excited about realizing that video models and world models are generalization machines. What that means is that they can do things in different domains in ways that we humans might not. The beauty of large scale video models is that they can generalize, and generalization comes in many different forms. The model learns how to make really compelling videos for Hollywood and films, but that same technology now powers robots.I think what our focus needs to be on how do you scale generalization so you can enable more use cases.
Nat Rubio-Licht: What has been the biggest challenge that Runway has faced in building and scaling its model?
Cristóbal Valenzuela: I think building a world-class team is hard. It's really hard. Building the engineering know-how of how to train these models at scale is also hard. But with any discipline, if you want to become the best at it, you need to continuously be doing it for as long as you can. We've been doing this for like almost 10 years. I think we've learned a lot about what works and what doesn't. The challenge has been being able to consistently learn and scale that know-how to where we are right now.
There are some more practical, real-world challenges, like compute constraints that are preventing us and many others from moving faster, but that's something beyond our direct control.
Nat Rubio-Licht: I know that Runway's stance on language models is that we've reached a saturation point. But as the industry broadly starts to move towards these systems taking autonomous actions, where do agents fit into Runway’s strategy?
Cristóbal Valenzuela: We just launched a new product that's an agent for ads, that allows you to create end-to-end ads, and it not only creates, but it edits, publishes it, and tracks the performance.I think agents are going to become much more commonplace in every part of how you use Runway. Similar to how I think language models are progressing, like coding models are now just agent models, and you can just text them in like iMessage, and they will do the work for you. I think there's a big chance that there's similar interaction happening with just pixels, and so ads are our first stepping stone towards that. But there's more I want to do in fiction, where you can just have an agent that's a creative partner that just helps you do whatever you need to do.
Nat Rubio-Licht: Let's dive more into the creative side. Of course, one of the biggest initial applications of an AI video generator is creative fields like entertainment or marketing. Have you run into any challenges in deploying into these scenarios?
Cristóbal Valenzuela: I think that most of the challenges were about people's perception or understanding of what the models could do. That was maybe a year, year and a half ago. The first video model that we ever put out was in 2023. Companies and people had to go through the idea that you're using AI to generate every part of your production stack, and that just takes time internally. People were not really ready for it, and were not really getting it three years ago. Now, I think there's no company in the world that's not thinking about this. To be honest, I think if you're not, you're probably not going to make it.
I think the challenge right now is more about scale. How do you increase the reach? Hollywood studios were very thoughtful and careful about how they used AI two years ago. I can tell you by evidence and data that, now, they're not. They finally understand what it means, how you can use it, and so it's just about how to get more people to use it faster.
Nat Rubio-Licht: I was at the Runway film festival a few months ago in LA, and something I spoke with your CCO Jamie Umpherson about was the shifting sentiments toward AI-generated content. How have you perceived that shift? Are some creatives starting to warm up to bringing the tech into their workflows?
Cristóbal Valenzuela: There are five stages of grief. And I'm bringing that up because I think there's a version of that sentiment going on specifically within media and Hollywood. Every company on the media side went through those versions of those stages. Over the last year and a half, they were in denial, thinking "No, it's not gonna be possible. It'll never match the performance of the things that we humans do. It'll never understand intent and taste." Now we're at the point where I can show you a video that an AI made, and you're gonna like it. You're gonna have fun, and you're gonna laugh, and you're gonna cry, and it's gonna be good. I think every studio out there has now moved to the acceptance stage.
Nat Rubio-Licht: What about individual creators? Where are they in the stages of grief?
Cristóbal Valenzuela: Not everyone is at the same stage. I guess it depends on where you're coming from. I think they are perpetually people in stages of negation, and I think some people will never get it. It's not a zero-sum game. A lot of creators might want to just create things manually and to go through the pain. That's fine. There's nothing wrong with it, but many people have already moved on to needing a faster version.
Nat Rubio-Licht: Do you think there are any places where AI doesn't belong in the creative process?
Cristóbal Valenzuela: No. Creativity is a state of mind. Creativity is a way of looking at the world, and you can manifest that creativity in any way you want. The tools are there for you to use in any shape, form, or way you want them to be used. I think it's no one's place to argue if you can use something in a particular way, and I think we shouldn't prohibit people from using the tools in whatever shape, form, or way they want.
Nat Rubio-Licht: I want to talk about the strategy Runway is taking to build world models. Why have you bet on scaling video models as a means to achieve this? And how has that approach aided Runway in building Praxis-1 and Solaris?
Cristóbal Valenzuela: Well, it works. Video models work... It's similar to language models where you're doing next token prediction: These models are doing next frame prediction, and the same scaling laws apply here. So we're doubling down on the belief that we're still early on both the market and the capabilities of these models. And going back to the generalization component, if you keep scaling the performance of the models, you're going to get a lot of different use cases.
Nat Rubio-Licht: How has Runway approached the data problem with physical AI?
Cristóbal Valenzuela: I don't think there's a lack of data. You need diverse data, first of all, and getting diverse data means being very creative about who you work with, what kind of data you can get. We work with like any company you can imagine to help us gather and collect more data. So I don't think there's a data shortage. It's more about the time it takes to gather the data, and cleaning it and processing it and annotating it. That supply takes time.
Nat Rubio-Licht: Physical AI is one of the most promising applications of world models. How close do you think we are to really scaling deployment of physical AI systems, or even to achieving generalization? And what are the barriers to getting there?
Cristóbal Valenzuela: I think it's still early, but we're seeing really good progress. I think we're maybe 12 to 18 months away from deployments that are fully live in production environments, where the policy models are just fully world models. I think the challenge is that it needs to be fast. You need to be able to squeeze performance and optimization out of these models, so you can run them in fast environments. There's a lot of engineering work that has to be done to make that possible, and then improving the models. But still, this is the first generation of Praxis-1, and we'll probably have more coming up very soon. I think the next 12 to 16 months will have a good amount of progress on the physical AI side.
Nat Rubio-Licht: Robotics and physical AI aside, what are some of the other promising applications of world models? For instance, I know Runway a few weeks back debuted Solaris for building interfaces and apps on the fly.
Cristóbal Valenzuela: I think you can take the idea of Solaris a step forward and think about generating not only interfaces, but entire operating systems that could be just powered by video models. It's still early, it's still speculative, but I think we're seeing incredibly interesting products that are just building on top of the ability to simulate or generate any pixel at any given time.
Nat Rubio-Licht: It's been a very exciting few months, or even weeks, in the world model space, for instance with Fei-Fei Li's lab being acquired by AMD. What do you think is in store for this space in the next 12 to 18 months?
Cristóbal Valenzuela: I think China has been world model-pilled. America is still LLM-pilled. I think the U.S. is spending a lot of time on coding and language models, whereas China is spending most of the time on video. Most of the consumption of tokens in China, I think the estimates are around 70%, is just for video models. People watch a lot of videos in China. Most information gets shared in short-form content. Also, they're spending a lot of time in robotics, way more than the US is. I think world models are going to become the backbone of a lot of the next frontier of intelligence. So if you want to build physical intelligence systems, like autonomous embodiments in the world, you're probably going to need to have a model that runs like that for American or Western technology. There's a lot more to do there.