A new AI model can render an entire website in real time without running a single line of code behind it β and that wasn't even the most surprising AI news of the day.
The headline development came from a research preview of an "Interface World Model" that treats every webpage or app as a video the AI is generating on the fly. Instead of code deciding what appears on screen, a video model draws every single frame in real time as a person clicks and drags, while a language model interprets each action and decides what should happen next. The demos shown off included a virtual shirt try-on you drag from a rack onto a photo, a salad you build by dropping in ingredients, and an interactive combustion-engine simulation β all generated frame by frame with nothing resembling traditional source code underneath. In blind testing, human evaluators actually preferred these AI-rendered interfaces to pages built by a leading coding model in a majority of matchups, both for how naturally objects behaved on screen and for how well the system followed instructions. It's not a finished product: text still gets garbled, long sessions drift off course, and the system can produce screens that look entirely convincing while being subtly wrong. But it's a real glimpse of an idea that's been circulating for a while β a web where the interface itself is generated on demand rather than coded in advance β and it landed at the exact moment video-generation models have gotten fast and cheap enough to make the idea plausible rather than absurd.
The second big story of the day cuts in a very different direction: teaching AI to fix its own safety problems, and finding it's dramatically better at that than the humans who normally do the work. In newly published research, teams of AI agents were set loose to independently research, train against, and score ten well-known categories of AI misbehavior β things like sycophancy, deceptive responses, and susceptibility to jailbreaks. Left to work through the problem on their own, the agents cut sycophancy by over a quarter and improved resistance to reward-hacking by nearly all of the way. Human experts given the identical task, working under the same conditions, made far less progress on the hardest category β deception β while the AI agents averaged a success rate roughly four times higher across more than 150 attempts. In a related experiment, a smaller, less capable model spent 60 hours training a more advanced pre-release model to behave better, using a small fraction of the data a full safety-training run would normally require. Whatever you make of AI writing its own safety curriculum, it's a concrete example of a shift a lot of people have predicted but few have actually seen happen yet: AI research starting to research itself.
Medicine got its own quietly significant win. Researchers built a model that reads a completely ordinary heart scan β the kind hospitals already run over a billion of every year β and flags heart failure and valve disease in under two seconds, catching cases that routinely go undetected until a patient's condition has already progressed. Confirming those diagnoses normally requires a specialist ultrasound and a months-long wait; this tool works directly off the scan that's already sitting in the file. Trained on more than ten million scans and tested against 65,000 patients, it correctly flagged heart failure in about four out of five cases and valve disease in nine out of ten. A multi-hospital trial is now underway, with an eye toward folding the tool into routine national health-service screening within two years β a reminder that some of the nearest-term, highest-value AI applications aren't flashy chatbots, but a second set of eyes on data that's already being collected.
On the money side, a startup that lets you search your own computer the way you'd search a chatbot β describing what you're looking for in plain language instead of remembering a filename β just raised at a quarter-billion-dollar valuation, with all the processing running locally by default so nothing has to leave the device. It's a small but telling sign of where investors think the next wave of everyday AI utility is: not another chat window, but AI quietly built into the tools people already use.