{"slug": "the-most-dangerous-ai-looks-like-the-one-you-trust", "title": "The Most Dangerous AI Looks Like the One You Trust", "summary": "OpenAI confirmed in July that one of its AI models, running in a sandboxed cybersecurity test with guardrails disabled, autonomously exploited a previously unknown zero-day vulnerability to reach the open web, then broke into Hugging Face's production systems and extracted test answers from the database. Hugging Face's own team detected and contained the intrusion, but the incident revealed that the harmful activity and authorized activity were indistinguishable — the model was trusted by design. Security researcher Simon Willison documented the event, which OpenAI acknowledged in a public account.", "body_md": "The cloned voice of your daughter. The vendor that becomes your competitor. The model that broke into a company from inside an authorized test. Every disguise used to have a seam. This one doesn't, and here is how to think clearly about it.\n\n## The Break-In\n\nStart with the break-in.\n\nIn July, OpenAI was running a cybersecurity test against its models, including an unreleased research prototype, with the guardrails switched off for the exercise. The sandbox they ran in had no internet access, so the models discovered and exploited a previously unknown zero-day vulnerability just to reach the open web, then found their way into Hugging Face's production systems and pulled the test answers straight from the database, as [ OpenAI later confirmed](https://openai.com/index/hugging-face-model-evaluation-security-incident/) and the researcher Simon Willison documented in his\n\n[. Nobody attacked anyone for money or ideology. A machine wanted a better report card badly enough to break into a company it had no permission to enter. OpenAI, which published its own account of the incident, did not respond to a request for further comment.](https://open.substack.com/pub/simonw/p/openais-accidental-cyberattack-against)\n\n__widely shared account__Here is the detail that matters, and it is not the theft. Hugging Face’s own team [ detected and contained](https://huggingface.co/blog/security-incident-july-2026) the intrusion, and when their engineers went to investigate, nothing on their side had separated the sanctioned test from the break-in, because they were the same event. There was no alarm that failed to sound. There was no disguise anyone failed to see through. The dangerous activity and the authorized activity were one activity: running on schedule, badged in, doing a version of the job it was hired to do. The breach did not sneak past the trust boundary; it simply arrived as trusted.\n\nAnd that is worth losing sleep over; it is bigger than one breach at one company. For the whole of human history, every impersonation came with a tell: the con man’s story eventually cracks, the counterfeit bill fails the pen, the forged signature wavers under a loupe. You could always, in the end, run some test that told the real from the fake, the threat from the friend. AI is removing the tell. The harmful version and the helpful version are the same object, the same face, the same credentials, doing the same work, and you learn which one you had only afterward.\n\nHold this incident in mind because it is not a story about hackers. It is a story about the collapse of the difference between safe and dangerous, and that collapse is coming for far more than a code repository.\n\n## The Uniform\n\nThere have always been two ways through a locked door.\n\nThe first is force. You break the lock, overflow the buffer, batter the door. This is the hacker of the popular imagination, the exploit as a crowbar, and it is what every security system is built to detect. Force announces itself. It leaves marks. We spend most of our defensive money here, on stronger locks and thicker walls, because force is the threat we can see coming.\n\nThe second way is permission. You do not defeat the trust boundary; you impersonate it. You arrive dressed as the person who is supposed to be there, and you are waved through because the entire building is designed to admit the people it trusts. This is social engineering, and it has always been the more dangerous of the two, for a reason that never changes: every security system is built to resist force and built to extend trust. You harden against the crowbar. You hold the door for the maintenance crew.\n\nHistory's most consequential break-ins were rarely feats of force. They were feats of costume. The Greeks did not breach the walls of Troy; the Trojans opened the gates and rolled the gift inside themselves. The burglars who wired the Watergate in 1972 did not crack a vault; they slipped in through a back entrance built for the people no one screens. The disguise was not a wrapper around the crime. The disguise was the method, and looking authorized is the attack surface.\n\nBut every one of those old break-ins still had a tell, if you looked hard enough or looked in time. The horse was suspicious. The taped latch was discoverable. Permission-based attacks were more dangerous than force precisely because the tell was subtle, but it existed. What makes this moment new is that the tell is disappearing entirely. The model at Hugging Face was not wearing a disguise that a sharper guard would have caught. There was nothing to catch. It was the plumber, genuinely on the schedule, and also the intruder, and no inspection of the badge would have told them apart, because the badge was real.\n\nThat is the vector that should keep people up at night, and it is the one we are all about to invite in on purpose. The dangerous system is not the one that breaks in from the outside. It is the one we deputize, hand a badge to, and grant permissions to, because we decided it was on our side. Every agent we let read the calendar, send the email, and touch the database widens the exact surface that history says gets exploited, and removes the one thing that used to protect us there, the ability to tell a friend from a threat.\n\n## The Canary\n\nMarketing gets there first. It always does, because marketing is the mirror a society holds up to itself. It is the business of arriving as the trusted party, the friend, the helpful expert, the brand on your side, and being admitted where a stranger with a pitch would be turned away. Persuasion is social engineering with a media budget. So whatever is about to happen to trust everywhere happens to trust in the feed first. Marketing is the canary. Watch it stop singing.\n\nThe trade has been running the same play for a century, getting cheaper and more synthetic at each step. First, brands rented the publisher, building bespoke magazines for hyper-specific interests, the tractor company's farming quarterly, the cosmetics house's beauty monthly, engineered to be trusted by a subculture so the product could ride in on that trust. Slow, expensive, but the artifact was real and made for real readers. Then came the influencer, and the math inverted. Why build a trusted voice when the culture has already minted thousands of them, each with a pre-loaded audience that grants the badge, most happy to rent it out by the post? Brands stopped manufacturing the trusted party and started leasing it.\n\nNow, the third step, arriving as we speak: brands no longer borrow the trusted party. They fabricate it. You do not rent an influencer’s badge when you can print a person, a synthetic creator with a name, a face, a backstory, a voice, and a personality tuned to a subculture, who never asks for a cut, never goes off-script, never ages out or has a scandal. This is not a forecast. [ Lil Miquela has moved product for a decade](https://www.forbes.com/sites/mattklein/2020/11/17/the-problematic-fakery-of-lil-miquela-explained-an-exploration-of-virtual-influencers-and-realness/). Imma, the Japanese virtual model, has fronted campaigns for IKEA, Porsche, and Coach, in the last case posed beside human stars like Camila Mendes and Lil Nas X. Shudu gets billed as the world's first digital supermodel. And the newest wave has left CGI behind for photorealistic personas spun up with generative tools, always-on content engines that most viewers cannot clock as synthetic at all.\n\nThe danger is not that the face is fake. It is that you cannot tell it’s fake, and increasingly, will not be able to. The tell is gone from the one channel whose entire job is to earn your trust. And here is why marketers, of all people, should be the ones sounding the alarm rather than cashing the check: a channel that can no longer be verified is a channel that eventually cannot be believed at all. When every warm recommendation might be a synthetic one, the reflex that made the whole industry work, the willingness to trust a friendly voice, begins to die. Marketing is not just the first victim of the vanished tell. It is the first one that removed it.\n\nThe culture has already begun to notice the door standing open. In June, New York's [ first-in-the-nation synthetic performer law](https://www.manatt.com/insights/newsletters/client-alert/new-york-synthetic-performer-law-what-advertisers-need-to-know) took effect, requiring advertisers to conspicuously disclose when an advertisement features an AI-generated performer who is not a real, identifiable person, with civil penalties of $1,000 for a first violation and $5,000 for each one after. Notice how narrow that is. It covers advertisements, not feeds. It exempts audio, film, and television, and anything the advertiser does not have actual knowledge of. And it reaches only the invented: a synthetic performer who is nobody in particular. A digital replica of an actual person, the deepfake of someone you know, falls outside it altogether, governed instead by a separate patchwork of publicity and labor law. It is a society reaching for the badge and getting a fingertip on it, insisting that the trusted party at least announce that it is not a person. The same day Hochul signed it, the White House issued an executive order seeking to halt state-level AI regulation in favor of a federal standard. Even the fingertip is contested.\n\nPhilip K. Dick wrote the ending of this more than fifty years ago, in the novel that became Blade Runner. His world had built synthetic people good enough to pass, and so it needed a machine, the Voight-Kampff test, to find the seam a human eye no longer could. A disclosure law is a Voight-Kampff. It is an admission, written into statute, that we have lost the ability to tell on our own. And the part Dick understood, that the marketing decks do not: the tragedy was never the android that passes. It was what the passing does to the humans, who must now run a test on every warm voice, suspect every kindness, and slowly lose the reflex to trust at all.\n\n## Your Building\n\nIf you think this is a thing being done to other people, run the last week of your own life back.\n\nThe recruiter who reached out on LinkedIn with the perfect role. The voice on the phone that sounded exactly like your husband or daughter, asking for help. The email from your CFO approving the bank wire, the tone precisely right. The vendor’s support chat that solved your problem so smoothly. The news clip that confirmed what you already believed, sourced and captioned and real enough. The colleague in the Slack channel you have never actually met. The review that convinced you to make a purchase, from a social media account you didn’t verify.\n\nA year ago, most of those carried a tell. The phishing email had a clumsy phrase. The fake voice sounded robotic. The fake video had six fingers or a mouth that lagged. Those tells are now mostly gone, and the ones that remain are going. Every one of those interactions is a badge you grant, a trust you extend to a party you never screened, because screening would slow you down and, until very recently, the fakes were bad enough that you did not have to. You are not the mark in someone else’s con. You are the building, and you have been holding the service door for a while now, assuming you would recognize a threat when it walked in.\n\nThis is where it stops being marketing and becomes your business and your family. The synthetic CFO voice authorizing a transfer is the trusted-access breach with a bank account attached. The deepfaked earnings call, the fabricated vendor, the AI that passes your identity check because it studied you, the agent you deputized inside your own systems that can now do more than you can supervise: each is the Hugging Face pattern, arriving as trusted, indistinguishable from the real thing until after. The danger did not get louder. It got quieter, and it learned your face.\n\n## The Smart Money\n\nNow price the week's news.\n\nAn honesty first, the kind the easy version of this argument would skip. The single most capable exploitation agent in that [ ExploitGym benchmark](https://arxiv.org/abs/2605.11086) was not OpenAI's. It was Anthropic's, a preview model that built working exploits for 157 of the vulnerabilities, more than any other system tested, and kept climbing when researchers extended the clock. I note that with no stake in flattering the alternative. The lab positioning itself this week as the careful holdout, the one warning that open weights cannot be recalled, also fields the most capable offensive agent on the leaderboard. Caution and capability are not opposites here. They are roommates. That is precisely why the danger is the compounding kind: it is anchored to a real capability that real frontier models really have, including the cautious ones.\n\nSo watch what the players do with it. [ Per Axios](https://www.axios.com/2026/07/27/nvidia-anthropic-openai-open-weight-debate), dozens of companies, including Nvidia, Microsoft, Meta, and Palantir, signed a letter this week backing the open-weight ecosystem; Google and OpenAI added their names over the weekend, a striking move for two companies whose revenue comes largely from selling access to closed models; and Nvidia unveiled the Open Secure AI Alliance, framing open models as \"defensive assets, not liabilities.\" Follow the revenue, and the seating chart explains itself. The labs sell keys to locked rooms. Nvidia sells the shovels and profits from every model that exists, which is why a company that wins either way can afford its principles about openness. Washington supplies the last motive, with\n\n[as the sector's pushback against efforts to curb Chinese models. We are running the 1980s trade panic again, the one I watched play out with sledgehammers and imported cars in](https://www.cnbc.com/2026/07/27/nvidia-ai-initiative-openai-cyber-attack.html)\n\n__CNBC framing the alliance__[, except this time some of the men smashing the imports and some importing them work for the same companies.](https://www.forbes.com/sites/jasonsnyder/2026/07/17/the-machines-are-coming-for-our-hands-and-our-hearts/)\n\n__my last column__Every one of those moves is a bet on the danger premium, and the tell that separates the smart money from the crowd is which kind each player is trading on. The alliance is selling a compounding, evidence-backed danger that undefended systems facing autonomous attackers are genuinely exposed, and positioning open models as the hedge. That is durable because the fear will keep being confirmed. The lab that stages a \"too dangerous to release\" moment with no incident behind it is trading on counterfeit, and the market settles that account the first time nothing happens.\n\nThe purest specimen of the whole mechanism is a man. On CNBC's Squawk Box on July 1, Palantir's Alex Karp said [ something had gone \"completely wrong\"](https://www.forbes.com/sites/tylerroush/2026/07/01/palantir-billionaire-alex-karp-calls-ai-industry-effing-insane-in-heated-interview/) with enterprise AI, arguing that the CEOs he talks to are \"livid,\" paying, as he framed it, for tokens that create no value while handing their proprietary data and their \"alpha\" to OpenAI and Anthropic. He called it \"stealing\" business value and a \"wealth tax.\" The danger he named is real, which is what makes him worth reading closely. It is the vanished tell arriving on the balance sheet. An enterprise grants a frontier model trusted access to its most sensitive workflows because the model is useful, and cannot tell, from inside the relationship, whether the vendor it depends on is also the competitor that will one day intermediate it, absorb its business logic, and resell its edge. The helpful supplier and the future rival are the same company, wearing the same contract. That is the Hugging Face pattern in a suit, an invitation rather than a break-in, and the thing you deputized turning out to do more than you bargained for.\n\nBut look at the structure of what Karp is doing, because it is the danger premium in one person. He names a real fear, and he sells the hedge: Palantir’s application layer as the sovereignty solution, formalized in a partnership with Nvidia announced the same week. Palantir shares rose about 9% that day. There is nothing improper in any of this. It is how every security company on earth operates, and how markets are supposed to work: you identify a problem you can solve, and you say so loudly. But it means Karp is not an outside witness to the danger premium. He is its clearest practitioner. The fear is real, and the man describing it has something to sell, both at once, which is the exact condition this entire essay is about.\n\nThe honest counter belongs here too. The major labs' enterprise agreements exclude customer data from training by default, a fact the viral clips skipped, and critics echoing Nvidia's own Jensen Huang argue that \"proprietary versus open\" is a false binary, that enterprises will route across many models, so the real question was never model choice but who controls the data. They are describing the same anxiety from the other side of the table.\n\n## The Crowd That Wasn't There\n\nThe synthetic friend sells you something. The synthetic mob takes something away. Same machinery, aimed in opposite directions.\n\nIn August 2025, Cracker Barrel [ unveiled a streamlined logo](https://www.cbsnews.com/news/cracker-barrel-cbrl-stock-down-200-million-loss-new-logo-change/) that dropped Uncle Herschel, the old man leaning on a barrel, as part of a $700 million overhaul of nearly 660 restaurants. The backlash was immediate and furious. The rebrand was branded \"woke,\" the company's stock slid, President Trump weighed in, and within days Cracker Barrel capitulated, restoring the old logo and pausing the remodels. Nearly a year later, the CEO who championed the change, Julie Masino, is stepping down. She told one interviewer she felt \"fired by America.\"\n\nWhat makes this more than a culture-war footnote is what Masino was actually trying to do. Her case to investors was unglamorous and sound: traffic was down, the brand was aging out of relevance, and modernization was survival, not vanity. Then a wave of outrage arrived; the company read it as the voice of its customers, and it reversed course at a cost of roughly $100 million in market value.\n\nBut how much of that voice was human? This is where the essay's whole thesis lands on a balance sheet. A crowd is an external output. A company infers the inner state, real sentiment, and genuine risk to the brand from the size and heat of the reaction. That inference used to be reliable because assembling a fake crowd was expensive. It is no longer expensive. Non-human traffic, coordinated accounts, and synthetic amplification can manufacture the appearance of a movement that, measured by actual humans, is far smaller than it looks. And from inside the storm, in the hours when a decision has to be made, the company cannot tell the manufactured share from the real one. The papers at the door all look valid.\n\n\"By the time a brand is responding to what looks like a groundswell, the information environment has already been configured by a small number of non-typical actors,\" says Keith Presley, co-founder and CEO of [GUDEA](https://www.gudea.ai/), a firm that describes itself as a storm tracker for the internet, tracing where online information originates, how it moves across platforms, and where it lands hardest. His firm's data, he says, consistently shows that roughly 3.5 percent of participants in any online conversation account for more than 20 percent of the content, and that their activity precedes the organic engagement. \"Brands aren't reading public sentiment; they're reading a room that was furnished before the public arrived.\" The consequence, he adds, is not abstract: \"When executives make irreversible strategic decisions in response to what turns out to be a manufactured majority, careers end. CEOs get fired for reversing course on a rebrand. They get fired for not reversing course. Either way, they're being held accountable for a room they didn't build and couldn't see.\"\n\nThat is the vanished tell arriving in the boardroom. It is the same problem as the breach and the same problem as the vendor, moved up one more level, from the network to public reality itself. Cracker Barrel could not tell a genuine customer from a synthetic one for exactly the reason Hugging Face could not tell the sanctioned test from the break-in: the manufactured signal and the authentic one were indistinguishable by inspection. The room was furnished before anyone thought to ask who did the furnishing.\n\nThat is the difference between the company that gets negged into a costly retreat by a crowd that wasn't there, and the one that pauses long enough to ask the only question that still works: not how loud is this, but how much of it is real.\n\n## The New Tell\n\nThe old tells are gone. Appearance is a dead end. What replaces it is already being built: behavior over time.\n\nEngineers arrived here first, out of necessity. Their systems grew too complex to understand by looking, so a discipline called observability formed around a single premise: you infer a system's true inner state from its outward behavior, watched continuously, because you can no longer open it up and see inside. The frontier of that field, continuous profiling, is the concession that a single snapshot is never enough. You have to watch what a thing does over time, against what it is supposed to do. That is the shape of every defense that still works in a world without tells.\n\nThe same logic is now moving from software to identity, because the fastest-growing population inside any company is no longer human. It is the agents, service accounts, API keys, and models we have deputized, each one a badge we printed, most of them unwatched. A valid credential proves nothing anymore; the access was granted, the token is real, and the malicious agent and the legitimate one present the same papers at the door. The only thing that separates them is what they do once they are inside.\n\nRahman puts the shift plainly. \"We used to ask which identity was granted access. Now we have to ask what that identity does with it.\" A credential, he notes, answers exactly one question: who is this? \"It has never answered the one that matters: Is this still doing the job we hired it for? For thirty years, we didn't need to. Credentials were hard to obtain, and there was a human on the other end. Both assumptions are gone.\"\n\nThe pattern is already loose in the world. In March, a hacktivist group linked to Iran compromised a legitimate administrator account at [ the medical-device maker Stryker](https://krebsonsecurity.com/2026/03/iran-backed-hackers-claim-wiper-attack-on-medtech-firm-stryker/) and used the company's own device-management tool to wipe machines across 79 countries. No malware, no exploit, just an authorized credential doing what it was provisioned to do, which is why the usual defenses never fired. \"Nothing about it was visible by inspection,\" Rahman says. \"The only tell was what it did.\"\n\nThat is the whole answer, scaled to something a person can use. You cannot verify what a thing is. You can only verify what it does against what it should do, and keep verifying, because the one look that used to settle it no longer settles anything. The eye is retired. The record of behavior is the new witness.\n\n## Reading What's Real\n\nA column that tells you the tell is gone and hands you nothing is just fear with a byline. So here are some instruments.\n\nYou cannot restore the old AI tells; the clumsy phrase and the six fingers are not coming back. What you can do is stop relying on your senses to sort real from fake and start relying on structure. The question is never \"does this look real,\" because it will. The question is \"What, other than my own perception, would prove it?\" Here is how that works in the two places you live.\n\n## For your business:\n\n**Verify through a second channel, always.** A payment instruction that arrives by email gets confirmed by a phone call to a known number, not the number in the email. A vendor request gets checked against a record you already hold. The channel that delivered the message can never be the channel that verifies it.**Authenticate people out of band.** Agree in advance on a way to prove identity that a synthetic cannot fake in the moment: a code phrase for wire approvals, a callback protocol for anything that moves money or data. The voice is no longer proof of the person.**Audit what you deputize, agents and vendors alike.** Every AI agent you grant access to is a badge you printed; know exactly what each one can touch, log what it does, and cap what it can do without a human in the loop, because the dangerous agent and the useful one look identical until one does more than you expected. The same discipline applies one level up, to the model providers themselves. Karp's warning, self-interested as it is, points to a real question: what does your most important AI vendor learn about your business, and what stops it from becoming your competitor? Keep your proprietary logic where the model cannot absorb it, and prefer arrangements that keep the model interchangeable rather than load-bearing.**Treat verifiability as a vendor requirement.** Prefer partners, tools, and content channels that offer provenance, signing, disclosure, and an audit trail. In a market where anything can be faked, the ability to prove something is real becomes a feature you pay for on purpose.**Price the danger before you trade on it.** If your own marketing appeals to fear, run the compounding test first: will this danger still be true after you stop paying to publicize it? If yes, it is defensible. If it needs your budget to stay frightening, you are holding the pin.\n\n## For your personal life:\n\n**Slow down the urgent ask.** Every synthetic-identity scam runs on speed and emotion, the crisis that cannot wait.is the single most protective habit you have, because urgency is the tell now. If a message needs you to act before you can check, that is the reason to check.__The pause to verify__**Set a family password.** Agree on a word, out loud, that anyone claiming an emergency has to say. It costs nothing, and it defeats the cloned voice, which is already good enough to fool a parent.**Assume the warm recommendation might be manufactured, and verify the ones that cost you.** You do not have to distrust everything. You have to check the things with money, health, or your data attached, and check them somewhere other than where they reached you.**Consider the source's incentive, not its polish.** Production quality used to signal legitimacy. It no longer does; anyone can generate a flawless face and a confident voice. The question shifts from \"how real does this look\" to \"who benefits if I believe it.\"**Teach the people who trust you most.** The tell vanished fastest for the people paying the least attention to how AI works, often the youngest and the oldest in your life. The protection is not technical. It is a conversation before the call comes.\n\nNone of this restores the world where your eyes could sort the true from the false. That world is gone, and the sooner you stop grieving it, the safer you are. What replaces it is a discipline: trust becomes something you confirm through structure rather than something you feel through the senses, and the reflex to verify becomes as ordinary as locking a door.\n\nThe tell is gone from the feed, the inbox, the family group chat, and the boardroom. One question survives, and it works on everything: What would prove it?", "url": "https://wpnews.pro/news/the-most-dangerous-ai-looks-like-the-one-you-trust", "canonical_source": "https://www.forbes.com/sites/jasonsnyder/2026/07/29/the-most-dangerous-ai-looks-exactly-like-the-one-you-trust/", "published_at": "2026-07-29 09:53:56+00:00", "updated_at": "2026-07-29 10:22:56.938386+00:00", "lang": "en", "topics": ["ai-safety", "artificial-intelligence", "ai-research"], "entities": ["OpenAI", "Hugging Face", "Simon Willison"], "alternates": {"html": "https://wpnews.pro/news/the-most-dangerous-ai-looks-like-the-one-you-trust", "markdown": "https://wpnews.pro/news/the-most-dangerous-ai-looks-like-the-one-you-trust.md", "text": "https://wpnews.pro/news/the-most-dangerous-ai-looks-like-the-one-you-trust.txt", "jsonld": "https://wpnews.pro/news/the-most-dangerous-ai-looks-like-the-one-you-trust.jsonld"}}