AI’s Top Labs Are Breaking Science—And They Know It More than half of AI unicorns have never led a single scientific publication, according to a peer-reviewed preprint posted July 16 on bioRxiv and covered by Science/AAAS, with lead author John Ioannidis—the Stanford metascientist who exposed Theranos—finding that these billion-dollar companies account for just one in every 1,000 AI papers published in 2025. The study examined 317 unicorn AI companies from 1998 to 2025 and found that over half had no first or last author publications, mirroring the pattern of unverifiable claims that led to Theranos's collapse, while safety justifications for secrecy are undermined by the fact that safety claims themselves require scientific verification. Meanwhile, the Stanford HAI Foundation Model Transparency Index average score fell from 58 in 2024 to 40 in 2025, with 84% of the 95 most notable models launched in 2025 shipping without training code, and Chinese labs like Moonshot AI producing more verifiable science than their US counterparts. A peer-reviewed preprint posted July 16 on bioRxiv—covered this week by Science/AAAS—contains a statistic that should unsettle every developer building on top of foundation models: more than half of AI unicorns have never led a single scientific publication. Collectively, these billion-dollar companies account for just one in every 1,000 AI papers https://www.science.org/content/article/ai-s-top-startups-are-barely-publishing-their-research published in 2025. The lead author is John Ioannidis, the Stanford metascientist who publicly called out Theranos for the exact same practice a decade ago. If you don’t know how that story ended, look it up. The Theranos Pattern, at Scale Ioannidis examined 317 unicorn AI companies from 1998 to 2025, searching for publications where a company researcher held a first or last author position—any leading role in the scientific record. More than half had none. This is not a case of busy startups prioritizing shipping. These are companies with multi-billion-dollar valuations, thousands of engineers, and research teams that publish blog posts, technical reports, and model cards instead of peer-reviewed science. Theranos raised $900 million on claims that turned out to be unverifiable because there was no published methodology to verify. OpenAI, Anthropic, and their cohort operate at orders of magnitude larger scale and are embedding their models into critical infrastructure, enterprise software, and developer tooling worldwide. The difference is scope. The pattern—extraordinary claims, zero documentation—is identical. Safety as an Excuse The labs have a ready answer for this: safety. Publishing methodology creates misuse vectors. Open research arms bad actors. Secrecy is a feature, not a bug. This argument has a fatal flaw. Safety claims are themselves scientific claims. When a lab says its model is “safe for deployment” or “below a threshold of concern,” that statement is not just marketing—it’s an assertion about real-world behavior that affects what developers build and what users trust. A paper published at NeurIPS this year called it “evidential inversion” https://arxiv.org/html/2605.08192v1 : the labs are withholding the very artifacts needed to evaluate the claims they’re making. You cannot simultaneously say “trust our safety evaluation” and “you cannot see our safety evaluation.” That’s not a safety posture. That’s a press release. The 2026 International AI Safety Report, led by Yoshua Bengio, put it plainly: “reliable pre-deployment safety testing has become harder to conduct” and “models now distinguish test from deployment contexts.” These aren’t open-source advocates saying this. This is the scientific consensus from the people most concerned about AI risk. The Numbers Say the Quiet Part Out Loud Stanford HAI’s Foundation Model Transparency Index tracks how much major AI developers disclose. After rising from 37 to 58 between 2023 and 2024, the average score fell to 40 in 2025 https://hai.stanford.edu/news/transparency-in-ai-is-on-the-decline —an 18-point drop in a single year. Meta’s score went from 60 to 31. Mistral’s from 55 to 18. Google, Anthropic, and OpenAI have all stopped disclosing dataset sizes and training durations for their latest flagship models. Of the 95 most notable models launched in 2025, 84% shipped without training code. Stanford’s summary is precise and brutal: “the most capable models now disclose the least.” As these models become more capable, the scientific record explaining how they work is getting thinner, not thicker. China Is More Scientifically Open Than Silicon Valley Here is the sharpest irony of the current moment. The labs most often cited as geopolitical risks—the Chinese ones—are producing more verifiable AI science than their US counterparts. Moonshot AI released Kimi K3 https://the-decoder.com/moonshot-ai-releases-kimi-k3-open-weights-and-infrastructure-after-shaking-up-the-frontier-model-race/ on July 27: a 2.8 trillion parameter mixture-of-experts model with open weights on Hugging Face, open-sourced infrastructure including high-performance attention kernels, and published methodology. It currently ranks third globally on Artificial Analysis’s Intelligence Index, behind only Claude Fable 5 and GPT-5.6 Sol. The gap between open-weight models and closed frontier models has compressed from six to nine months to three to five months. Meanwhile, the US labs calling open weights “dangerous” continue publishing blog posts and capability demos while the Chinese labs they’re worried about release the actual science. The argument that openness is a risk when your competitors are being more open than you is an argument that only protects your moat, not the world. What Developers Should Actually Demand The practical implication is uncomfortable: if you’re building on a closed foundation model today, you are building on an unaudited claim. Not an evaluated system—a claim. The safety documentation in a model’s terms of service is not engineering. The capability benchmarks published by the lab that created the model are not independent verification. The minimum bar for responsible deployment should include published methodology not a technical report written by the lab’s own marketing team , reproducible safety evaluations, and third-party audits with actual access to training artifacts—not just inference-time access to the finished model. Congress agrees on the direction: the AI Foundation Model Transparency Act of 2026 https://www.congress.gov/bill/119th-congress/house-bill/8094/text/ih is moving through the House, and NeurIPS has now established reproducibility as an official conference track. The field is slowly recognizing that “trust us” is not a methodology. Open weights alone are not enough either. A model can release weights while withholding training data provenance, RLHF methodology, and safety evaluation procedures. Open weights plus published science is the standard worth demanding. Right now, almost no one meets it. As ByteIota covered last week, even Anthropic’s open-weights position comes with significant caveats https://byteiota.com/anthropics-open-weights-position-not-a-ban-but-a-catch/ —a reminder that “open” is doing a lot of work in a sentence that still protects the moat. Key Takeaways - More than half of AI unicorns have never led a peer-reviewed publication; they collectively authored 0.1% of AI papers in 2025. - Safety claims require reproducible evidence—labs that withhold proof of their own safety assertions are making a marketing argument, not a scientific one. - AI transparency scores fell 18 points in 2025; 84% of top models shipped without training code. - Chinese labs Moonshot, DeepSeek are releasing more verifiable science than the US labs that cite them as dangers. - Developers should treat closed model safety claims as unverified assertions and push for published methodology and independent audits as table stakes.