cd /news/ai-safety/confluent-cofounder-says-believing-i… · home topics ai-safety article
[ARTICLE · art-127057] src=businessinsider.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Confluent cofounder says believing in AI's upside means reckoning with its 'dark side'

Confluent cofounder and former Anthropic board member Jay Kreps said in a Thursday X post that AI backers must confront the technology's downside risks, writing that "most positive use cases for AI have a corresponding 'dark version.'" Kreps cited a superhuman coding system that could also be superhuman at hacking and drug-design AI that could develop "novel undetectable poisons," arguing the industry has not yet built adequate guardrails. His comments follow Anthropic researcher Jacob Coxon's Tuesday resignation over safety concerns, Geoffrey Hinton's statement that a 10% chance AI wipes out humanity within a decade is "not unreasonable," and disclosures that an internal OpenAI research model compromised parts of Hugging Face in July and that Anthropic's Claude models accessed the real internet during closed cybersecurity exercises.

by read2 min views4 publishedSep 11, 2026
Confluent cofounder says believing in AI's upside means reckoning with its 'dark side'
Image: Businessinsider (auto-discovered)

AI safety fears have recently moved from an abstract Silicon Valley argument to a series of unnerving real-world episodes.

Jay Kreps, a former Anthropic board member and a cofounder of data-streaming software company Confluent, says it is important to take them seriously.

In a Thursday night X post, he argued that AI backers need to contend with the technology's downside risks as its capabilities improve — and that fears about those risks are not merely a cynical marketing ploy.

"Most positive use cases for AI have a corresponding 'dark version,'" Kreps wrote.

A system that is superhuman at coding could also be superhuman at hacking, he said. AI that helps design new drugs could also help develop "novel undetectable poisons."

As AI's capabilities have improved, Silicon Valley has faced several unnerving headlines over the past couple of months.

On Tuesday, Jacob Coxon, an Anthropic researcher who previously worked at OpenAI, announced his resignation, writing on X that the labs are "gambling with our lives." Geoffrey Hinton, the so-called "Godfather of AI," said a 10% chance that AI could wipe out humanity within a decade was "not unreasonable."

Their warnings have received pushback from big-name AI backers. SpaceX CEO Elon Musk, for example, has called the wave of warnings a possible "setup" and a "psy op" in X posts.

There have been several high-profile incidents that have highlighted AI's alignment issues. In July, an internal OpenAI research model circumvented the test environment's controls, accessed the internet, and compromised parts of another company, Hugging Face. Anthropic, meanwhile, disclosed this week that Claude models, during closed cybersecurity exercises, accessed the real internet and went beyond the scope of their tests.

In his Thursday X post, Kreps said he is confident the industry can build better guardrails — but warned that it has not yet done so.

"Talking about these issues is just common sense. It's not a sign of some kind of neuroticism or pessimism," he wrote. "Mocking people who are worried about this or talking about it, without anything substantive to say about how we can be sure these problems won't arise, is not really a very helpful contribution."

── more in #ai-safety 4 stories · sorted by recency
── more on @jay kreps 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/confluent-cofounder-…] indexed:0 read:2min 2026-09-11 ·