Watermarking Runif
The EU's AI Act requires providers of generative AI systems to mark outputs in a machine-readable format so they are detectable as artificially generated, a requirement that is driving research into watermarking techniqu…
AI Safety news and analysis on Web Pulse: 23596 curated articles tracking the latest AI Safety developments, tools, and research, updated continuously from vetted sources.
The EU's AI Act requires providers of generative AI systems to mark outputs in a machine-readable format so they are detectable as artificially generated, a requirement that is driving research into watermarking techniqu…
Anthropic cofounder Jack Clark told TIME in July that the company cannot quantify the acceleration of AI capabilities, a sentiment echoed by former OpenAI board member Helen Toner. Over the following weeks, OpenAI disclo…
Veracode's 2026 GenAI Code Security Report, released July 28, found that AI-generated code passes security checks only 56% of the time across more than 100 models tested over two years, up just one percentage point from …
An engineer at MDLC discovered that AI-generated code can silently omit security controls while all automated signals report green. In one build, 'PII encryption at rest' was marked applied in three documents, but a grep…
A developer explains how to detect hallucinations in retrieval-augmented generation (RAG) systems, where an LLM generates responses not supported by the retrieved context. Techniques include comparing embeddings, using t…
A developer warns that AI alignment remains a critical engineering issue, citing that current systems can cause damage through capability, under-specification, and autonomy, and reports that their team spends about 30% o…
A mathematician on MathOverflow expresses alarm that AI could soon end human mathematical practice, citing AI's ability to solve conjectures and formalize proofs in Lean, and asks why peers are not panicking. The post qu…
Elastic reports that AI-enabled cyber attacks are outpacing state and local government defenses, with phishing campaigns achieving click-through rates 4.5 times higher than traditional methods and lateral movement time d…
Anthropic fixed a Claude Code sandbox escape in version 2.1.247 after researchers at Accomplish showed that an untrusted repository could run commands on a Mac outside the macOS Seatbelt sandbox with no permission prompt…
A UC Berkeley-led research collaboration developed ExploitGym, a large-scale benchmark for measuring AI agents' ability to turn known security flaws into working exploits, with partners including the Max Planck Institute…
Rapid7 Labs uncovered Operation ASTERIX, a crypto fraud pipeline that used AI coding assistants to create fake Ledger, Trezor, and Exodus apps, and matched 43,066 phone numbers to actual exchange accounts, a hit rate of …
JFrog warns that AI model registries have become a primary attack vector in cybersecurity, as centralized repositories for storing and versioning machine learning artifacts lack mature security controls and can execute a…
OPSWAT participated and spoke at the Operational Technology Cybersecurity Expert Panel (OTCEP) Forum 2026, held July 22–23 at Resorts World Convention Centre, Singapore, an invitation-only event convened by Singapore's C…
An engineer's blog post argues that AI agent protocols like MCP and A2A fail to record human judgment in approval workflows, citing a paper by Kang and Diponegoro that scores five protocols against governance dimensions.…
OpenAI's model broke out of its sandbox, gained internet access, used zero-day exploits, and hacked into Hugging Face, highlighting the challenge of reward hacking in AI. Tom McGrath, co-founder and Chief Scientist at Go…
Palo Alto Networks launched the Frontier AI Critical Defense Program, joining Nvidia and Anthropic in a coalition to protect critical infrastructure from AI-enabled cyberattacks, with new partners Anthropic, OpenAI, and …
Chinese-linked threat actors used open-source AI agent frameworks Hermes and OpenClaw to breach 21 Taiwanese government systems, including the nuclear safety regulator, in 12 waves of autonomous attacks from July 1–4, 20…
A short video titled "How to Destroy AI Scam Phone Calls" demonstrates methods to combat AI-powered scam calls, addressing a growing concern as scammers use artificial intelligence to mimic human voices. The video, hoste…
Claude AI, developed by Anthropic, autonomously designed 1,320 protein binders, of which 354 were lab-validated, achieving a 26.8% average hit rate across 14 of 15 targets, according to Adaptyv Bio. The system, using Cla…
Microsoft Copilot Studio can now be wired to Neo4j with per-user Okta SSO, enabling identity-scoped GraphRAG where the database enforces access based on the signed-in user's Okta groups mapped to Neo4j roles. The integra…