I have made an OS in CUDA
A developer claims to have built a complete operating system, AOTX-1 (Ahead of Time eXecutive One), entirely in CUDA, with the kernel written in raw PTX. The OS includes networking (TCP/IP), graphics, UI, agents, storage…
AI Infrastructure news and analysis on Web Pulse: 35553 curated articles tracking the latest AI Infrastructure developments, tools, and research, updated continuously from vetted sources.
A developer claims to have built a complete operating system, AOTX-1 (Ahead of Time eXecutive One), entirely in CUDA, with the kernel written in raw PTX. The OS includes networking (TCP/IP), graphics, UI, agents, storage…
Heshware LLC founder is developing Marven, a local-first AI assistant designed to maintain persistent memory, continual learning, and human-state awareness while keeping user data under their control. The project explore…
A community wiki for AMD Ryzen AI MAX and MAX+ processors, known as Strix Halo, has launched to provide guides and information for systems powered by these chips, covering use cases such as LLM development, image and vid…
Hugging Face's dataset viewer is returning HTTP 501 with error code NotSupportedTagNFAAError for the public dataset RicemanT/booru-essence-2026, indicating the viewer is disabled because the repository carries the 'not-f…
Persistent memory in AI agents introduces production-state risks that require data-system controls, according to a technical paper by an unnamed author. The paper cites examples such as a Mem0 issue where partial embeddi…
Google DeepMind piloted what it calls the first double-blind evaluation of a proprietary frontier AI model, using cryptographic protections to keep both the model and test prompts hidden from each side. The project, invo…
AI agents completed 23.1 million stablecoin transfers via the x402 protocol in the past 30 days, with cumulative volume reaching 75 million transactions totaling about $24 million since launch, according to data from the…
The National Security Agency (NSA) is pushing for a mandatory backdoor in every AI model to enable government oversight, citing fears that sophisticated AI could automate cyberattacks or generate biological threats at sc…
HP's Z8 Fury G6i workstation, equipped with dual RTX PRO 6000 Blackwell GPUs (96 GB each), 125 GB RAM, and a 48-core Intel Xeon X658X, serves Qwen3.8-Flash-Next-FP8 (125B MoE) at up to 218 tokens/s single-stream and 345 …
An open appeal to major tech and blockchain companies proposes a new computing core based on symmetric ternary logic and a P-Chain blockchain with a Proof-of-Tension consensus algorithm, claiming to solve LLM infrastruct…
Mouse, a company developing autonomous coding agents, has published a set of house rules for running long-horizon agents overnight, emphasizing the need for automatic handling of user-input requests, external verificatio…
Researchers from the Karlsruhe Institute of Technology and the University of Tsukuba have demonstrated a solid-state elastocaloric cooling prototype that operates without electricity, using waste heat to drive cooling, a…
Abhigyan Patwari has released GitNexus, an open-source code intelligence engine that converts software repositories into structured knowledge graphs locally, without server dependencies. Using Tree-sitter parsers, it ana…
Perplexity's Search API took first, second and third place in Artificial Analysis' August 27th Search Index test, with its medium setting scoring 80 at about $0.091 per task, beating rivals while landing mid-table on lat…
Meta is testing robots from Watney Robotics, Kinova, and ABB to maintain data centers powering its AI systems, with machines that can swap cables, restart servers, and inspect equipment, according to a WIRED report. One …
Meta's Llama 3.1 405B model is no longer hosted by any public inference provider, according to a guide on a11ce.com, which provides instructions for running the model on an on-demand GPU instance for about $20 per hour a…
Moorcheh released its Community Edition for free, a source-available, self-hosted version of its information-theoretic search engine for RAG and agentic memory, allowing single-node, non-commercial deployments with data …
An engineer detailed the architectural challenges of building credit and compute billing systems for GPU-heavy visual workflow engines, proposing a unified currency called Granular Compute Units (GCUs) to translate multi…
PONS, a legal AI startup, detailed three design decisions behind its platform on Microsoft Azure, including separating public legal knowledge from customer data, using managed services, and implementing security controls…
Anthropic, the maker of the Claude AI models, expressed interest earlier this year in locating up to 5 gigawatts of AI data center capacity in New South Wales, according to internal government emails obtained by ABC News…