First Steps Toward Automated AI Research
Recursive's automated AI research system achieved state-of-the-art results on three benchmarks: fixed-budget language model training, small-model training speed, and GPU kernel optimization. The syste…
Recursive's automated AI research system achieved state-of-the-art results on three benchmarks: fixed-budget language model training, small-model training speed, and GPU kernel optimization. The syste…
A developer explains how the Transformer architecture works, from self-attention to modern LLMs. The key innovation is that Transformers compare tokens directly via attention rather than processing se…
Researchers developed a Transformer-based scheduling policy for the open shop scheduling problem (OSSP) using deep reinforcement learning. The model, trained on small Taillard benchmark instances, gen…
A developer argues that the current AI revolution is fundamentally different from past waves of enthusiasm, citing the convergence of large-scale labeled data, GPU computing, and deep network architec…
Jürgen Schmidhuber, a pioneer in artificial intelligence whose lab developed foundational ideas like LSTM, world models, and artificial curiosity decades before they became mainstream, argued in a new…
NVIDIA released Nemotron 3 Ultra, a 550-billion-parameter open Mixture-of-Experts model designed for long-running AI agents that plan, use tools, and reason across many turns. The model uses a hybrid …
A developer has demonstrated that the 1982 Hopfield associative memory update rule and the 2017 Transformer scaled dot-product attention mechanism are mathematically identical operations, with one equ…
Michigan Senate candidate Mallory McMorrow unveiled a detailed artificial intelligence policy agenda last week, proposing an AI Workforce Reinvestment Fund and a "token tax" on companies using AI to f…
A systematic study of query, key, and value (QKV) projection variants in Transformers found that sharing the key and value projections (Q-K=V) performs on par with or better than the standard three-pr…
Five foundational research papers explain how large language models work, covering the Transformer architecture, in-context learning, scaling laws, and instruction tuning with human feedback. The pape…
Researchers have developed TopoMamSurv, a Graph Mamba survival analysis framework that uses topology-aware ordering to address computational bottlenecks in Whole Slide Image analysis. The framework in…
A study evaluating encoder-only Transformer and LSTM frameworks for streamflow prediction in ungauged basins found that the LSTM outperformed the Transformer across both upstream-only and combined con…
Researchers introduced Geometry-Aware Tabular Diffusion (GATD), a method that improves tabular data synthesis by feeding pairwise angles and lengths from column value differences into diffusion denois…
Anthropic employees have donated more than $880,000 directly to political campaigns favoring stricter AI regulations, with OpenAI staff contributing an additional $300,000, according to FEC data analy…
Kog Team researchers introduced Delayed Tensor Parallelism (DTP), a Transformer architecture that hides communication overhead behind computation and weight streaming to accelerate batch-size-one infe…
Researchers have developed a new interpretation method for Transformer models that use heterogenous attention structures, which process information from different sources. The approach addresses chall…
Sreeni Ramadorai draws a parallel between Transformer models' residual connections and Agile software development sprints, arguing both systems preserve foundational knowledge while incrementally addi…
Prompt caching can reduce LLM input costs by 50–90% and improve time-to-first-token by 3–10× without quality loss, according to a 2026 developer guide. The optimization stems directly from Transformer…
Researchers developed HRVConformer, a deep learning architecture that classifies neonatal hypoxic-ischemic encephalopathy (HIE) by directly processing raw heart rate signals, achieving an AUC of 83.23…
A developer built a Transformer-based time-series forecaster using PyTorch to predict blood glucose levels 30 minutes in advance from Continuous Glucose Monitoring (CGM) data. The model, which uses se…