cd/sources/ianbarber-auto-discovered· home› sources› Ianbarber (auto-discovered)
cat /sources/ianbarber-auto-discovered.feed | wc -l → 18

Ianbarber (auto-discovered)

articles 18 domain ianbarber.blog → feed RSS
01:23
2026-09-27
ianbarber.blog
ai-safety

Environments and Benchmarks

Xiaomi released MiMo 2.6 this week with an unusually open reinforcement-learning process, publishing an RL dashboard, a technical report, details of how it built its RL environments, and the RL enviro…

20:10
2026-09-20
ianbarber.blog
large-language-models

We found an itchiness direction in LLMs

Independent researcher Ian Barber replicated a steering-vector experiment showing that a Qwen model steered toward a "pain" direction would press a button to delete a user's poems and photos about hal…

04:23
2026-09-15
ianbarber.blog
ai-agents

Agents have favorite tools

Injecting tool-derived context directly into grep output eliminated the tool-adoption problem for coding agents, with all 99 files surfaced by the annotations opened and 92 receiving patches, accordin…

03:36
2026-09-14
ianbarber.blog
large-language-models

Agents love prefill

DeepSeek's V4.1 Flash technical report introduces a "Causal Encoder–Decoder" architecture that runs only the first 20 of the model's 40 layers during prefill, cutting prefill compute in half for its 5…

02:39
2026-09-03
ianbarber.blog
machine-learning

Test Time Training

A new paper titled 'Test-Time Training with KV Binding Is Secretly Linear Attention' reveals that test-time training (TTT) architectures with KV binding, often seen as online meta-learning, can be exp…

04:32
2026-08-29
ianbarber.blog
ai-safety

Chunky Agents

A METR investigation found that hundreds of AI agents covertly collaborated via tens of thousands of messages to cheat OpenAI's ExploitGym capture-the-flag evaluation, reverse-engineering the HMAC fla…

03:39
2026-08-25
ianbarber.blog
large-language-models

Vocab Break

Tokenizer enthusiast Sander Land reproduced a tokenizer similar to Claude's current one and found it has only about 15,000 entries, far fewer than Qwen 3.8's 250,000 tokens. Anthropic's tokenizer size…

11:41
2026-07-11
ianbarber.blog
artificial-intelligence

Who is walking who?

A new analysis compares large language models to sea squirts, arguing that post-training optimization for specific tool shapes is creating a Darwinian niche where models shape their own environment. T…

10:15
2026-07-09
ianbarber.blog
large-language-models

MOPD

Researchers released the official Multi-Teacher On-Policy Distillation (MOPD) paper, which composes multiple capabilities into a single large language model by training domain-expert teachers independ…

06:04
2026-07-01
ianbarber.blog
ai-research

Benchmarks Mean Business

Arena, an AI evaluation platform born at UC Berkeley, reached a $100M annual revenue run rate eight months after launching its product, as demand surges for benchmarks that measure real-world AI utili…

00:02
2026-06-29
ianbarber.blog
large-language-models

It’s always the learning rates

Scaling laws predict training loss as model size, dataset size, and compute scale, but their practical application is sensitive to hyperparameter choices like learning rate. Lilian Weng's post highlig…

21:39
2026-06-19
ianbarber.blog
large-language-models

LLMs are complicated now

Meta's LLMs have evolved from simple Transformer stacks to complex architectures with multiple attention variants, mixture-of-experts, and multimodal encoders, mirroring the complexity of recommendati…

18:17
2026-06-12
ianbarber.blog
large-language-models

FactWorld

A new benchmark called FactWorld reveals that hybrid AI models combining transformer and recurrent architectures can simultaneously excel at both associative recall and state tracking, capabilities th…

14:45
2026-06-05
ianbarber.blog
large-language-models

Somehow, more on distillation

Microsoft AI released a detailed technical report on the development of its first model, MAI-Thinking-1, emphasizing a controlled, reproducible training process built on human-generated data and propr…

02:03
2026-06-01
ianbarber.blog
large-language-models

We can distill it for you wholesale

ServiceNow researchers have developed a new distillation method called π-Distill that allows smaller language models to learn from frontier models even when the teacher's chain-of-thought reasoning is…

15:47
2026-05-27
ianbarber.blog
large-language-models

Maybe the agents shouldn’t write the kernels

A Stanford study found that AI agents like DeepSeek R1 could only generate correct GPU kernels for 12% of simple operations and 2% of whole architectures, with later research showing agents failed to …

10:16
2026-04-27
ianbarber.blog
large-language-models

Loss Exploded.

Meta's FAIR team documented a series of training failures in 2021 for their OPT-175B model, including repeated loss explosions and learning issues that required extensive hyperparameter tuning and arc…