cd/entity/HumanEval· home› entities› HumanEval
grep -l @humaneval /news/*.json | wc -l → 57

HumanEval

mentions 57 type Organization page 1/3 feed RSS

// recent coverage 57 mentions

04:07
2026-10-07
arxiv.org
ai-safety

Backdooring Sparse Autoencoders

A paper submitted to arXiv on 5 October 2026 introduces a decoder-only sparse autoencoder (SAE) backdoor that induces attacker-chosen behavior when the modified SAE is inserted into the forward pass o…

04:00
2026-10-05
arxiv.org
large-language-models

Large Language Continuous Diffusion Models

Researchers introduced Sigma, described as the first large-scale continuous diffusion language model at 3B and 8B parameters, built on steerable low-dimensional ODE/SDE latent trajectories and trained…

02:48
2026-10-04
arxiv.org
machine-learning

Decoding Looped Transformers Better for Almost Free

Researchers posted a paper on arXiv on 1 October 2026 introducing LoopCD, a training-free contrastive decoding framework for looped Transformers that contrasts the final prediction with an earlier rec…

06:04
2026-10-01
dev.to
large-language-models

The cost of proving it works

Arize's 2026 cost model breaks production evaluation spend into traffic volume × sampling rate × evaluation surfaces × evaluator cost plus human review and retention, while offline evaluation is datas…

00:00
2026-09-13
mindstudio.ai
large-language-models

Edge0-35B Benchmarks: What 4-bit Quantization Really Costs You

Edge0-35B-A3B, a 35B-parameter sparse mixture-of-experts model built on Qwen3.5-MoE, loses an average of 3.9 points across five benchmarks when quantized to 4-bit through its edge0 inference pipeline …

14:28
2026-08-28
frontierroles.com
artificial-intelligence

Solutions Architect (6812) — MetroStar

MetroStar is hiring a Solutions Architect in Indianapolis, IN, offering $78,000–$115,000 per year to support secure, scalable cloud architectures that modernize legacy Marine Corps systems and enable …

18:54
2026-08-24
spock.is
artificial-intelligence

Simulating cosmic rays to lobotomize LLMs

A new experiment by an unnamed researcher found that flipping a single bit in the Qwen2.5-Coder-3B large language model can reduce its coding accuracy from 85% to near zero, simulating the effect of c…

01:00
2026-08-22
smarterarticles.co.uk
artificial-intelligence

95% Solved: Why AI Code Still Ignores Your Instructions

A team led by Ming Zhong at the University of Illinois Urbana-Champaign and Google DeepMind has introduced SWE-IF, a framework that aligns code evaluation with human preference, revealing that instruc…

15:14
2026-08-20
promptcube3.com
artificial-intelligence

Rethinking LLM scaling after Jie Tang's latest breakdown

Jie Tang's team at Z.ai argues that LLM scaling should optimize for fixed inference budgets, favoring smaller, deeper architectures trained longer over wider models. Their ablation shows a 7B model tr…

04:00
2026-08-20
machinebrief.com
artificial-intelligence

What is Missing from AI Post-Training AI: An Empirical Analysis

A new empirical analysis from arXiv (2608.19072v1) finds that large language model (LLM) agents post-training an LLM lock in their training strategy at the very beginning and spend the remaining budge…

page 1 / 3 next →
// co-occurs with top 8 entities
// topics top 6 topics