cd/sources/machinebrief-auto-discovered· home› sources› Machinebrief (auto-discovered)
cat /sources/machinebrief-auto-discovered.feed | wc -l → 5116

Machinebrief (auto-discovered)

articles 5116 domain machinebrief.com → page 184/256 feed RSS
07:39
2026-07-15
machinebrief.com
artificial-intelligence

AI Benchmarks: When Is Enough Truly Enough?

A study examining partial evaluations on AI agent benchmarks including SWE-bench, AppWorld, and tau-bench finds that partial budgets are only valid when they replicate the full benchmark's final decis…

06:40
2026-07-15
machinebrief.com
artificial-intelligence

Evidence Selection in RAG with QUBO

A new approach using Quadratic Unconstrained Binary Optimization (QUBO) for evidence selection in retrieval-augmented generation (RAG) systems achieves competitive exact-match and token-F1 performance…

06:40
2026-07-15
machinebrief.com
artificial-intelligence

The Mysteries of Cross-Language Comprehension

Researchers have developed a Bayesian model that explains how speakers understand unfamiliar languages through intercomprehension, outperforming larger zero-shot language models in aligning with human…

06:40
2026-07-15
machinebrief.com
large-language-models

When AI Converges on 'Serendipity'

A study of 44 language models found that 41% of the time, AI models choose the word 'serendipity' when asked to pick any word, revealing a strong tendency toward conformity. In seven out of 31 categor…

06:39
2026-07-15
machinebrief.com
artificial-intelligence

Automating Financial Report Evaluation: A New Benchmark

Researchers have introduced a new pipeline that automates the evaluation of financial reports by generating rubrics without human experts, producing 14,450 candidate rubrics from 104 user queries. The…

06:39
2026-07-15
machinebrief.com
artificial-intelligence

Semantic Decoding: A New Fusion Framework at the Forefront

Researchers have developed a multi-feature fusion framework that combines static lexical representations (Word2Vec) with dynamic contextual representations (GPT) via a non-linear cross-attention mecha…

06:38
2026-07-15
machinebrief.com
large-language-models

Why Context Still Tricks Big Language Models

Large language models (LLMs) handle irrelevant context well in aggregate but remain fragile on specific examples, with random gibberish sometimes boosting performance and other times dragging it down,…

06:38
2026-07-15
machinebrief.com
artificial-intelligence

WikiSTAR: The AI That Knows Wikipedia's Secrets

Researchers have developed WikiSTAR (Scientific Tracking of Article Revisions), an AI tool that uses an LLM classifier to analyze how scientific articles evolve on Wikipedia by tagging changes such as…

06:38
2026-07-15
machinebrief.com
artificial-intelligence

South Korea's Bold Move: A National AI Chatbot for All

South Korea's government has issued a tender for private companies to develop a universal AI chatbot and a government services agent, backed by up to 256 Nvidia B200 GPUs for successful bidders who ma…

06:37
2026-07-15
machinebrief.com
artificial-intelligence

ASML's AI Chip Bet: A Wager on the Future

ASML, a heavyweight in semiconductor manufacturing, raised its forecast for the second time in 2023, betting on surging demand for AI chips as tech giants expand production capacities. The company's a…

06:25
2026-07-15
machinebrief.com
large-language-models

LLMs in Cancer Care: Promise Meets Skepticism

A study testing large language models on 72 real breast cancer cases found that the top performer, Claude Opus 4.8 with a D&C+SA pipeline, achieved a global score of only 0.594 ± 0.025, revealing pers…

06:25
2026-07-15
machinebrief.com
artificial-intelligence

Humans vs. AI: The Search for Semantic Soul

A new study comparing human semantic memory retrieval with that of large language models GPT-4o, Gemini 2.5-Pro, and Claude-Sonnet-4.5 found that humans exhibit higher entropy, larger semantic steps, …

06:25
2026-07-15
machinebrief.com
artificial-intelligence

AgentLens: The Code Agent Benchmark We've Been Waiting For

AgentLens, a new open-source benchmark for interactive code agents, evaluates not just task completion but also instruction-following, verification, and error recovery by blending formal verification …

05:40
2026-07-15
machinebrief.com
machine-learning

Neural Operators and Multiscale Challenges: A New Approach

Researchers have introduced LOD-MSNO (LOD-Multiscale Neural Operator), a hybrid model that combines the Localized Orthogonal Decomposition method with neural operators to improve accuracy on multiscal…

05:40
2026-07-15
machinebrief.com
machine-learning

Why Task-Specific Synthetic Data Beats Image Quality

A new method called Class-Contrastive Influence (C2I) quantifies synthetic sample usefulness by gradient-based influence on classifiers, outperforming traditional realism-focused approaches in few-sho…

05:40
2026-07-15
machinebrief.com
artificial-intelligence

Temporal Video Models: Why Order Matters

A new evaluation technique called reversal-drop reveals that temporal video models often rely on positional encoding rather than understanding visual sequences, with Molmo2 failing to distinguish even…

05:40
2026-07-15
machinebrief.com
artificial-intelligence

Fin-Analyst: Revolutionizing Equity Trading with AI

Fin-Analyst, an AI trading agent, achieved a 13.51% return on Tesla (TSLA), outperforming the Buy-and-Hold strategy by 28.33 percentage points, according to its developers. The agent uses an eight-spe…

← prev page 184 / 256 next →