Deep Learning Is Not So Mysterious or Different
A paper by Andrew Wilson and colleagues, submitted to arXiv on March 3, 2025, argues that deep learning's generalization behaviors, such as benign overfitting and double descent, are not unique to neu…
A paper by Andrew Wilson and colleagues, submitted to arXiv on March 3, 2025, argues that deep learning's generalization behaviors, such as benign overfitting and double descent, are not unique to neu…
A new study introduces PLCBench, the first real-PLC hardware-in-the-loop framework for evaluating whether autonomous large language model (LLM) agents can convert network access to programmable logic …
A paper submitted to arXiv on April 2, 2023, titled 'Eight Things to Know about Large Language Models,' surveys evidence for eight surprising points about LLMs, including that they predictably get mor…
A new study from arXiv (submitted 26 Aug 2026) introduces SwarmWorld, a simulation where initially homogeneous language-model agents self-organize into evolving technological societies without assigne…
A new arXiv preprint (2608.25880) reveals that Office Open XML (OOXML) documents can yield different evidentiary content when displayed in Microsoft Office versus when extracted for large language mod…
A new arXiv paper (2608.27141) demonstrates that safety monitors for autonomous large language model (LLM) agents fail to detect attacks whose evidence is spread across multiple iterations, because tr…
A new paper submitted to arXiv on August 27, 2026, argues that the Gaussian kernel, also known as the squared exponential or radial basis function kernel, should never be used as a default in Gaussian…
A developer compared the research platforms OpenAlex and Valyu, detailing their respective strengths in scholarly metadata and full-text retrieval. OpenAlex offers a graph of 322 million works with ci…
Researchers introduced WikiSkill, a framework that co-evolves AI agent skills with a persistent knowledge base (wiki) by separating raw execution experience, accumulated knowledge, and executable skil…
A developer compared the Semantic Scholar and Valyu research APIs for building research agents, noting that Semantic Scholar provides metadata and citation graphs across 214 million papers, while Valy…
A developer built a tool modeling how AI providers throttle their models under high server load, finding that the common practice of throttling once user counts exceed a threshold can backfire by prom…
An independent researcher is seeking an arXiv endorsement for a paper on learned query optimization, which studies how graph-based query plan representations provide complementary optimization signals…
Hugging Face user @furkanyllmz's authorship claim on the paper 'MoganColBERT-TR: A Late-Interaction Multi-Vector Retrieval Model for Turkish' (arXiv 2608.26344) has failed approval for the second time…
A pre-registered audit of a frozen pedagogy judge, sealed before its 990 calls, found that difference-in-differences on a censored rating scale can manufacture an effect, with the registered primary e…
Netflix researchers introduced GenRec, an LLM-backed recommendation ranker built on an in-house foundational LLM, in a paper submitted to arXiv on Aug 10, 2026. In a large-scale A/B test, GenRec achie…
A new framework called GROUND (Governed Retrieval Over Unified Normalized Definitions) reduces hallucinations in LLM-based enterprise analytics by constraining generated SQL to a governed semantic lay…
A new hybrid scoring system for automated second-language speaking assessment, combining interpretable speech-timing features with a text-LLM fluency judgment, achieved a Spearman rho of 0.818 against…
A new arXiv paper (2608.26481v1) finds that using a single shared critic across parallel environments in reinforcement learning causes value mismatch that degrades learning, and shows that conditionin…
A new arXiv paper (2608.26363v1) proposes a unified mathematical framework linking information propagation in convolutional neural networks (CNNs) to relativistic physics, showing that symmetric filte…
Researchers propose CG4AI, a column generation framework that builds a convex combination of AI models while enforcing linear constraints on outputs, using a master linear program and pricing subprobl…