GML5 IndexCache
Researchers from Tsinghua University and Z.ai have proposed IndexCache, a method to reduce the computational cost of DeepSeek Sparse Attention (DSA) in GLM-5.2. IndexCache exploits the observation that adjacent layers in…
Neural Networks news and analysis on Web Pulse: 1261 curated articles tracking the latest Neural Networks developments, tools, and research, updated continuously from vetted sources.
Researchers from Tsinghua University and Z.ai have proposed IndexCache, a method to reduce the computational cost of DeepSeek Sparse Attention (DSA) in GLM-5.2. IndexCache exploits the observation that adjacent layers in…
Meta introduces Brain2Qwerty v2, a non-invasive brain-computer interface that decodes sentences from magnetoencephalography (MEG) recordings with up to 78% word accuracy, aiming to restore communication for people who ca…
Teens' inattention is often due to normal brain development, not defiance, as their prefrontal cortex is still building neural networks for focus. Parents can help by identifying 'attention robbers' like novelty and soci…
Meta's FAIR lab, in collaboration with Spain's BCBL, unveiled Brain2Qwerty, a non-invasive AI system that decodes typed sentences from brain activity with up to 80% character accuracy using magnetoencephalography (MEG). …
Nansense, a new interactive PyTorch debugger, allows developers to pause training, step batch-by-batch, and time-travel to different epochs while visualizing activations, gradients, weights, and optimizer state. The tool…
Anomaly puzzles challenge the brain's pattern-recognition system by presenting violations of expectation, illustrating how prediction errors drive anomaly detection. The puzzles, inspired by Sherlock Holmes's reasoning, …
Researchers accidentally discovered that ordinary CMOS transistors can function as artificial neurons and synapses, potentially enabling neuromorphic computing that is vastly more energy-efficient than current AI hardwar…
Meta AI's Brain2Qwerty system decodes brain signals into text using non-invasive MEG sensors, achieving a 32% character error rate, but the technology is not yet real-time or commercially viable. The research demonstrate…
Researchers at arXiv have discovered that activation patching, a key tool in mechanistic interpretability, contains hidden interaction effects (INT) that can distort causal attributions. These effects, which measure how …
Researchers propose Transformation-Aware Decoupling (TAD), a framework for 3D scene graph generation that improves robustness to viewpoint changes by decoupling relation reasoning into stable and directional components. …
Researchers propose vMFProto, a distributional part-prototype framework that models each class as a mixture of von Mises-Fisher components on the hypersphere for interpretable classification. The method achieves state-of…
Researchers introduce the Prism Transformer, a new architecture that progressively increases head counts across layers to create a hierarchical attention processing structure. This design improves performance on zero-sho…
Researchers developed a geometry-conditioned Fourier neural operator (FNO) to approximate the solution operator for the cubic nonlinear Schrödinger equation on two-dimensional tori with varying aspect ratios. The model c…
Researchers propose HIA-GAT, a heterogeneous graph attention network for frame-level traffic conflict risk prediction on freeways. The model outperforms baselines on NGSIM datasets, achieving AUC up to 0.867, and provide…
Researchers developed a distribution-based deep multiple instance learning method for tumor proportion scoring in non-small cell lung cancer, using only slide-level labels to predict zero-inflated beta parameters. The ap…
Researchers propose HybridCodec, a novel approach combining discrete tokens with continuous residuals to improve speech representation in language models, reducing information loss and autoregressive steps while retainin…
Researchers introduced the context-ready transformer, a recurrent neural network architecture that pre-contextualizes tokens before they enter a transformer block, achieving faster inference and competitive performance. …
Researchers used a developmental approach to study how neural language models learn statistical patterns, finding that Transformers first acquire abstract global statistics before local dependencies, with early over-gene…
Researchers at Baylor College of Medicine found that patients under general anesthesia can process language at a sophisticated level, distinguishing parts of speech and predicting upcoming words, challenging traditional …
An independent research project analyzed the internal dynamics of small and medium-sized language models, revealing that functional properties become linearly decodable in hidden representations and that models cluster i…