CLI Agent Failures: Why Early Detection Is Key
A study analyzing 1,794 valid coding trajectories from a dataset of 3,843, generated by seven leading models across three coding-agent scaffolds (OpenHands, MiniSWE, and Terminus2), found that coding …
A study analyzing 1,794 valid coding trajectories from a dataset of 3,843, generated by seven leading models across three coding-agent scaffolds (OpenHands, MiniSWE, and Terminus2), found that coding …
VGGT, a 3D reconstruction model, uses its internal layer L17 to differentiate co-visible image pairs without explicit supervision, solving a key challenge in 3D vision. Building on this, Co-VGGT adds …
A new framework called VEXAIoT uses LLM agents to exploit IoT vulnerabilities with a 95.0% overall success rate across ten attack scenarios, achieving 94.5% in IoTGoat and 96.7% in Metasploitable2. De…
Researchers have developed a parameter-efficient AI framework based on CLIP models that enhances animal re-identification by adapting to visual shifts over time, using continuous metadata conditioning…
Researchers have introduced STEEL, the first open-source implementation of FlashAttention optimized for XDNA-like neural processing units (NPUs), achieving significant energy and speed advantages for …
Google Research's SensorFM, a foundation model trained on over a trillion minutes of Fitbit and Pixel Watch data from five million users, beats existing benchmarks on 34 of 35 health and behavioral ta…
Beijing-based Zhipu released GLM-5.2, an open-source AI model with a 1 million-token context window that competes with premium models like OpenAI's GPT-5.5 at no cost, though it suffers from speed iss…
A new study from Northeastern University professor Christoph Riedl finds that managers often devalue employees' work once they learn AI played a part, creating a 'AI penalty' where workers face conseq…
Apple is suing OpenAI and former employees for alleged intellectual property theft, accusing them of unauthorized access to confidential data and unreleased product information. The lawsuit highlights…
Anthropic is extending Claude Fable 5 in subscription plans until July 19, 2026, allowing subscribers to use up to 50 percent of their weekly limit on the model. The move is a tactical response to pri…
A Broadcom survey finds 56% of enterprises are transitioning to or planning private cloud environments for AI, driven by concerns over data sovereignty, security, and governance. Oliver Rowell of Xtra…
Price cuts by Meta, OpenAI, and SpaceXAI on AI models within eight days have not lowered enterprise agent bills, as hidden costs such as integration, training, and data preparation keep overall expens…
The European Union's AI Act will require AI systems to disclose critical information about their operations and decision-making processes starting August 2, 2026, a mandate aimed at transparency but c…
China is rapidly deploying AI across sectors, from medical avatars to food-delivery drones on the Great Wall, while expanding state surveillance with facial recognition and data tracking. The governme…
Meta pulled its new AI image tool, Muse Image, just 72 hours after its July 8 launch due to backlash over privacy concerns, with the feature automatically altering public images without explicit conse…
A large-scale study of public repositories reveals that software engineering activities are increasingly being packaged into reusable AI agent skills, covering diverse tasks across the development lif…
GenCeption, a new computer vision model from an unnamed research team, uses text-to-video generation to achieve state-of-the-art performance on tasks like depth estimation and 3D keypoint prediction, …
Researchers have demonstrated a 128-qubit Quantum Convolutional Neural Network (QCNN) that achieves efficient image classification on the MNIST dataset without exponentially growing hardware demands. …
ConceptSMILE, a model-agnostic auditing framework, assesses the reliability of concept-based AI explanations by extending SMILE's perturbation-based logic. In a case study on retinal fundus images, Me…
Urban mining's future depends on AI that supports human auditors through defensible, not just accurate, pre-demolition assessments, according to a new analysis. The integration of explainable AI and d…