ORA: Smaller Models. Same Intelligence
Ora Computing launched an automated LLM compression engine that reduces model size by up to 70% with minimal accuracy loss, enabling deployment on edge devices, on-prem servers, or cloud infrastructur…
Ora Computing launched an automated LLM compression engine that reduces model size by up to 70% with minimal accuracy loss, enabling deployment on edge devices, on-prem servers, or cloud infrastructur…
A developer replaced 2.5 hours of daily busywork with a $0 AI agent setup running on a Mac Mini M4. The system uses local LLMs (Ollama with Qwen models), Python scripts, and cron jobs to automate emai…
Anthropic accused Alibaba of conducting the largest known distillation attack on its AI models, with 28.8 million exchanges through nearly 25,000 fraudulent accounts. Anthropic's policy head called fo…
A researcher replicated Anthropic's concept-injection experiments on 14 open-weight language models and found that the models do not satisfy criteria for genuine introspection, instead exhibiting stat…
Off Grid AI Desktop is a free, open-source app that runs large language models locally on Apple Silicon Macs, leveraging unified memory for GPU inference. The app supports models like Gemma and Qwen, …
Anthropic accused Alibaba of executing the largest known distillation attack against its Claude AI model, involving 28.8 million queries from 25,000 fraudulent accounts between April and June 2025. Th…
Researchers at arXiv found that jailbreak attacks on large language models can be detected by analyzing entropy dynamics in intermediate layers, rather than final outputs. The study shows that monoton…
Pangram Labs researchers explored the internal representations of their AI detection model Pangram 3.3.2 using document-level analysis of activations across layers, aiming to understand what the model…
Zhipu AI released GLM 5.2, a 744B parameter open-weight mixture-of-experts model that rivals Claude Opus in coding and visual design quality at lower inference cost. The model's MoE architecture activ…
Anthropic accused operators linked to Alibaba's Qwen AI lab of using nearly 25,000 fraudulent accounts to extract capabilities from its Claude AI model, generating over 28.8 million exchanges between …
Anthropic accused Alibaba's Qwen AI lab of conducting the largest distillation attack on its Claude model, using nearly 25,000 fraudulent accounts to make 29 million exchanges between April 22 and Jun…
Cli-Modelarium 0.1.4 adds support for Alibaba's Qwen models (via DashScope) and Z.AI's GLM models, bringing the total to 10 cloud LLM providers. The command-line tool enables side-by-side comparison o…
Anthropic accused Alibaba's Qwen lab of using nearly 25,000 fake accounts to conduct 29 million exchanges with its Claude AI model, the largest distillation campaign yet. The accusation, shared with U…
Alibaba's Qwen team released Qwen-AgentWorld, a language world model that simulates environment responses to agent actions, outperforming GPT-5.4 and Claude Opus 4.8 on the AgentWorldBench across seve…
Zhipu's GLM 5.2 open-weight model, available on Synthorai at roughly one-sixth of frontier per-token prices, achieves frontier-level benchmarks but its per-task cost varies by over an order of magnitu…
Alibaba Group slashed prices for its Qwen AI models by up to 80% on its Qoder coding platform, targeting US workday hours to attract global developers amid fierce competition with US and Chinese rival…
Alibaba's Qwen team released Qwen-AgentWorld, a language model that simulates complex environments natively, replacing external simulators for training AI agents. The model, trained on over 10 million…
The most upvoted papers on Hugging Face reveal a trend of AI shifting from answer models to action models, focusing on agents, simulation environments, GUI/mobile interaction, and benchmarks for real-…
A developer benchmarked 15 AI models through Global API's unified endpoint, measuring time to first token and tokens per second. Step-3.5-Flash from StepFun topped the speed leaderboard with 80 tokens…
A developer has documented methods for overseas developers to access Chinese LLMs such as Qwen, DeepSeek, and GLM without requiring a Chinese phone number. The most practical solution is using a gatew…