Shapelearn Qwen 3.8 27B (13.1 GB VRAM)
ByteShape released full ShapeLearn GGUF quantizations for Qwen 3.8 27B, with all five models sitting on the measured quality-speed frontier across six GPU comparisons, the company reported. The releas…
ByteShape released full ShapeLearn GGUF quantizations for Qwen 3.8 27B, with all five models sitting on the measured quality-speed frontier across six GPU comparisons, the company reported. The releas…
A developer built a minimal Python web server that turns Cerebras-hosted inference of Alibaba's Qwen 3.8 27B model into a live operating system with zero apps on disk, streaming tokens at 1,950 tokens…
Alibaba's Qwen team released Qwen 3.8 27B, a 27-billion-parameter model distributed as a 17GB GGUF file that scores 52 on the Artificial Analysis Intelligence Index, matching OpenAI's cloud-hosted GPT…
A developer community analysis argues that 2026's leading large language models are deliberately trained to minimize stored factual knowledge in favor of compact, reusable reasoning procedures, citing…
AkitaOnRails published Part 2 of its LLM Benchmark v4, retesting 39 large language models on a seven-sprint Rails app suite seeded with 14 real CVE-based sabotages, after the author rejected the entir…
A homelab user identified as gessha detailed a home setup that uses two Nvidia RTX 3090 GPUs in an i5-8600K workstation to host Qwen 3.8 27B and Qwen 3-VL 4B models via llama-server, alongside three m…
A user running hybrid GPU+CPU inference on 4x Nvidia V100 GPUs reported beating an Nvidia RTX 5090 in token generation (TG) for Qwen 3.8 27B and outperforming two DGX Spark systems with Qwen 3.8 Flash…
Perplexity released Portable Computer for Windows, a local AI agent that runs on NVIDIA GeForce RTX or RTX PRO GPUs with at least 24GB of VRAM and uses a Qwen 3.8 27B model post-trained for the Perple…
Perplexity released Portable Computer for Windows on compatible NVIDIA GeForce RTX PCs and RTX PRO Workstations, a local version of its Perplexity Computer agent that runs local models accelerated by …
SwarmLLM, an open-source project by developer Nehanth, enables a Qwen 3.8 27B model to run across browser tabs on multiple devices, with each device holding a slice of the model and passing 10 KB acti…
Cerebras added Alibaba's Qwen 3.8 27B to its public API, delivering 1,500 tokens per second, enabling a 300-word response in under half a second. The model scores 61.7% on SWE-bench Pro, ahead of Clau…
OpenAI's Astra launch has locked out paying Plus and Pro subscribers while enterprise customers received first access, prompting CEO Sam Altman to apologize on X for the 'messy' rollout. Separately, r…
Mistral announced Agentic Search, a feature enabling autonomous search workflows with a data opt-out guarantee for training. The company also highlighted inference optimization for agentic use cases, …
OpenAI shipped GPT-6 Astra on Thursday at $10 per million input tokens and $50 per million output, 2.5 times the price of GPT-5.6 Sol, and rated it 'Critical' for cyber capability under its Preparedne…
Cerebras Systems now offers Qwen 3.8 27B at 1500 tokens per second on its wafer-scale hardware, reducing end-to-end latency to approximately 18 milliseconds per token and enabling real-time multi-step…
Cerebras Systems announced that the Qwen 3.8 27B model is now available on its platform at 1500 tokens per second, with all public models served unpruned and using selective weight-only quantization f…
Qwen 3.8 27B, running locally on a 16GB RAM MacBook Pro, outperformed OpenAI's GPT 5.6 Luna Max on the DABstep benchmark at over 17 times lower cost, with electricity costs under $0.50 versus over $8.…
Qwen 3.8 27B, an open-weight model, outperformed OpenAI's GPT 5.6 Luna Max on the DABstep benchmark while running locally on a laptop, costing under $0.50 in electricity versus over $8 for Luna Max, a…
Perplexity launched Portable Computer, a local version of its agentic AI software that runs on-device to keep data off the cloud by default, available initially to Pro and Max subscribers on Nvidia's …
Perplexity launched Portable Computer on August 25, moving its Computer agent to run fully local on Nvidia's DGX Spark desktop, which costs $4,699, with zero per-token costs for local jobs. The initia…