Running qwen 3.6 / 3.8 on 3090+3080 over RPC?
A user running Qwen 3.8 27B on a 3090 reports that switching to ninfer, a Qwen-focused inference engine, doubled their token generation speed from 35-40 to 55-60 tokens per second. The user is conside…
A user running Qwen 3.8 27B on a 3090 reports that switching to ninfer, a Qwen-focused inference engine, doubled their token generation speed from 35-40 to 55-60 tokens per second. The user is conside…
A user running ninfer's Qwen 3.8 27B model on an RTX 3090 is seeking guidance on creating an abliterated version of the model, noting that ninfer uses a custom file format that prevents simple configu…
Perplexity released Portable Computer on August 25, a local version of its Computer agent platform that runs AI agents on users' own Nvidia hardware, starting with the DGX Spark and RTX GPU PCs, to re…
Perplexity AI Inc. launched Portable Computer, an on-device AI agent for desktops with Nvidia Corp. silicon, following reports that Nvidia is considering an investment valuing the startup at over $30 …
Perplexity launched Portable Computer, an on-device AI offering that runs the Nvidia DGX Spark with Qwen 3.8 27B or PPLX 27B, keeping data local and charging only when tasks escalate to the cloud. The…
Perplexity launched Portable Computer on Tuesday, moving its Computer agent stack onto NVIDIA's DGX Spark so files, models and most tool execution can stay on user-controlled hardware, with work compl…
A developer reports that running Qwen 3.8 27B at Q5_K_M quantization on a dual AMD Radeon RX 7900 XT and 7800 XT setup achieves 20 tokens per second with a 256k context window, enabling autonomous mul…
Perplexity has released Portable Computer, a local-first build of its agentic Computer platform that runs on NVIDIA DGX Spark and RTX GPUs with 24 GB of VRAM, packaging the local model, inference engi…
Perplexity launched Portable Computer, a fully local version of its agentic Computer platform that runs on Nvidia DGX Spark desktops and Linux machines with Nvidia RTX GPUs, eliminating monthly token …
Perplexity launched Portable Computer, a local AI agent that runs open models on desktop hardware, with Nvidia's DGX Spark as a partner platform. The agent runs a post-trained version of Alibaba's Qwe…
A hardware enthusiast is weighing two multi-GPU inference builds, comparing Nvidia RTX 4000 Pro Blackwell GPUs against AMD R9700 Pro GPUs, with the Nvidia option consuming 140 watts per GPU versus 300…
DeepSeek V4 Flash, a mixture-of-experts model, ran at roughly 10 to 11 tokens per second on a single RTX 3090 with 192GB of system RAM in tests by FreeToken's desktop app, while a dense Qwen 3.8 27B m…
Anthropic's flagship Fable model is struggling to attract paying users as cheaper competitors dominate the mass market, according to the Financial Times, with aggressive rate limits, no ZDR option for…
Alibaba's Qwen 3.8 27B open-weight vision model, released under Apache 2.0, takes 21 minutes and 22,276 reasoning tokens to draw a pelican riding a bicycle as an SVG, according to Simon Willison's tes…
A user reports running Qwen 3 Coder 30B A3B, Qwen 3.6 27B, and Qwen 3.8 27B on a local machine with a 7800X3D, 64GB DDR5, and an RTX 3090 24GB, achieving about 70 tokens per second on Qwen 3.8 27B, an…
A user reports running large language models on a 128GB Bosgame M5 Strix Halo mini PC purchased used for 1800€ on eBay, achieving 30 tokens per second decode with Qwen 3.8 27B and 50 tokens per second…
Qwen 3.8 27B, a local LLM from Alibaba's Qwen team, achieves milestone-level coding performance on consumer hardware, with a user reporting it as the first local model that convinced them to integrate…
VPIPE, a new macOS app from developer tgo-app-dev, runs MiniMax H3, a 33B multimodal model generating video and audio, on Apple Silicon Macs with as little as 16 GB RAM, achieving a 5-second 0.5 MP 24…
A quality study from Kingy AI, a local-AI benchmarking blog, found that Qwen 3.8 27B, a 27-billion-parameter vision model, achieves roughly 92% top-token agreement with the full-precision reference at…
Alibaba Cloud's Qwen 3.8 27B open-weights model delivers roughly double the intelligence and capability of last year's OpenAI gpt-oss models on the same hardware, according to software engineer Willia…