cd/entity/RTX 3090· home› entities› RTX 3090
grep -l @rtx 3090 /news/*.json | wc -l → 82

RTX 3090

mentions 82 type Person page 1/5 feed RSS

// recent coverage 82 mentions

00:00
2026-10-02
nijho.lt
artificial-intelligence

Self-hosting AI does not save money, and I do it anyway

Self-hosting the open-weight Qwen3.8 27B model costs $619 per full Artificial Analysis Intelligence Index run through the cheapest zero-data-retention provider, versus $67 for OpenAI's GPT-6 Luna, acc…

11:44
2026-09-29
x.com
large-language-models

Qwen3.8-Flash-Next Is on TensorFold with Speed Boosts

TensorFold 0.3.6.2 delivered decode speeds over 62 tokens per second on a single stream and 119 tokens per second across five concurrent streams running Qwen3.8-Flash-Next on a single Nvidia DGX Spark…

05:49
2026-09-21
forum.level1techs.com
ai-infrastructure

Creating a custom homelab machine for LLM usage on a budget

A homelab builder running Mistral 7B and DeepSeek-R1 8B locally, and Mistral Large in the cloud, is comparing GPU options for a budget local LLM machine, primarily for software debugging. The listed p…

12:58
2026-09-16
twitter.com
large-language-models

Operating System powered by Qwen 3.8 27B at 1950 tokens/SEC

A developer built a minimal Python web server that turns Cerebras-hosted inference of Alibaba's Qwen 3.8 27B model into a live operating system with zero apps on disk, streaming tokens at 1,950 tokens…

00:00
2026-09-15
doug.sh
ai-infrastructure

Keeping vLLM's Prefix Cache Warm Between Agent Turns

Tuning vLLM 0.28.0's prefix cache raised the share of prompt tokens served from cache from 55% to 95% on a local Qwen3.8-27B coding agent, cutting average time to first token from 26-28 seconds to 7.3…

23:13
2026-09-12
agentsearchengine.app
ai-research

local-deep-research shipped an update

The open-source project local-deep-research reported approximately 95% on SimpleQA (n=500) and 77% on xbench-DeepSearch (n=100), which it says makes it the first fully-local open-source project to hit…

22:05
2026-09-12
promptcube3.com
large-language-models

Gemini Forum, Qwen Coder local setup

A developer resolved CUDA out-of-memory crashes running Qwen2.5-Coder-32B on a 24GB RTX 3090 by capping the context window at 8k tokens instead of 32k via a custom Ollama Modelfile, cutting latency fr…

22:22
2026-09-10
gist.github.com
generative-ai

Aurora worfklow

A developer released "Aurora," a three-shot alpine video generated in a single unbroken H3 ref2va generation on an RTX 3090, using Z-Image to draw three storyboard frames that anchor cuts at 00:02.700…

13:00
2026-09-02
vettedconsumer.com
large-language-models

How Much RAM Do You Need to Run a Local LLM in 2026?

Running large Mixture-of-Experts local LLMs in 2026 requires at least 64GB of system RAM, ideally 128GB or more, because sparse MoE models keep rarely-used expert weights in system memory rather than …

03:13
2026-08-29
forum.level1techs.com
artificial-intelligence

MoE with little models

A forum user comparing CUDA and ROCm for local AI inference reports that ROCm issues have diminished and is considering an all-AMD build with dual R9700 GPUs by 2027, citing AMD's lower cost. Another …

page 1 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics