cd/entity/RTX 3090· home entities RTX 3090
grep -l @rtx 3090 /news/*.json | wc -l → 57

RTX 3090

mentions 57 type Person page 3/3 feed RSS

// recent coverage 57 mentions

14:16
2026-06-16
byteiota.com
large-language-models

Local LLMs vs Claude for Coding: The 70% Problem

A Hacker News thread on June 16 revealed that local LLMs like Qwen 3.6 35B-A3B handle about 70% of daily coding tasks but fall short on complex multi-file reasoning, creating a gap akin to a junior ve…

02:20
2026-06-16
discuss.huggingface.co
ai-agents

Assimetric parallel inference using consumer RTX PC

A user with a 24GB RTX 3090 and i5-10400 PC is experimenting with asymmetric parallel inference to reduce model looping and agent freezing, using their gaming PC as a platform to learn basic agentic A…

18:37
2026-06-15
discuss.huggingface.co
large-language-models

Unusual parallel inference using consumer RTX rig

A technical report proposes using a consumer RTX 3090's integrated GPU (iGPU) to run a small language model as a 'Sentinel' for monitoring and validating outputs from the primary GPU-bound model. The …

00:00
2026-06-15
glukhov.org
large-language-models

Cost Optimization for LLM Systems: Where the Money Actually Goes

LLM costs scale linearly with usage, and enterprises spending over $10,000 annually can optimize by implementing token budgets, choosing between API and local inference, and using fallback strategies.…

11:39
2026-06-14
runtimewire.com
artificial-intelligence

Pearl's AI mining pitch faces a 112 MW usefulness test

Pearl Research Labs claims to have built a live GPU-secured blockchain without proving the work is useful AI computation, facing a 112 MW usefulness test. A preprint estimates Pearl's network runs at …

09:55
2026-06-13
imil.net
large-language-models

RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8

A user combined an RTX 5080 and RTX 3090 on an Asus Prime X570-Pro motherboard to run Qwen 3.6 27B Q8 at over 80 tokens per second. The setup required disabling CSM, enabling Above 4G Decoding and ReS…

02:56
2026-06-09
vettedconsumer.com
ai-products

RTX 5090: A 32GB AI Powerhouse — or an Expensive Way to Game?

NVIDIA's RTX 5090, priced at $1,999, offers 32 GB of GDDR7 memory with 1,792 GB/s bandwidth, making it the fastest consumer GPU for both gaming and local AI inference. The card excels at running 32B-c…

19:30
2026-06-05
gilesthomas.com
machine-learning

JAX backends and devices

JAX defaults to loading data directly onto GPU memory when a CUDA-enabled version is installed, causing out-of-memory errors for large datasets that would fit in system RAM. The framework's `jax.devic…

07:08
2026-06-04
dev.to
ai-infrastructure

His AI Said 'Swap the PSU.' He Said 'One More Test.'

A homelab engineer known as Marco spent weeks debugging an RTX 3090 that hard-reset his entire machine every time it ran inference, leaving no logs behind. The GPU crash, triggered by a routine NVIDIA…

01:13
2026-05-30
dev.to
large-language-models

Used RTX 3090 Buying Guide for Local LLM in 2026

A used RTX 3090, now three generations old and costing under $900, remains the best value for running 30B+ parameter LLMs locally in 2026, as its 24GB of VRAM is the only single-GPU solution that fits…

03:39
2026-05-23
dev.to
large-language-models

BeeLlama v0.2.0: 164 tok/s on a 27B model, one RTX 3090

BeeLlama v0.2.0 demonstrates that speculative decoding can achieve a 4.4x to 4.93x throughput multiplier on a single RTX 3090, running 27B and 31B parameter models at 37-36 tokens per second baseline …

← prev page 3 / 3
// co-occurs with top 8 entities
// topics top 6 topics