cd/entity/RTX 3090· home› entities› RTX 3090
grep -l @rtx 3090 /news/*.json | wc -l → 82

RTX 3090

mentions 82 type Person page 2/5 feed RSS

// recent coverage 82 mentions

13:32
2026-08-26
forum.level1techs.com
large-language-models

Make a custom qwen 3.8 27B abliterated for ninfer?

A user running ninfer's Qwen 3.8 27B model on an RTX 3090 is seeking guidance on creating an abliterated version of the model, noting that ninfer uses a custom file format that prevents simple configu…

00:00
2026-08-25
mindstudio.ai
artificial-intelligence

DeepSeek V4 Flash on One RTX 3090: Real Tokens-Per-Second Numbers

DeepSeek V4 Flash, a mixture-of-experts model, ran at roughly 10 to 11 tokens per second on a single RTX 3090 with 192GB of system RAM in tests by FreeToken's desktop app, while a dense Qwen 3.8 27B m…

00:00
2026-08-25
mindstudio.ai
artificial-intelligence

Escha-W2: 2-Bit Quantization That Shrinks a 27B Model to 10GB

Escha Labs Inc. released Escha-W2, a 2-bit quantized build of Qwen3.8-27B that compresses the 27-billion-parameter model to 10.15GB, enabling 128k context on a single 24GB GPU while matching FP8 quali…

00:00
2026-08-25
mindstudio.ai
artificial-intelligence

Run Qwen3.8-27B-Escha-W2 on a 24GB GPU with SGLang

Escha Labs released Qwen3.8-27B-Escha-W2, a 2-bit quantized build of Qwen3.8-27B that fits 10.15 GB of weights on a 24 GB consumer GPU, enabling 64k context out of the box and up to 128k with tuning v…

03:27
2026-08-23
forum.level1techs.com
large-language-models

Running qwen 3.6 / 2.8 on 3090+3080 over RPC?

A user reports running Qwen 3 Coder 30B A3B, Qwen 3.6 27B, and Qwen 3.8 27B on a local machine with a 7800X3D, 64GB DDR5, and an RTX 3090 24GB, achieving about 70 tokens per second on Qwen 3.8 27B, an…

00:00
2026-08-20
runagentrun.co.uk
large-language-models

Dual 3090s: the bottleneck isn't the GPU

Two benchmarks of Qwen3.8-27B on a single RTX 3090 show a 3.2x performance gap: 41.49 tok/s with llama.cpp (build b10088) versus 132 tok/s with vLLM using a DFlash2 block drafter, according to Insider…

22:34
2026-08-18
promptcube3.com
large-language-models

Llama 3.

Meta's Llama 3.1 70B model can now run on a single 24GB consumer GPU using GGUF or EXL2 quantization, achieving 5-10 tokens per second on an RTX 3090, according to a deployment guide. The guide recomm…

14:00
2026-08-18
kdnuggets.com
artificial-intelligence

Run Qwen3.8-27B as a Local AI Coding Agent in Just 3 Commands

Ollama and OpenCode now enable running Qwen3.8-27B as a local AI coding agent with just three terminal commands, according to a guide from OpenCode. The process involves installing Ollama, pulling the…

18:48
2026-08-17
promptcube3.com
artificial-intelligence

Can I run a GPT-5 Codex review locally using Ollama?

A developer tested running coding models locally via Ollama and found that DeepSeek-Coder-V2 on an RTX 3090 delivers 0.2s time to first token and 45 tokens per second, compared to GPT-4o's 1.1s and 60…

00:00
2026-08-11
mindstudio.ai
artificial-intelligence

Meta Muse Glimmer 30B: How to Run It Locally and Is It Worth It?

Meta released Muse Glimmer, a 30 billion parameter open-weight language model under Apache 2.0, designed for agentic tasks and positioned as a competitor to Qwen 3.6 27B. The model is available as an …

11:48
2026-08-05
gist.github.com
machine-learning

Cloud Training on RunPod: A Field Guide to the Edge Cases

AlphaPebble Labs engineers detailed a field guide for training AI models on RunPod's rented GPU infrastructure, highlighting edge cases such as the SSH gateway acting as a console rather than an exec …

14:20
2026-07-26
promptcube3.com
artificial-intelligence

Llama local deployment for secure code reviews

A developer reports that running Llama 3.1 8B locally on an RTX 3090 enables secure code reviews with full control over context and system prompts, achieving 1.2-second response times on 50-line snipp…

19:46
2026-07-25
promptcube3.com
large-language-models

DeepSeek-R1 Local Deployment: My Hardware Struggles

A user reports that deploying the full 671B parameter DeepSeek-R1 model locally requires over 100GB of VRAM and is impractical on consumer hardware, with CUDA out-of-memory errors occurring even at sm…

10:01
2026-07-25
promptcube3.com
artificial-intelligence

Voice Cloning: Quality vs. Length

Sample quality, not length, is the primary driver of voice cloning realism, with reverb being the most damaging artifact because it becomes permanently embedded in the speaker embedding, according to …

← prev page 2 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics