cd/entity/RTX 5090· home entities RTX 5090
grep -l @rtx 5090 /news/*.json | wc -l → 72

RTX 5090

mentions 72 type Person page 3/4 feed RSS

// recent coverage 72 mentions

20:26
2026-07-10
machinebrief.com
large-language-models

ARCQuant: Redefining Efficiency in LLM Inference with NVFP4

ARCQuant, a new framework for Large Language Model inference, uses the NVFP4 numerical format to achieve up to 3x speedup on GPUs while maintaining accuracy comparable to full-precision baselines. The…

17:04
2026-07-10
sourcefeed.dev
artificial-intelligence

Why Mini PCs Run 70B Models That Discrete GPUs Can't

Mini PCs with unified memory architectures, such as those using AMD's Strix Halo chip, can run 70-billion-parameter language models that exceed the VRAM capacity of discrete GPUs like the RTX 5090, bu…

08:59
2026-07-10
techpowerup.com
ai-chips

NVIDIA Readies GeForce RTX 5090 SE Graphics Card

NVIDIA is preparing the GeForce RTX 5090 SE graphics card, a cut-down version of the flagship RTX 5090 based on the same GB202 silicon but with fewer cores and a 500W TGP, targeting a $1,500 MSRP. Unl…

10:32
2026-07-04
dev.to
large-language-models

Solving the GPU Pinning Saga and Gemma's Meta-Commentary

Glad Labs fixed a GPU pinning issue where LiteLLM 1.89.2's global api_base override prevented per-model routing, causing vision tasks to cold-load onto the wrong GPU. The team also hardened content gu…

12:42
2026-07-01
news.ycombinator.com
ai-tools

Ask HN: Move to Private Models?

A Hacker News user asks the community for advice on moving from cloud-based AI services to private models for sensitive data analysis and reasoning, considering options like adding dual RTX 5090 GPUs …

23:39
2026-06-30
latent.space
large-language-models

Ahmad Osman on why local AI is catching up

Ahmad Osman, founder of Osmantic, argued at the AI Engineer World's Fair that local AI is rapidly catching up to proprietary frontier models, driven by shrinking gaps in open-source LLMs and improved …

20:14
2026-06-30
dev.to
large-language-models

GLM Is the New Hotness, So Let's Test It On the Homelab

GLM, the model family from Z.ai (formerly Zhipu AI), is gaining attention for its open weights and strong benchmark numbers. A developer tested GLM-4.7-Flash, a 30B-A3B MoE model, on a single-GPU home…

03:50
2026-06-29
gladlabs.io
artificial-intelligence

Automating AI Content Workflows

Glad Labs has developed an AI-operated content pipeline that treats content generation as an engineering problem, moving beyond simple prompting to autonomous agents with human oversight. The system u…

04:59
2026-06-28
gladlabs.io
artificial-intelligence

The Operational Cost of Manual Content

A new AI content pipeline called Poindexter uses agent infrastructure, RAG pipelines, and open-source LLMs to automate technical publishing, reducing the human loop and enabling solo developers to pro…

20:17
2026-06-16
github.com
artificial-intelligence

Show HN: cuTile Rust: Safe, data-race-free GPU kernels in Rust

NVIDIA Research released cuTile Rust, a tile-based system for writing memory-safe, data-race-free GPU kernels in Rust. The project extends Rust's ownership model to GPU programming, achieving up to 92…

16:04
2026-06-16
simonwillison.net
large-language-models

Quoting Georgi Gerganov

Georgi Gerganov, creator of llama.cpp, reported that the Qwen3.6-27B model is highly capable for local coding tasks, using it daily on his M2 Ultra and RTX 5090 systems. He noted the model's utility f…

21:53
2026-06-11
letsdatascience.com
generative-ai

Diffusion Decoders Replace VAE Decoders for 4K Images

NVIDIA Research and Tencent YoutuResearch have released open-source systems that replace the VAE decoder bottleneck in latent diffusion models, enabling faster 4K image generation. NVIDIA's PiD reform…

← prev page 3 / 4 next →
// co-occurs with top 8 entities
// topics top 6 topics