cd/entity/NVIDIA A100· home entities NVIDIA A100
grep -l @nvidia a100 /news/*.json | wc -l → 10

NVIDIA A100

mentions 10 type Person feed RSS

// recent coverage 10 mentions

13:42
2026-08-31
promptcube3.com
artificial-intelligence

Stop treating your infrastructure as a separate layer from your

A new technical guide argues that AI-native applications require rethinking infrastructure as a core part of the workflow, not a separate layer, citing the probabilistic nature of LLMs and the need fo…

12:16
2026-08-14
dev.to
large-language-models

vLLM vs Ollama: Production Serving 2026

A developer's comparison of vLLM and Ollama for LLM serving in 2026 shows that while Ollama is simpler for single-user local use, vLLM outperforms it dramatically under concurrency, with throughput up…

13:30
2026-08-10
cast.ai
artificial-intelligence

LLM Inference Cost Optimization: Run AI Inference for Less

Cast AI benchmark testing shows that continuous batching at batch size 8 reduces Llama 3.1 70B inference cost on a single H100 from approximately $0.60-$0.80 per million tokens to $0.15-$0.25 per mill…

07:53
2026-07-27
snipvote.com
artificial-intelligence

AgentKVShift cuts agentic memory prefill latency 2-3.5x on A100

AgentKVShift, a training-free method for agentic memory retrieval, achieves 2–3.5x prefill speedups on a single NVIDIA A100 GPU by refreshing only 10–30% of the KV cache while maintaining near full-re…

01:46
2026-06-16
github.com
large-language-models

Show HN: Locket – Robust feature-level access control for LLMs

Researchers from Aalto University and the University of Waterloo introduced Locket, a feature-locking technique for large language models that enables pay-to-unlock schemes by restricting specific mod…

00:00
2026-06-11
huggingface.co
machine-learning

Profiling in PyTorch (Part 2): From nn.Linear to a Fused MLP

PyTorch's `nn.Linear` module transposes its weight tensor before performing matrix multiplication and addition, as revealed by profiler traces showing an `aten::t` operation that only modifies tensor …

// co-occurs with top 8 entities
// topics top 6 topics