Cache-Control for LLMs
Anthropic's Claude Sonnet 5 charges $2.00 per million fresh input tokens, $2.50 to write into a five-minute cache, and $0.20 per read, making two identical calls cost $2.70 with caching versus $4.00 w…
Anthropic's Claude Sonnet 5 charges $2.00 per million fresh input tokens, $2.50 to write into a five-minute cache, and $0.20 per read, making two identical calls cost $2.70 with caching versus $4.00 w…
Together AI introduces ThunderAgent, a system for high-throughput agentic inference that achieves up to 2.5× higher single-node throughput and 2.4× speedup on an 8-node cluster with near-linear scalin…
Mingxin FX100 achieves 90% line-rate utilization on a single 100GbE port, delivering approximately 11.25 GB/s effective bandwidth in measured tests. This eliminates the network as a primary bottleneck…
AWS introduced disaggregated prefill and decode (DPD) for LLM inference on SageMaker HyperPod, separating compute-bound prefill and memory-bound decode onto different GPU pools connected via EFA RDMA …
LMCache introduces a novel KV cache optimization layer to accelerate LLM inference, enabling faster local deployment on consumer hardware. AllenAI releases olmo-eval, a workbench for evaluating open l…
Tensormesh announced $20 million in seed extension funding led by AMD Ventures, with participation from CoreWeave, NVentures, Valley Capital Partners, and Laude Ventures, bringing total capital raised…