Inference Engineering, Co-op — Inferact
Inferact, founded by the creators and core maintainers of vLLM, is hiring University of Waterloo co-op students for an on-site Inference Engineering co-op in San Francisco, with pay not published. The…
Inferact, founded by the creators and core maintainers of vLLM, is hiring University of Waterloo co-op students for an on-site Inference Engineering co-op in San Francisco, with pay not published. The…
A developer documented step 03 of a manual "Kubernetes the Hard Way" homelab build on Proxmox, working through SSH root-login restrictions on Ubuntu cloud images and distributing a dedicated keypair t…
Futurum research cited in The Great Unification found 20% of enterprises were actively running the Model Context Protocol (MCP) in production in the first half of 2026, with another 26.9% building pil…
Cast AI's workload-level rightsizing engine, PrecisionPack, reduces provisioned CPU footprint by approximately 50%, according to Cast AI's 2026 State of Kubernetes Optimization Report. PrecisionPack, …
A developer describes how adopting Claude across their company's development workflow shifted their focus from writing code to evaluating it, after noticing the AI would hardcode values, ignore establ…
Cast AI creates low technical lock-in but moderate operational lock-in, according to a Cast AI analysis of its own Kubernetes node-provisioning product. The company states that Cast AI provisions stan…
Envoy AI Gateway has been renamed Agent Router and has joined the Agentic AI Foundation, giving agent builders a single integration point for every model and MCP tool. The project grew out of a 2024 c…
Thoughtbot published five actionable software development tips for healthcare teams, drawn from a panel featuring David Pace, Executive Director at Merck Research Labs IT Enablement and User Experienc…
Together AI launched public preview of preemptible compute for Together GPU Clusters, offering interruptible GPU nodes at a flat 50% of the on-demand rate on Kubernetes clusters in all regions. Preemp…
Kubernetes does not accelerate LLM inference but orchestrates GPU clusters by placing, monitoring, replacing, and scaling inference servers, according to an article on AI workload management. The piec…
Kubernetes 1.37 “Garhwal,” released August 26, 2026, includes gang scheduling in Beta, which prevents partial-job deadlocks that can leave idle GPUs costing around $467 per event. The feature, detaile…
Kubernetes v1.35 (Timbernetes) introduces 60 enhancements focused on production readiness and AI workloads, including GA for in-place Pod resizing, stable Pod generation fields, built-in mTLS, and a n…
A new beginner's guide from Tigera walks engineers through deploying an AI agent on Kubernetes, covering containerization, secret management, health probes, resource limits, and network egress control…
A technical analysis of the AI inference stack in 2026 maps the path of a single request from application to GPU, detailing layers such as API gateway, inference gateway, distributed serving, inferenc…
A developer outlines a foundational learning path for DevOps engineers that prioritizes Linux, networking, Git, Docker, and cloud concepts before tackling Kubernetes. The author argues that mastering …
Anthropic's Model Context Protocol (MCP) is an open standard that enables AI agents to interact with external systems such as Kubernetes clusters, observability stacks, and ticketing systems, addressi…
A developer proposes a unified mental model for distributed compute systems, arguing that frameworks like Kubernetes, Slurm, Ray, and Spark share fundamental challenges in scheduling, resource managem…
Tracarbon, a Python library that tracks device energy consumption and calculates carbon emissions, now supports monitoring GPU power and carbon output while running local large language models. It det…
Kubernetes 1.37 “Garhwal,” released August 26 with 67 enhancements, now natively supports scaling workloads to zero replicas via the Horizontal Pod Autoscaler (HPA) in beta, eliminating the need for K…
Stacklok has released ToolHive, an open-source platform under Apache 2.0 that runs Model Context Protocol (MCP) servers inside containers, isolating them from host credentials and network access to en…