Scality's AI Inferencing Factory accelerates inferencing Scality launched its AI Inference Factory, an on-premises inference stack that pools a shared KV cache across GPU servers and separates prefill from decode to improve GPU efficiency, CEO Jérôme Lecat said. The stack combines validated open-weight models, a disaggregated vLLM/Dynamo serving layer, a control plane that authenticates, meters and routes requests, and Scality's Autonomous Data Infrastructure (ADI) object storage, which tiers data across TLC, HDD and optionally tape in a single namespace and serves as a multi-petabyte KV cache accessed via S3-over-RDMA. Scality positions the offering as a supported alternative to cloud inference for enterprises, government agencies and neo-cloud providers seeking data sovereignty and protection from per-token pricing and provider-controlled model version and quantization changes. OBJECT Scality's AI Inferencing Factory accelerates inferencing Scality’s https://www.blocksandfiles.com/object/2026/09/10/scalitys-maestro-manages-artesca-fleets/5295470 AI Inference Factory provides a shared KV cache pool across on-premises GPU servers, catching up with Cloudian and MinIO, but adding prefill and decode separation for better GPU efficiency. This AI Inference Factory brings together validated open-weight models, a disaggregated inference-serving layer, and Scality AI Data Infrastructure ADI , which uses policy-governed autonomous operations to manage massive datasets across the AI lifecycle.It gives enterprises, government agencies and neo-cloud providers a supported alternative to cloud-based AI services, without having to assemble, integrate and maintain the entire software stack themselves. It puts ADI object storage under a vLLM/Dynamo serving layer as a shared KV cache, with S3-over-RDMA and a native KV connector. Scality co-founder and CEO Jérôme Lecat said: “With AI moving into mission-critical production environments, organisations need greater control over where inference runs, how their models are managed and what happens to their data. For 15 years, Scality has built data infrastructure that thousands of customers around the globe rely on to operate 24/7. AI Inference Factory brings that experience to on-premises AI, giving organisations the reliability and sovereignty they need to run critical AI workloads on their own terms.” The company says on-prem inference has data sovereignty and potential cost advantages as, with the cloud-based inference model and per-token pricing, the more useful an AI workflow becomes, the more it costs. Also model versions and quantisation can change at the provider's discretion, affecting the workflows built on them. Scality’s AI Inference Factory supports specialised chatbots, AI-assisted software development, agentic applications and other inference-intensive workloads. It combines 4 software layers in a stack: - Validated open-weight models maintained as part of the supported stack; - Disaggregated inference serving that separates prefill from decode so each can scale independently; - A control plane that authenticates and meters requests, routes them to the appropriate GPUs holding the relevant context and schedules workloads against SLA targets; and - Scality ADI , which provides high-performance object storage for models, enterprise data and inference state. Scality’s ADI https://www.blocksandfiles.com/object/2026/05/12/scalitys-autonomous-data-infrastructure-does-agent-driven-tiering-and-more/5238809 Autonomous Data Infrastructure was launched in May as an enterprise-focussed data infrastructure management product called ADI Autonomous Data Infrastructure to place data in four performance, cost, and protection storage tiers with policy-driven AI agent workers. ADI integrates with standard AI stacks and deploys at multi-petabyte scale on NVMe, with RDMA access to hard disk drives for virtually unlimited KV cache capacity. A single namespace automatically tiers data across TLC, HDD and, optionally, tape, allowing organisations to align storage performance and economics with different stages of the AI data lifecycle. It says ADI serves as the shared storage layer for the AI Inference Factory, extending key-value KV cache beyond GPU high-bandwidth memory HBM capacity limits. This enables an organization’s AI inference infrastructure to persist and retrieve model context instead of requiring GPUs to recompute it, improving GPU utilisation and reducing inference costs. Scality clams ADI provides a shared, multi-petabyte cache that GPUs read at latency the same order of magnitude as GPU memory, meaning the operation’s latency is within roughly the same factor-of-ten range as a GPU’s memory access, but obviously slower, “allowing them to retrieve existing context and remain focused on inference.” We understand this is the same as every other KC caching scheme. Co-founder and CTO Giorgio Regni said: “The KV cache on Scality ADI is fast enough to sit in the serving path. Restoring a context from ADI is the same order of magnitude as GPU memory, and 14 times faster than recomputing it, with the GPUs staying busy the whole time. Storage is no longer the reason to keep the KV cache inside the GPU server.” He added: “That opens the door to disaggregated serving. One pool of GPUs handles prefill, processing the prompt and writing the KV cache to ADI. A second pool handles decode, reading the cache back and generating tokens. With Scality ADI as the shared cache, any decode GPU can pick up any context, and there is essentially no size limit. Prefill no longer interrupts decode, and both pools run at full load.” As Regni says, the AIIF architecture separates AI Inferencing prefill from decode, with ADI acting as the shared cache between GPU pools. Prefill computes the KV cache for a new prompt, and decode generates the answer, one token at a time. When both run on the same GPUs, every new prefill interrupts the decodes already running. Once split, each pool does one job full time and scales on its own, with both working from the same shared KV cache. This allows available decode GPUs to access existing context without tying it to a specific GPU server. Scality cites two examples of this separation improving inference performance: - DistServe OSDI 2024 measured up to 7.4 times more requests served within the same latency, moving the cache directly between GPUs. In Scality's design, prefill writes the KV cache to Scality ADI, and decode reads it back, so any decode GPU can pick up any returning context, and the tier puts essentially no limit on the size of this cache. - Moonshot AI's Mooncake reported 75 percent more requests on Kimi's production traffic by adding a shared KV cache pool. Scality testing demonstrated that ADI can reduce GPU consumption while maintaining the same level of performance: - 1.9-second load time for Gemma-3 27B, almost 10x faster than local NVMe, with the model streamed in parallel across the cluster over RDMA; - 166 ms warm time-to-first-token on a 14K-token context restore from ADI, only 83 ms behind HBM; - 14x faster KV cache retrieval than recomputation on a 14K-token context and up to 72x faster on a 439K-token context; - A KV cache more than 80x larger than a single GPU’s memory, keeping up to 1,000 concurrent sessions resumable without recomputation; - No measurable impact on token generation, as context is restored before the first token; - 97 percent of network line rate for data transfer between GPUs and storage. Scality validates and maintains the complete stack. It’s compatible with open-source harnesses and agentic frameworks such as OpenCode, Hermes, Goose, LangGraph and Pydantic AI, and runs validated open-weight models such as Mistral, Gemma, gpt-oss, Qwen, Kimi, GLM, and DeepSeek. The software deploys on standard servers from Dell, HPE, Lenovo and Supermicro. The stack components are open code, meaning customers can inspect how inference state is stored and moved, and submit contributions for Scality review. Scality AI Inference Factory packages the whole serving stack behind one endpoint. It includes vLLM, Dynamo, the Scality AI Connector, Scality ADI, and more. This AI Inference Factory https://www.scality.com/ai-inference-factory is available now as a software license or as a fully managed service. Check out a technical deep dive into the AI Inference Factory architecture and test results here https://www.solved.scality.com/shared-kv-cache-ai-inference . Nvidia KV caching Virtually all filesystem storage suppliers are supporting Nvidia's STX KV Caching https://www.blocksandfiles.com/ai-ml/2026/03/30/nvidia-and-its-partners-kv-cache-extenders/5209284 scheme. This specifies hardware configurations requiring BlueField-4 DPUs, Spectrum-X networking, and dedicated CMX https://www.blocksandfiles.com/ai-ml/2026/03/30/nvidia-and-its-partners-kv-cache-extenders/5209284 flash storage tiers. Think DDN, Everpure, HPE, Hitachi Vantara, IBM, NeApp, Nutanix, VAST Data and WEKA. Object storage S3 Nvidia CMX partners include Cloudian and MinIO with its petabyte-scale MemKV https://www.blocksandfiles.com/ai-ml/2026/05/12/minio-adds-petabyte-scale-memkv-cache-for-nvidia-gpu-inference/5238593 caching system. Cloudian’s HyperScale AI Data Platform https://www.blocksandfiles.com/object/2026/06/09/cloudian-closes-gap-between-enterprise-ai-ambitions-and-messy-production-deployments/5252816 AIDP is a turnkey, on-premises appliance that’s an alternative to public cloud S3 data sources. It’s an S3-compatible, RDMA-based, data store for AI models and agents running on Nvidia Blackwell GPU hardware and software and BlueField DPUs, and complying with Nvidia’s AI Data Platform reference architecture. It lets customers run production AI on their own IT infrastructure, keeping sensitive data under their direct control. Cloudian says it saves up to 60 percent on cost by avoiding the recurring token, egress, and inference fees that make public cloud AI expensive at scale. This is the same kind of message as Scality’s but without the prefill-decode separation. MinIO says that, with MemKV, an entire GPU cluster can access a common pool of context at microsecond latencies that keep pace with inference, rather than waiting on millisecond-latency. Again, this is the same kind of message as Scality’s but without the prefill-decode separation. Bootnote Mooncake is the serving platform for Kimi https://kimi.ai/ , an LLM service provided by Moonshot AI https://www.moonshot.cn/ . Both the Transfer Engine and Mooncake Store are open-sourced. Mooncake features a KV Cache-centric disaggregated architecture that separates prefill and decode clusters. It also leverages underutilized CPU, DRAM, and SSD resources in GPU clusters to build a disaggregated KVCache pool.