Scality today launched Scality AI Inference Factory, an open-code software stack for running open-weight models on infrastructure the customer owns. Scality validates, ships, and maintains the whole stack, which places its AI Data Infrastructure (ADI) object storage under the inference layer as a shared key-value (KV) cache, and it is aimed at enterprises, government agencies, and neo-cloud providers that want an alternative to per-token cloud inference.
Scality frames the launch around cost and control. Per-token pricing grows with every workflow that proves useful, providers can change model versions and quantization on their own schedule, and regulated organizations have to send prompts, documents, and code to infrastructure they do not run. The stack targets coding assistants, specialized chatbots, and agentic applications, and Scality says tools such as Claude Code, Codex, and OpenCode keep working once developers point them at the local endpoint.
Disaggregated Serving Architecture and KV Caching Over RDMA #
The stack combines validated open-weight models, a serving layer built on vLLM and Dynamo that splits prefill from decode so each GPU pool scales on its own, and a control plane that authenticates and meters requests, routes them to GPUs holding the relevant context, and schedules work against SLA targets. Underneath sits ADI, which stores models, enterprise data, and inference state, connected to the GPUs through the Scality AI Connector.
ADI’s main job is holding the KV cache outside GPU high-bandwidth memory. The prefill pool computes the cache for a new prompt and writes it to ADI, and the decode pool reads it back, so any decode GPU can pick up any returning session. vLLM breaks the cache into fixed-size named blocks that ADI stores as objects, and on that path the AI Connector replaces the S3 control plane with a stripped-down one that uses single-pass authentication and drops XML and multipart framing.
On a read, the GPU reserves memory and tells ADI which object to place there. The storage server moves the object from TLC NVMe into its RAM over PCIe, and the network card writes it directly into GPU memory over RDMA, so the data bypasses the CPU, the kernel, and the software stack. Scality says that path reaches 97% of network line rate between GPUs and storage.
Scality CTO Giorgio Regni described what that changes for serving: “The KV cache on Scality ADI is fast enough to sit in the serving path. Restoring a context from ADI is the same order of magnitude as GPU memory, and 14 times faster than recomputing it, with the GPUs staying busy the whole time. Storage is no longer the reason to keep the KV cache inside the GPU server.”
Benchmarks and Internal Performance Metrics #
Scality ran its tests in a deliberately small lab: one eight-GPU server, five storage servers a few generations old, and four 100Gb/s RDMA links, which made the network the bottleneck. Scality says the product itself ran on a 400Gb/s fabric, and it did not name the GPU in the test server.
Warm time to first token on a 14K-token context restored from ADI measured 166 ms, 83 ms behind serving the same context from HBM. Recomputing that context took 2.3 seconds, so retrieval was 14 times faster. On GLM-5.2 with a 438,764-token context, recompute took 529 seconds and the read-back took 7.3 seconds, a 72x gap, and that single session produced 189GB of KV cache, about 430KB per token, more than the HBM left over with the model loaded across all eight GPUs.
In a second run with one vLLM instance per GPU, ADI held 8TB of KV cache, more than 80 times the memory of one GPU and about ten times that of the whole server. First-token latency stayed between 355 and 373 ms at every cache size from 69GB to 8TB, and inter-token latency did not move because the fetch completes before the first token. Scality says that capacity supports up to 1,000 concurrent sessions resumable without recompute.
The same path loads model weights, and Gemma-3 27B’s 54.86GB of weights loaded into one GPU in 1.9 seconds, which Scality puts at almost 10x faster than the GPU server’s local drives because the read spreads across many drives in the storage tier. That’s about 28.9GB/s, roughly 58% of the four 100Gb/s links combined, and Scality says a GPU can swap models up to a thousand times a day depending on load.
Scality also says it sent its fixes upstream. vLLM’s KV offload connector was fetching about five times the cache blocks each restore needed, and Scality corrected the connector without touching the vLLM core. It also extended the open-source ModelExpress to pull weights straight from object storage, and it says the shipping stack contains no Scality forks of any component.
Hardware Compatibility and Tiering #
Scality AI Inference Factory deploys on standard servers from Dell, HPE, Lenovo, and Supermicro. ADI runs at multi-petabyte scale on NVMe, with RDMA access to HDD for bulk KV cache capacity, and a single namespace automatically tiers data across TLC flash, HDD, and optional tape, so model weights, enterprise data, and cold cache can each sit on the media that matches how often they are read.
Validated model families include Mistral, Gemma, gpt-oss, Qwen, Kimi, GLM, and DeepSeek, and the stack is compatible with harnesses and agent frameworks including OpenCode, Hermes, Goose, LangGraph, and Pydantic AI. Every component ships as open code, so customers can inspect how inference state is stored and moved and submit changes for Scality’s engineers to review.
Scality AI Inference Factory is available now as a software license or as a fully managed service. Scality is demonstrating the stack and its reference architecture at Scality Day in Paris today.