Democratizing Managed Lustre with lower cost and frictionless development Google Cloud launched a Dynamic Tier for its Managed Lustre parallel filesystem priced at 6 cents/GB per month, offering sub-ms latency for hot data and scaling linearly to 80 PB and tens of thousands of clients. Google Cloud says the tier delivers roughly 300µs average read latency, up to 4x better responsiveness than alternative distributed file systems, and a 67% improvement in aggregate throughput for many clients reading the same file. The tier keeps all data in a single namespace, promoting hot data to an SSD High-Performance Cache while demoting older checkpoints to an HDD Capacity Pool with 10 to 30 ms average read latency. This is the first of a two-part series exploring how Google Cloud is bringing the foundational values of a high-performance parallel filesystem–TB/s throughput, sub-ms latency at high client scale, and POSIX support–to a broader set of use cases and users. Historically, due to the cost and special purpose nature of parallel filesystems, colder data had to be stored outside of the filesystem and AI developers have had to maintain separate, slower environments for writing code, compiling libraries, and managing repositories. This fragmentation increases the toil of manual data staging, dataset copying, and managing disjointed namespaces. Google Cloud Managed Lustre is solving these problems through our 6 cents/GB month Dynamic Tier and by optimizing Managed Lustre performance for a range of development tasks and workloads – making Managed Lustre a “One-Stop Shop” for high-performance AI and HPC workloads. The Managed Lustre Dynamic Tier provides sub-ms latency for hot data, which allows you to store all of your data in a single namespace, and costs only 6 cents/GB month . Throughput, capacity scale and client scale: Throughput scales linearly with capacity up to 80 PB, while sub-ms latency for hot data remains stable as you scale to tens of thousands of clients. Single-flat fee: Predictable pricing. No independent charges for disk media types, data movement within the namespace, or metadata IOPS. Read Latencies: Sub-ms latencies for High-Performance Cache SSD . The Capacity Pool “HDD” is built on Google Cloud Hyperdisk throughput, which has an average read latency of 10 to 30 ms https://docs.cloud.google.com/compute/docs/disks/hd-types/hyperdisk-throughput . Multi-Epoch Training and/or Training with Optimized Fetch Sizes: Hot data is promoted to the High Performance Cache SSD after the first run. Larger data prefetch will allow you to take advantage of the Dynamic Tier cost structure and gain from low-latency SSD. Write-Heavy Checkpointing: Bursty checkpoint writes land directly in the High Performance Cache. Older checkpoints are transparently demoted to the Capacity Pool HDD . Rapid Checkpoint Restore: New checkpoints are written to the High Performance Cache, enabling low-latency checkpoint restores. Interactive Snappiness for Developers: Low-latency tasks like git cloning, compiling libraries, or running notebooks benefit from a local-disk feel ~300µs average read latencies on the same shared workspace hosting large training sets. In addition to Managed Lustre’s scalability for large AI and HPC workloads checkpoint/restart/data-loading , it also meets the demands for interactive work, meaning developers can start on Managed Lustre and stay on Managed Lustre throughout the entire workload lifecycle: Consolidates the AI and HPC lifecycle into a single namespace, providing a "local disk" feel for interactive work Read more about the latency benefits of Managed Lustre experienced by Salesforce and others https://cloud.google.com/products/managed-lustre . Latency: ~300µs average read latency—delivering up to 4x better responsiveness than alternative distributed file systems. Accelerated Setup: Untar the Linux kernel in ~2 minutes 4.7x faster than alternative file solutions , run a 20-worker parallel git clone of Python in ~40 seconds, compile Python in ~200s. Managed Lustre maximizes GPU ROI by preventing storage bottlenecks during cluster initialization. When thousands of worker nodes attempt to read the exact same file simultaneously such as a shared model checkpoint, base weights, or container layer , traditional distributed file systems can choke on localized hotspotting, leaving high-cost GPU clusters idle for minutes. Improves Aggregate Throughput for a large number of clients reading the same file: Demonstrates a 67% improvement over alternative file solutions. Parallel Loading: Imports libraries like PyTorch across 4,000+ processes in under 60 seconds. Here is the code for the tests we’ve run, so that you can perform your own testing. We used fio https://github.com/axboe/fio to emulate small, low-concurrency reads and writes: