WEKA, which supplies high-speed unstructured data access for AI and accelerated computing, has a deal with Backblaze to store data needed for re-use in its B2 cloud storage.
The two companies say GPUs need extremely fast access to data to stay fed, while the datasets, checkpoints and outputs surrounding those workloads are growing toward enormous scale and that data can need retaining. They are combining WEKA NeuralMesh for performance-sensitive AI and accelerated computing workloads with Backblaze B2 cloud object storage as the capacity layer for large datasets and retained AI assets. In effect, hot data in WEKA and cold data in B2's cloud. Data can be kept on the appropriate tier as it moves through ingestion, training, checkpointing, inference and downstream workflows. You can't keep once hot AI data on flash for ever.
Gleb Budman, Backblaze’s CEO, said: “AI teams need their GPUs fed and an infrastructure with the performance and capacity to support the full AI data workflow. WEKA has mastered the performance tier. We've spent nearly two decades doing the same for capacity storage. Together, AI teams get a qualified, complete solution to ensure fast and efficient production,.
WEKA Chief Strategy Officer Nilesh Patel, neatly put forward WEKA’s point of view: "AI workloads are stretching storage in two directions at once. GPUs need microsecond access to data to stay fed, while datasets and checkpoints are growing to exabyte scale. Our collaboration with Backblaze gives customers a validated path to both - without the cost of building and testing that integration themselves. Speed where it matters, scale wherever you need it.”
For example, raw training data (training sets, media libraries, source files) can reside in B2, move into NeuralMesh when needed for a performance-sensitive workload, and checkpoints and outputs can then return to B2 for retention and future reuse, or in situations where recovery to an earlier stage of testing is necessary. WEKA’s Snap-to-Object capability has tested with Backblaze, enabling teams to recover checkpoints or saved inference data from the B2 capacity tier. Backblaze recently signed a $335 million strategic agreement with CoreWeave for that Neocloud to use B2 cloud storage for retained AI data. Coreweave is the fourth major AI cloud infrastructure company to contract with Backblaze in this way.
Certification of B2 Cloud Storage for NeuralMesh is underway. Customers can contact Backblaze or WEKA to get started.
Bootnote
Scality and WEKA set up a similar deal to the WEKA-Backblaze one, using WEKA’s NeuralMesh high-performance storage with Scality RING’s cost-efficient object tier, in February this year.
Backblaze competitor Wasabi is also providing cloud storage for AI. It has formed a dedicated AI business, and appointed semiconductor and storage veteran Pinaki Mukherjee as SVP and GM to lead it.
It says it stores hundreds of petabytes of AI data for frontier model labs, generative AI startups, and Neocloud compute providers worldwide. The new AI business will bring dedicated strategy, partnerships and go-to-market focus for this demand as AI infrastructure matures.
Wasabi raised $70 million for AI cloud storage funding in January, and subsequently arranged a $250 million credit facility.