cd /news/ai-infrastructure/ai-ready-data-anywhere-a-new-archite… · home › topics › ai-infrastructure › article
[ARTICLE · art-140871] src=lakefs.io ↗ pub= topic=ai-infrastructure verified=true sentiment=↑ positive

AI-Ready Data, Anywhere: A New Architecture for Control, Access, and Infrastructure

LakeFS, Attimis, and Seagate detailed a joint architecture that combines lakeFS's data control plane, Attimis OneBucket's single logical S3 bucket spanning clouds, on-prem, and edge storage, and Seagate Exos FUSE on-prem storage to deliver distributed AI-ready data without copying or migrating it. The three vendors validated the stack end to end and published a working sample on GitHub, positioning lakeFS as a governed state layer that supports Parquet, Iceberg, Delta Lake, and unstructured data such as images, video, and PDFs.

by read9 min views5 publishedSep 10, 2026
AI-Ready Data, Anywhere: A New Architecture for Control, Access, and Infrastructure
Image: Lakefs (auto-discovered)

How lakeFS, Attimis OneBucket, and Seagate Exos bring together data control, global access, and on-prem infrastructure to deliver trusted, distributed, AI-ready data. #

Enterprise AI doesn’t stall because the models are missing. It stalls because the data isn’t ready: fragmented across clouds, regions, and on-prem systems; hard to version; hard to govern; hard to reproduce. The teams who need it (data engineers, ML engineers, platform owners) spend their week moving and reconciling data instead of using it.

We’ve been working together on a stack that addresses this directly. It separates three concerns that most teams have been forced to entangle:

Data control — versioning, lineage, governance, safe promotion. This is lakeFS, the control plane for AI-ready data.

Data location — one logical storage bucket that spans physical storage across clouds, on-prem, and edge, without copying or migrating. This is Attimis OneBucket™, global access to data anywhere.

Data platform — dense, durable, on-prem storage where data lives when economics, performance, or sovereignty require it. This is Seagate Exos FUSE storage.

Below I’ll walk through each layer, what the three of us validated end to end, and how to run it yourself. A working sample lives on GitHub, linked at the bottom.

Data control: lakeFS, the control plane for AI-ready data #

When we describe lakeFS as the control plane for AI-ready data, the concrete meaning is this: lakeFS sits as a governed state layer between any compute engine (training jobs, RAG pipelines, agents, analytics) and your object storage. It doesn’t copy data. It doesn’t require migration. And it works with the formats your teams already use, from Parquet, Iceberg, and Delta Lake to unstructured data like images, video, and PDFs. We went deeper on that governed state layer in Trust But Verify: Securing Data Access for AI Agents.

For an AI or data platform team, that translates into capabilities that look familiar to anyone who’s used Git, only applied to data: Isolated branches for experimentation, so a team can validate a pipeline or model against production data without risking it.

Atomic commits and reproducible snapshots: every model trained against a precise commit, every experiment re-playable months later for an audit.

Lineage and auditability built in (lakeFS Enterprise), so the evidence model governance asks for comes from the system of record instead of a second tool bolted on beside it.

Write-Audit-Publish workflows and branch protection rules that block bad data before it ships, the data equivalent of a CI/CD gate.

Instant rollback when something does ship wrong, without restoring from backup.

AI-ready data isn’t only clean data — it’s data that can be safely changed. Most data stacks are good at storing and moving data, and bad at managing change. lakeFS closes that gap.

But a control plane is only as useful as the storage it governs. Because lakeFS speaks S3 to its object store, it can govern data on any S3-compatible object store. This is where lakeFS, Attimis, and Seagate come together, combining data control, location-independent access, and infrastructure to support distributed AI without forcing customers into a single place to keep their data.

Data location: Attimis OneBucket #

Stop moving data. Start using it. The constraint most AI teams hit isn’t capacity. It’s location. Training data sits in one cloud. Inference runs in another region. Sensitive data has to stay on-prem. An edge site collects telemetry that eventually needs to be joined with corporate datasets. Each is solvable with enough plumbing: pipelines, replication jobs, custom caching, sync tooling, time and money. But this is how AI projects stall or die.

Attimis OneBucket presents one logical S3 bucket that spans the underlying data storage. Built on Attimis’s Adaptive Data Fabric™, OneBucket presents a standard S3 interface that provides local access to data anywhere, across clouds, on-prem, and the edge, with no application changes and no rearchitecting of workflows. It deploys where the customer needs it: on-prem, air-gapped, in a VPC, or as SaaS. There’s no central bottleneck, because a stateless architecture scales to hundreds of petabytes.

For AI workloads, this has immediate value: Data stays where it should live. You don’t have to migrate data to start using it. Point lakeFS to OneBucket and all data becomes accessible as one global namespace.

Compute moves to the data. A training job in US-East reads petabytes of data sitting in Frankfurt, with no copy, no pipeline, and no wait. OneBucket handles locality and caching beneath it.

Sovereignty and residency become a policy decision rather than an architecture project. OneBucket keeps data where it belongs with no architecture changes or application rewrites.

Pairing OneBucket with lakeFS finally separates two concerns that AI teams have been forced to entangle. lakeFS gives the AI team operational control over what the data is. OneBucket gives them flexibility over where it lives. The two concerns finally get to be independent, which is exactly what distributed AI organizations need.

Data platform: Seagate Exos #

Cloud-only is no longer the default for enterprise AI. Data gravity, unpredictable cloud economics, and rising sovereignty requirements are forcing organizations to make deliberate decisions about where data lives, and where AI runs. As datasets grow into the petabyte and exabyte range, moving data becomes the bottleneck. The hard part is now placing data where it can be used efficiently.

Seagate Exos is built for that reality. Across the Exos portfolio, Seagate delivers systems purpose-built for AI-scale data infrastructure:

  • Exos SCALE provides extreme-density capacity for large-scale data lakes and training datasets.
  • Exos PROTECT adds RAID-based, hardware-enforced protection with five-nines availability for critical data assets.
  • Exos FUSE puts compute and storage in the same chassis, so applications and AI pipelines run where the data already sits.

This portfolio approach lets organizations match infrastructure to workload. A centralized training cluster and a remote edge site can run different Exos systems without splitting into two operating models.

Why Exos FUSE for this stack

For this AI-ready architecture, Exos FUSE is especially relevant because it converges compute and storage in a space-efficient system designed for hybrid, edge-based workloads. It ships as one consistent platform, so customers running AI-centric software like Attimis and lakeFS deploy faster and have less to operate. That compute-to-data proximity reduces latency, simplifies deployment, preserves software flexibility, and keeps sensitive data sovereign and controlled. As AI pipelines become more distributed, that proximity stops being an optimization and starts being a requirement. Exos is the dense, durable, cost-efficient storage layer where AI data can persist and scale, whether the requirement is a compact edge deployment or a larger sovereign data environment. With Exos FUSE, customers can run software close to the data to speed decisions and reduce infrastructure sprawl. With the broader Exos systems portfolio, they can align storage architecture to workload needs without locking themselves into a single operating model.

Where Exos sits under OneBucket and lakeFS

In this three-layer design, Seagate anchors the physical edge data infrastructure, Attimis OneBucket adds a location-independent S3 namespace above it, and lakeFS delivers governed versioning, auditing, lineage, and promotion. The result is a practical AI-ready data plane that helps enterprises keep data sovereign, reduce unnecessary movement, and support AI pipelines with infrastructure built for scale and control.

Attimis already runs on Seagate at multiple shared customers, so the bottom two layers of this stack are already in production. What’s new is what the validation with lakeFS adds at the top.

What we validated together #

To confirm the full stack integrates end-to-end, we built and tested a compact POC that brings together lakeFS for data control, Attimis OneBucket for location-independent access, and Seagate Exos FUSE as the on-prem compute and storage foundation.

The POC exercises the core lakeFS workflow against OneBucket-backed storage:

  1. Upload a raw dataset (a CSV file) to the main branch through the lakeFS S3 Gateway.
  2. Commit the change. The objects are written to OneBucket; the version metadata lives in lakeFS.
  3. Create an experiment branch offmain .
  4. Write a transformed version of the dataset to the experiment branch and commit.
  5. List and read objects on each branch independently. The main branch sees only the raw data, while theexperiment branch sees both raw and transformed: same repository, same underlying storage, two different logical views.

This is Git-like branching and commit semantics for data, with the actual bytes living on OneBucket (and on Seagate underneath that) the whole time. Nothing was copied. Nothing was migrated. The application layer talked to lakeFS; OneBucket handled the object I/O.

A few interoperability notes for anyone setting this up themselves:

Use the S3 endpoint issued with your OneBucket account. The sample’s .env.example shows where it goes. Put only the hostname there, not the bucket name. Path-style addressing is required.

lakeFS needs an explicit region configuration, because region auto-discovery via GetBucketRegion is not currently available.

Chunked / streaming transfer encoding has to be disabled on the lakeFS blockstore connection.

We’ve packaged the working setup as a lakeFS sample so you can run it yourself: lakefs-onebucket in lakeFS-samples.

What this pattern gives you #

There’s no single new SKU to buy here. I’m writing this up because all three of us kept hearing the same architecture described back to us, independently, by our own customers:

Trusted AI data across distributed storage. Teams branch, validate, train, and promote data through lakeFS, access it globally through OneBucket, and run AI pipelines on Seagate Exos where the data already sits.

Reproducibility and governance without storage migration. Customers get lakeFS capabilities (lineage, auditing, tracking, branching, commits, rollback, and governed promotion through Write-Audit-Publish) without committing to a single cloud or moving petabytes.

Sovereignty and regulated environments, where data has to stay on-prem or in a specific region, get an AI-ready control plane that doesn’t change the deployment model.

Multi-cloud and hybrid by default. A consistent control plane and a single namespace across clouds and on-prem, with hardware economics that hold up at petabyte scale.

Try it #

Run the POC. The [lakeFS + OneBucket sample](https://github.com/treeverse/lakeFS-samples/tree/main/01_standalone_examples/lakefs-onebucket) is end-to-end runnable in Docker.

- Learn more about [lakeFS as the control plane for AI-ready data](https://lakefs.io) , and about[building an AI-ready data architecture](https://lakefs.io/blog/ai-ready-data-architecture/) .
- Learn more about Attimis OneBucket at [attimis.co](https://attimis.co) and[request an account](https://app.onebucket.io/sign-up) .
- Learn more about [Seagate Exos integrated storage](https://www.seagate.com/products/data-storage-systems/) .

If you’re running into the data-readiness wall on an AI initiative, we’d be glad to talk, together or individually.

Iddo Avneri is Chief Business Officer at lakeFS. With thanks to Shayne Stubbs, Co-Founder of Attimis, and Nick Jarvis, Sr. Tech Alliance Lead at Seagate, who contributed the OneBucket and Exos sections.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @lakefs 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-ready-data-anywhe…] indexed:0 read:9min 2026-09-10 · —