cd /news/artificial-intelligence/from-database-to-ai-data-platform-oc… · home topics artificial-intelligence article
[ARTICLE · art-102598] src=startupfortune.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

From Database to AI Data Platform: OceanBase's Roadmap to Rival Databricks

OceanBase CTO Charlie Yang outlined the company's roadmap to rival Databricks as an AI data platform at OceanBase Hours in Singapore, arguing that databases can serve as the foundation for AI data platforms just as data lakes do for Databricks. Yang and Forrester Principal Analyst Indranil Bandyopadhyay identified four forms of convergence—multimodel data, translytical processing, in-engine governance, and SQL-centric query surfaces—that will define platforms built for AI agents rather than human users. OceanBase, which handles roughly 500,000 transactions per second during Singles' Day, released its Lakebase engine in June to extend its distributed database into open storage and multimodal data, positioning against Databricks' Lakebase architecture.

read5 min views5 publishedAug 19, 2026
From Database to AI Data Platform: OceanBase's Roadmap to Rival Databricks
Image: Startupfortune (auto-discovered)

As companies like Databricks build AI data platforms from the data lake, OceanBase is demonstrating that the database can serve as an equally powerful foundation. At this week's OceanBase Hours in Singapore, the company outlined its roadmap for this direction.

Charlie Yang, CTO of OceanBase, sat down with Indranil Bandyopadhyay, Principal Analyst at Forrester, for a conversation framed around a single question: what does a data platform need to look like once AI agents, not humans, become its primary users. Their answer was convergence, four distinct forms of it, and it doubles as a map of where OceanBase is trying to position itself against a much larger rival, Databricks, as an AI data platform.

Bandyopadhyay opened with a reversal he sees playing out across the industry. For decades, data moved to where compute lived. That assumption held because the consumer at the other end was human, tolerant of latency, working normal hours. Agents are not. They query continuously, at machine speed, and a copy of data that is a few seconds stale can already be wrong. Retrieval, inference, and governance are moving back down to the data layer itself, because that is the only place freshness, consistency, and authorization can all be enforced at once.

From there he traced four forms of convergence Forrester is tracking: data models collapsing into a shared multimodel core rather than separate engines for relational, document, and vector data; transactional and analytical systems merging into what Forrester calls "translytical," since an agent acting on a two-second-old account balance can make the wrong call; governance moving into the engine itself rather than sitting as a bolted-on layer, because agents act too continuously and unpredictably for periodic, human-paced access checks; and query surfaces consolidating around something like SQL, so an agent working around the clock is not juggling a different language for every data type it touches. Yang's response was that the same convergence is happening on the vendor side, just from different starting points. Databricks built out from the data lake and is now adding transactional and database capabilities, most notably through its Lakebase architecture, which keeps PostgreSQL compute stateless and elastically scalable while externalizing storage to independent services.

Databricks is choosing patience over a 2026 IPO Databricks CEO Ali Ghodsi says 2026 is not the right year for an IPO, even as the company sits at the center of enterprise AI demand. Its private funding gives it room to keep building without public-market pressure. - why databricks delayed its 2026 IPO plans - AI company postponing public market debut strategy

OceanBase is running the route in reverse, and arriving at a similar destination. It started as a shared-nothing distributed database built inside Ant Group in 2010, proven at a scale few competitors can claim, handling roughly half a million transactions per second during events like Singles' Day, and has since added HTAP so that transactions and real-time analytics can run against the same live data. Its own Lakebase engine, released this past June, extends that foundation further into open storage, multimodal data, and hybrid search, with a semantic layer on top made up of DataStudio for data engineers and DataPilot for analysts translating natural-language questions into queries against enterprise data.

Yang's advice on how buyers should justify any of this was pragmatic. Both he and Bandyopadhyay cautioned against choosing a platform purely on AI grounds. Given how difficult AI ROI still is to demonstrate, the safer path is selecting a platform for the workloads a company already has, then making sure it is ready for agentic AI rather than requiring a second migration later.

Yang offered three practical starting points: find a real business pain point rather than a purely technical one, pick either a net-new AI-native application or an existing analytical workload rather than attempting a full legacy migration up front, and start with the smallest slice of high-value data rather than everything at once. He pointed to Alipay as a real-world example. In the first phase, less than 5 percent of total data was migrated to OceanBase's AI data platform, yet it ended up serving more than 90 percent of hot queries, with the remaining cold data left untouched in the legacy system.

The comparison to Databricks is not incidental. Materials shared around the event, along with a Bloomberg report from late July, described OceanBase in talks to raise roughly $300 million to $443 million in funding, with investors said to be measuring it against Databricks as a benchmark. OceanBase's revenue reportedly crossed $200 million on an annualized basis this year, up 70 percent from 2025, and IDC has credited it with the largest share of China's distributed database market in 2025. Its client base, still concentrated among Chinese enterprises including the Industrial and Commercial Bank of China and China Mobile, is now expanding into Southeast Asia, Japan, India, and Latin America. Databricks, for comparison, said in February it was on track for $5.4 billion in annual revenue and has been raising capital at a valuation near $188 billion. The gap in scale is significant, but OceanBase's trajectory suggests it is building toward a similar destination.

What the fireside chat made clear is that OceanBase is not trying to bolt AI features onto an existing database. It is arguing that the database's original engineering demands, consistency, reliability, real-time access, and scale, become more important, not less, once an AI agent is the one relying on them. This strategy positions OceanBase not as a follower in the AI data platform race, but as a company building toward a converged architecture, one where the database serves as the foundation for both transactional workloads and the next generation of AI infrastructure, and where OceanBase itself could become the kind of platform that Databricks is today.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @oceanbase 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/from-database-to-ai-…] indexed:0 read:5min 2026-08-19 ·