Human writing, a luxury product
The United States is changing how it allocates research funding, with a focus on investing in individual scientists over legacy institutions and increasing funding to AI research, according to the Dir…
The United States is changing how it allocates research funding, with a focus on investing in individual scientists over legacy institutions and increasing funding to AI research, according to the Dir…
Apache Spark 4.2 introduces native vector similarity search, Metric Views for governed business metrics, and real-time Python streaming, positioning Spark as an AI-native data platform. The release ta…
Databricks raised funding at a $188 billion valuation, driven by Coatue's bet on enterprise AI governance tools that are outpacing the underlying AI models. The company, founded by UC Berkeley academi…
Apache Spark 4.2 is now available in Databricks Runtime 19 Beta, introducing governed metrics, vector and top-K primitives, an Arrow-first Python path, first-class change data capture, and stronger st…
VAST Data and Cloudera announced a strategic partnership to integrate Cloudera's data engineering, analytics, and AI services with VAST's AI Operating System, creating a combined AI factory architectu…
Researchers introduced Jailbreak, a system that uses large language models to regenerate database storage readers, bypassing traditional database drivers like JDBC and ODBC to read storage files direc…
Databricks co-founder Reynold Xin explains the limitations of traditional monolithic OLTP databases and introduces Lakebase, a serverless Postgres database that externalizes storage and compute. The a…
Google Cloud announced the preview of BigQuery's AI.AGG() function, which enables natural-language analysis of unstructured and multimodal data across millions of rows using a single line of SQL. The …
Azure Databricks provides a production-grade feature engineering pipeline for MLOps using Apache Spark, Delta Lake, and MLflow. The pipeline follows the Medallion Architecture with Bronze, Silver, and…
Databricks, a data and AI platform valued at $134 billion, is doubling down on its open-source strategy by launching OpenSharing, Omnigent, and DBRX for open LLMs at its Data + AI Summit. The company …
Databricks announced major updates to Unity Catalog at the Data+AI Summit 2026, including AI runtime governance with hard spend caps and contextual service policies, external engine writes to managed …
Databricks launched Genie ZeroOps, an autonomous AI agent that monitors production data pipelines and ML models, detects failures, diagnoses root causes, suggests fixes, and validates them in isolated…
Snowflake announced Snowflake AIM, a unified AI-powered platform to help enterprises modernize, migrate, and virtualize data and code workloads on Snowflake. The platform offers two paths: modernizati…
Intel is ending development of its open-source BigDL project, which enabled running large language models across Intel XPUs from laptops to data centers. The project will be archived on June 30, 2026,…
Google Cloud announced the general availability of Lightning Engine for Managed Service for Apache Spark, delivering up to 4.9x faster performance than standard open-source Spark and 2x the price-perf…
Google Cloud announced at Next '26 that Managed Service for Apache Spark clusters now include Lightning Engine, a native C++ vectorized execution engine that delivers up to 4.9x faster performance tha…
Google Cloud announced the general availability of its serverless Managed Service for Apache Spark runtime version 3.0, which introduces zero-setup onboarding and reduces startup times by 75%. The upd…
Quanton has launched a high-performance compute engine for Apache Spark that delivers faster execution through SIMD-vectorized processing and storage-aware planning, with no migration or code rewrites…
Python developers processing billions of rows or running distributed machine learning pipelines now have seven specialized libraries—including PySpark, Dask, and Polars—that handle datasets exceeding …
The article explains that data engineering is shifting from complex distributed clusters to single-node processing, driven by modern hardware with many cores and large memory, as well as new tools lik…