No.76

Data Founders

Reynold Xin

Co-Founder & Chief Architect, Databricks

The engineer who wrote much of Apache Spark and now architects Databricks' AI platform.

Score 79/100

Why they’re on the list

Xin was one of the principal creators of Apache Spark and has been Databricks' chief architect since its founding, giving him a rare combination of deep technical authorship and sustained influence over the direction of one of the world's most valuable data and AI platforms.

Reynold Xin is a co-founder and the chief architect of Databricks, and one of the principal engineers behind Apache Spark, the open-source distributed data-processing engine that underpins much of the modern big-data and AI stack. He holds a Ph.D. in computer science from the University of California, Berkeley, where he studied under Ion Stoica and Michael J. Franklin at the AMPLab, and a BA.Sc. from the University of Toronto.

As a doctoral researcher, Xin became one of the most prolific contributors to the Apache Spark project, taking a lead role in designing and building GraphX, Spark's graph-processing library, and driving 'Project Tungsten', a major initiative to rework Spark's execution engine for far greater memory and CPU efficiency. He also co-designed Spark's DataFrames API and served as release manager for the landmark Spark 2.0 release, work that helped Spark supplant Hadoop MapReduce as the default engine for large-scale data processing.

In 2013, Xin joined Matei Zaharia, Ion Stoica, Ali Ghodsi and other Spark creators in co-founding Databricks to commercialise the technology. As chief architect, he has remained close to the technical core of the company through its transformation from a Spark-hosting service into a full 'lakehouse' platform for data engineering, analytics and machine learning, and more recently into a key infrastructure provider for enterprise generative AI.

Xin's engineering credentials extend beyond Spark: in 2014 he led the Databricks team that set a world record in the Daytona GraySort benchmark, sorting 100 terabytes of data faster than any system before it, including Hadoop-based clusters many times larger. He has been repeatedly recognised as one of the most-cited researchers in the database systems field, reflecting both his academic output and the industry-wide adoption of the systems he helped build.

As Databricks has grown into one of the most valuable private technology companies, with valuations reported above $100 billion through 2025 and 2026, Xin has continued to speak publicly at conferences such as the Data + AI Summit and VLDB about the evolution of the lakehouse architecture, streaming data systems, and the company's Lakebase database product, cementing his role as one of the clearest technical voices in enterprise data infrastructure.

Career timeline

  1. 2012Shark, a Spark-based SQL engine he worked on, wins Best Demo at SIGMOD
  2. 2013Co-founds Databricks alongside the creators of Apache Spark
  3. 2014Leads Databricks team to a Daytona GraySort world record; GraphX merges into Spark
  4. 2016Serves as release manager for Apache Spark 2.0
  5. 2021Recognised as a top-cited scholar in the database systems field
  6. 2025Databricks valuation surpasses $100 billion
  7. 2026Speaks at VLDB 2026 on lakehouse, streaming and Lakebase innovations at Databricks

Sources

  1. Reynold Xin — Wikipedia
  2. Reynold Xin — Databricks Data + AI Summit speaker page
  3. Building for the AI Era: Lakebase, Streaming, and Lakehouse Innovations at VLDB 2026 — Databricks Blog
  4. Reynold Xin — Co-Founder & Chief Architect at Databricks — The Org
  5. Reynold Xin — Crunchbase

More in Data Founders