No.47

Open Source

Doug Cutting

Creator of Lucene and Hadoop; formerly Chief Architect, Cloudera

Built Lucene, Nutch and Hadoop, the trio that made web-scale big data possible.

Score 84/100

Why they’re on the list

Cutting created Lucene, Nutch and Hadoop, the projects that made distributed, commodity-hardware processing of web-scale data possible and launched the big-data industry.

Doug Cutting is the American software engineer whose successive open-source projects — Lucene, Nutch and, above all, Hadoop — gave the industry the foundational tools for indexing and processing data at web scale, effectively inventing the 'big data' era of the mid-2000s. A Stanford graduate, Cutting cut his teeth on search technology at Xerox PARC, Excite and Apple, where he was the primary author of the V-Twin text search framework, before releasing Lucene, a high-performance open-source text search library, which became the search engine underpinning countless later projects.

With collaborator Mike Cafarella, Cutting built Nutch, an open-source web crawler intended to create a transparent alternative to proprietary search engines. The pair's effort to make Nutch scale to web-sized crawls led them to implement the distributed file system and MapReduce processing model described in Google's landmark research papers, work that Cutting spun out in December 2004 as Apache Hadoop, reputedly named after his son's toy elephant.

Hadoop's distributed storage and processing model let ordinary organisations analyse datasets far beyond the capacity of a single machine, using commodity hardware rather than specialised supercomputing infrastructure. Cutting developed it full-time at Yahoo!, where it was battle-tested at enormous scale, before it was released as an Apache project and adopted throughout the industry, seeding an entire ecosystem of tools — HBase, Hive, Pig, Spark and others — and giving rise to companies including Cloudera and Hortonworks.

Cutting was a founding member of Cloudera, the commercial Hadoop distribution company, serving as its chief architect for many years while remaining deeply involved in Apache Software Foundation governance: he was elected to the ASF board in 2009 and served as its chairman from 2010. He received the O'Reilly Open Source Award in 2015 in recognition of his contributions to the movement.

Though the big-data landscape has since moved on from MapReduce toward newer processing engines, Hadoop's distributed-storage paradigm (HDFS) and the broader ecosystem it spawned remain deeply embedded in enterprise data infrastructure, and Cutting is widely credited as one of the people most responsible for making large-scale, commodity-hardware data processing an industry norm rather than the preserve of a handful of internet giants.

Career timeline

  1. 1985Graduates Stanford University
  2. 1997Begins developing Lucene, the open-source search library
  3. 2002Co-founds Apache Nutch with Mike Cafarella
  4. 2004Creates Apache Hadoop, building on Google's GFS and MapReduce papers
  5. 2006Joins Yahoo! to develop Hadoop full-time
  6. 2008Becomes a founding member of Cloudera as Chief Architect
  7. 2009Elected to the Apache Software Foundation board
  8. 2010Elected Chairman of the Apache Software Foundation
  9. 2015Receives the O'Reilly Open Source Award

Sources

  1. Doug Cutting - Wikipedia
  2. Chief Architect of Cloudera on growth of Hadoop - Opensource.com
  3. Hadoop creator Doug Cutting on evolving and succeeding in open source - Medium
  4. Apache Nutch - Wikipedia
  5. Doug Cutting - O'Reilly

More in Open Source