Apache Hadoop is a collection of open-source software utilities for reliable, scalable, distributed computing, providing a framework for distributed storage, the Hadoop Distributed File System, and distributed processing of very large datasets using the MapReduce programming model. Doug Cutting and Mike Cafarella created it starting in 2006, building on ideas published by Google about the Google File System and MapReduce, and naming the project after Cutting's son's toy elephant. Hadoop made it practical for ordinary organizations to store and process web-scale data across clusters of commodity hardware, and it became the foundation of the broader big-data software ecosystem that followed it. This description is adapted from Wikipedia contributors under CC BY-SA 4.0; changes were made. https://creativecommons.org/licenses/by-sa/4.0/
Facts
SignificanceAn open-source framework that made distributed storage and MapReduce processing of very large datasets practical on clusters of ordinary hardware. 1 Connections
Sources
1. Wikipedia: Apache Hadoop
Wikimedia FoundationLead sectionQuote, Lead section
Apache Hadoop is a collection of open-source software utilities for reliable, scalable, distributed computing. It provides a software framework for distributed storage and processing of big data using the MapReduce programming model.
View the Source Reader Challenges (0)
No disputes yet. Spotted an error or a better source? Open the first one.
Sign in to dispute this or suggest a correction.