Data mining is the process of extracting and finding patterns in massive data sets using methods at the intersection of machine learning, statistics and database systems. It is an interdisciplinary subfield of computer science and statistics whose overall goal is to extract information from a data set and transform it into a comprehensible structure for further use, forming the analysis step of the wider knowledge discovery in databases process; beyond the raw analysis it also involves database and data management, data pre-processing, model and inference considerations, interestingness metrics, complexity considerations, post-processing of discovered structures, visualization and online updating. This description is adapted from Wikipedia contributors under CC BY-SA 4.0; changes were made. https://creativecommons.org/licenses/by-sa/4.0/
Facts
Disputed
Origin YearThe term itself appeared around 1990 in the database community; the statistical and analytical methods it draws on trace back much further (Bayes' theorem, 18th century; regression analysis, 19th century). Core ConcernExtracting patterns from large data sets using methods from machine learning, statistics and database systems 1 Connections
Associated With
Bioinformatics, Fields Field-to-field association; both are standard neighbouring areas of computing
Data Science, Fields Field-to-field association; both are standard neighbouring areas of computing
Source Wikipedia: Data mining
Sources
1. Wikipedia: Data mining
Wikimedia FoundationLead section, para 1
Data mining is the process of extracting and finding patterns in massive data sets involving methods at the intersection of machine learning, statistics, and database systems.
Etymology section
The term data mining appeared around 1990 in the database community, with generally positive connotations.
View the Source Reader Challenges (0)
No disputes yet. Spotted an error or a better source? Open the first one.
Sign in to dispute this or suggest a correction.