Log in Sign up
Back to Discover
💻

Data mining

technology Maturity 11-13

Computers look at lots of facts.

Spurious correlations - spelling bee spiders.svg
Spurious correlations - spelling bee spiders.svg
They find patterns in the facts. This helps us learn new things. It can help a store sell more. It is like a treasure hunt. Do you like to find patterns?

43 words

Computers look at huge piles of facts.

Spurious correlations - spelling bee spiders.svg
Spurious correlations - spelling bee spiders.svg
They search for hidden patterns. This is called data mining. It is like a treasure hunt for information.

One way is finding groups. The computer finds things that are similar. Another way is finding odd things. It looks for facts that are not normal.

It can even find links. A store might see what people buy together. This helps them know what to sell.

Sometimes the computer makes a mistake. It might find a pattern that is not real. People must check the work to be sure.

Finding these patterns helps us learn. It turns many facts into smart ideas.

113 words

Computers can look at huge piles of facts. They search for hidden patterns in these sets. This work is called data mining.

Spurious correlations - spelling bee spiders.svg
Spurious correlations - spelling bee spiders.svg

Data mining is not like digging for gold. Instead, it finds patterns in large amounts of data. It uses math and computer science to help. One way is called cluster analysis. This finds groups of things that are similar. Another way is anomaly detection. This finds unusual or odd records.

Stores use this to learn about shoppers. They use association rule learning to find links. This shows which items people buy together. A computer can also use classification. This helps sort things, like marking spam emails.

Sometimes, the computer makes a mistake. It might find a pattern that is not real. This is called overfitting. The pattern might only work for one group. It may not work for others.

Spurious correlations - spelling bee spiders.svg
Spurious correlations - spelling bee spiders.svg

To fix this, experts use a test set. This is a new group of data. They check if the pattern works there too. If it works, they turn the patterns into knowledge. This helps us make better guesses about the world.

194 words

Data mining is a way to find hidden patterns in huge piles of information. Scientists use it to turn messy data into useful knowledge. This work sits at the meeting point of math, computer science, and database systems. It is a key part of a bigger process called Knowledge Discovery in Databases, or KDD. People often think it means digging for data, but that is not quite right. The real goal is to find the patterns hidden inside the data.

Spurious correlations - spelling bee spiders.svg
Spurious correlations - spelling bee spiders.svg

This process works in several clear steps. First, experts must collect a large set of data from a warehouse. Next, they perform pre-processing to clean the information. This means they remove mistakes or missing parts from the set. Then, the actual data mining happens using special computer tools. These tools look for things like unusual records or groups of similar items. Finally, experts check the results to see if the patterns are real.

Spurious correlations - spelling bee spiders.svg
Spurious correlations - spelling bee spiders.svg

People have looked for patterns in data for a very long time. In the 1700s, mathematicians used Bayes' theorem to help. Later, in the 1800s, they used regression analysis. The term "data mining" appeared around 1990 in the database community. Before that, some people used the term "database mining." However, a company in San Diego named HNC had a trademark on that phrase. Because of this, researchers started using the name we use today.

There are many ways computers find these patterns. One way is called cluster analysis, which finds groups of similar things. Another way is anomaly detection, which finds odd or unusual records. Some stores use association rule learning to see which items people buy together. For example, a shop might see that people buy certain foods at once. This is often called market basket analysis. Computers also use classification to sort things, like marking spam emails.

Spurious correlations - spelling bee spiders.svg
Spurious correlations - spelling bee spiders.svg

Sometimes, a computer can make a mistake called overfitting. This happens when a pattern works for one group but not for the whole world. It is like finding a fake link between spelling bee winners and spider deaths. To stop this, experts use a test set of new data. They apply their patterns to this new set to see if they still work. If the patterns are correct, they become useful knowledge. This helps people make much better predictions about the future.

Spurious correlations - spelling bee spiders.svg
Spurious correlations - spelling bee spiders.svg

408 words

Data mining is a specialized process used to extract and identify patterns within massive data sets. This field exists at the intersection of three major areas: machine learning, statistics, and database systems. While the name suggests digging for data, that is actually a misnomer. The real goal is to find knowledge and patterns hidden inside large amounts of information. This process is a critical analysis step within a broader framework known as Knowledge Discovery in Databases, or KDD. By transforming raw information into a structured and comprehensible format, data mining allows for much deeper insight and future use.

The KDD process follows a specific sequence of stages to ensure the results are useful. First, researchers perform selection to gather a target data set, often from a data warehouse or data mart. Next comes pre-processing, which is a vital step for cleaning the data. During cleaning, experts remove noise or observations that contain missing data. After the data is prepared, the actual data mining occurs using various mathematical algorithms. Finally, the process moves to interpretation and evaluation. This last step ensures that the discovered patterns are actually meaningful and not just random coincidences.

There are several distinct types of tasks that data mining algorithms perform. One common task is anomaly detection, which identifies unusual records or deviations from the standard range. Another is association rule learning, which looks for dependencies between different variables. A famous example of this is market basket analysis, where a supermarket studies which products are frequently bought together. Clustering is another method used to discover groups of similar data without using predefined structures. Additionally, classification can be used to sort data, such as an email program labeling a message as "spam" or "legitimate."

To ensure accuracy, data miners must be careful of a problem called overfitting. Overfitting occurs when an algorithm finds a pattern that works perfectly in a specific training set but fails to appear in the wider world. This can lead to results that look significant but cannot actually predict future behavior. To prevent this, experts use a test set of data that the algorithm has never seen before. By applying the learned patterns to this new set, they can measure the true accuracy of the model. If the patterns do not meet the required standards on the test set, the researcher must go back and change the pre-processing or mining steps.

The history of finding patterns in data stretches back centuries. In the 1700s, mathematicians used Bayes' theorem to identify relationships. By the 1800s, regression analysis became a common tool for studying data. The specific term "data mining" emerged around 1990 within the database community. Before this, the phrase "database mining" was used, but it was trademarked by a company called HNC in San Diego. As a result, researchers adopted the term "data mining" instead. While terms like "data dredging" or "data fishing" were used critically in the 1960s and 1980s, the modern term now carries a positive meaning in business and science.

Data mining is highly significant because of the sheer scale of modern information. As computer technology has grown more powerful, the ability to store and manipulate data has increased dramatically. This has allowed for the use of advanced tools like neural networks, genetic algorithms from the 1950s, and support vector machines from the 1990s. These technologies allow us to bridge the gap between pure mathematics and practical database management. By using spatial indices and other database techniques, we can apply complex learning algorithms to datasets that are far too large for manual analysis.

Understanding data mining also requires distinguishing it from general data analysis. Data analysis is often used to test specific hypotheses or models, such as checking if a marketing campaign worked. In contrast, data mining uses machine learning to uncover "clandestine" or hidden patterns that were not previously known. This makes it a powerful tool for predictive analytics and decision support systems. By identifying multiple groups within a dataset, data mining can help systems provide much more accurate predictions for the future.

673 words
🖼️ Images & Media (1)
File:Spurious correlations - spelling bee spiders.svg
Spurious correlations - spelling bee spiders.svg
Up Next
💻
Data science
Technology
More to explore

What is Nepedia?

A free, ad-free encyclopedia for children. Every article is written at five reading levels, so the same page works for a five-year-old and a fifteen-year-old — use the level switcher above to see this one change. No account needed to read.