Professional Documents
Culture Documents
SEMINAR 2005
A Technical Presentation On
DATA MINING
P.CHANAKYA. Y9MC95008, IV Semester, II MCA, St.Anns College of P.G. Studies, Chirala 523 187, Contact: funnychanu@gmail.com www.funnychanu.weebly.com/seminor
ABSTRACT
Commercial databases are growing at unprecedented rates. This explosive growth in stored data has generated an urgent need for new techniques and automated tools that can intelligently assist us in transforming the vast amounts of data into useful information and knowledge. Data mining is an automated or convenient tool for the extraction of patterns representing knowledge implicitly stored in large databases, data warehouses and other massive information repositories.
Here we describe basic ideas behind the Data mining and Warehousing, the need for Data mining & Warehousing, the actual process of Data mining, the architecture of Data mining system, the different techniques involved, the architecture of Data Warehouse, the advantages and problem areas in Data Warehousing.
Data mining:
Data Mining is the extraction of hidden predictive information from large databases.
automated, prospective analyses offered by Data mining move beyond the analyses of past events provided by retrospective tools typical of decision support systems. Data mining tools can answer business questions that traditionally were too time consuming to resolve. They scour databases for hidden patterns, finding predictive information that experts may miss because it lies outside their expectations.
The starting point in this architecture is Data Warehouse containing a combination of internal data tracking all customer contact coupled with external market data about competitor activity. An OLAP (OnLine Analytical Processing) server enables a more sophisticated enduser business model to be applied when navigating the Data Warehouse. The Data Mining Server must be integrated with the Data Warehouse and the OLAP server to embed ROI-focused business analysis directly into this infrastructure. Integration with the Data Warehouse enables operational decisions to be directly implemented and tracked. As the warehouse grows with new Decisions and results, the organization can continually mine the best practices and apply them to future decisions.
The phases depicted start with the raw data and finish with the extracted knowledge, which was acquired as a result of the following stages:
Selection
Selecting or segmenting the data according to some criteria, e.g. all those people who own a car, in this way subsets of the data can be determined.
Preprocessing
This is the data cleaning stage where certain information is removed, which is deemed unnecessary and may slow down queries e.g., it is unnecessary to note the sex of a patient when studying pregnancy. Also the data is reconfigured to ensure a consistent format as there is a possibility of inconsistent formats because the data is drawn from several sources. e.g. sex may recorded as f or m and also as 1 or 0.
Transformation
The data is not merely transferred across but transformed in that overlays may added such as the demographic overlays commonly used in market research. The data is made usable and navigable.
Data mining
data.
discover
construct
D1
subsets
descriptions
D2
D3
Clustering and Segmentation basically partition the database so that each partition or group is similar according to some criteria or metric.
2.2. Rule induction: A data mine system has to infer a model from the
database i.e. it may define classes such that the database contains one or more attributes that denote the class of a tuple i.e. the predicted attributes while the remaining attributes are the predicting attributes. Class can then be defined by condition on the attributes. When the classes are defined the system should be able to infer the rules that govern classification, in other words the system should find the description of each class.
Sales forecasting Industrial process control Customer research Data validation Risk management Target marketing etc.
Transparency Accessibility Consistent reporting performance Client/server architecture Generic dimensionality Dynamic sparse matrix handling Multi-user support Unrestricted cross dimensional operations Intuitative data manipulation Flexible reporting Unlimited dimensions and aggregation levels
10
Data
Warehouse:
Data
Warehouse
is
Subject-oriented,
Data Warehousing Systems: Legacy Systems: Any information system currently in use that was built
using previous technology generations. Most legacy Systems are
operational in nature, largely because the automation of transactionoriented business process had long been the priority of IT projects.
Source Systems: Any system from which data is taken for a data
warehouse. A source system is often called a legacy system in a mainframe environment.
contains subject oriented, volatile, and current enterprise-wide Reports detailed information.
EXTRACT CLEAN TRANSFOR M LOAD REFRESH
Data Mining
DW
Operational Systems
11
Data Warehouse Holds historic data Data is largely static Read only accesses Adhoc complex queries Analysis driven Subject oriented Used by top managers for analysis Denormalized data for model queries of the
(Dimensional model) Must be optimized for writes and Must be optimized small queries. involving a large
portion
warehouse.
Potential high Return on Investment Competitive Advantage Increased Productivity of Corporate Decision Makers
12
CONCLUSION
Comprehensive data warehouses that integrate operational data with customer, supplier, and market information have resulted in an explosion of information. Competition requires timely and sophisticated analysis on an integrated view of the data. However, there is a growing gap between more powerful storage and retrieval systems and the users ability to effectively analyze and act on the information they contain. Both relational and OLAP technologies have tremendous capabilities for navigating massive data warehouses, but brute force navigation of data is not enough. A new
13
technological leap is needed to structure and prioritize information for specific end-user problems. The data mining tools can make this leap. Quantifiable business benefits have been proven through the integration of data mining with current information systems, and new products are on the horizon that will bring this integration to an even wider audience of users.
REFERENCES
1. Building the Data Warehouse 2. Data Warehouse Architecture and Implementation 4. http://www.thearling.com/dmintro/dmintro.htm 5. http://www.sdgcomputing.com/glossary.htm - W.H. Inmon. - Barry Devlin.