You are on page 1of 4

INTERNATIONAL BURCH UNIVERSITY

LITERATURE REVIEW
Ali LAFCIOLU
lafcioglu@bosnasema.ba

EDUCATIONAL DATA MINING Student Retention and Attrition

Educational data mining (EDM) is an emerging discipline that focuses on applying data mining tools and techniques to educationally related data. (Baker, R., & Yacef, K., 2009). The discipline focuses on analyzing educational data to develop models for improving learning experiences and improving institutional effectiveness. Data mining can assist educational institutions with uncovering useful information in order to guide decision-making. (Kiron, D., Shockley, R., Kruschwitz, N., Finch, G., & Haydock, M., 2012). Educational data mining is a field of study that analyzes and applies data mining to solve educationally-related problems. Applying data mining this way can help researchers and practitioners discover new ways to uncover patterns and trends within large amounts of educational data. A literature review on educational data mining follows which cover topics student retention and attrition. Student attrition and retention have long been familiar terms in higher education. But until recently, most administrators have been content to acknowledge that the phenomena exist and to accept them. Elite institutions have tended to assume that attrition is a consequence of maintaining the competitive academic conditions upon which their reputation depends. For very different reasons, colleges with open-door policies have come to accept attrition as an inevitable consequence of their admission policy. In most cases, retention rates have not been a concern unless drastic attrition occurred in a particular department or program, thus signaling that some type of problem existed. (Lenning, 1980) Students who drop out of college often suffer personal disappointments, financial setbacks, and a lowering of career and life goals. Concern about the student has led to much research on college student attrition and retention. Student retention is an important issue for all university policy makers due to the potential negative impact on the image of the university and the career path of the dropouts. Research has shown that data mining can be used to discover at-risk students and help institutions become much more proactive in identifying and responding to those students. (Luan, 2002). Luan applied data mining tools as a way to predict what types of students would drop out of school, and then return to school later on. He applied classification and regression trees to educational data in order to predict which students are unlikely to return to school. In this case study, Luan applied both quantitative and techniques to uncover student success factors. This research is important because it demonstrated the successful application of data mining tools to assist retention efforts. As noted earlier, the case study method for EDM may often produce results that are not generalizable. However, the process by which researchers apply the data mining can be generalized and used in other contexts. It is simply the results of the data mining models that may not be generalized. The paper also discusses potential applications of data mining in higher education. The benefits of data mining are its ability to gain deeper understanding of the patterns previously unseen using current available reporting capacities. In a related study, (Lin, 2012) applied data mining as a way to improve student retention efforts. (Lin, 2012) was able to generate predictive models based on incoming students data. Lin model can produce short accurate prediction lists identifying students who tend to need the support from the student retention program most. The research study found that certain machine learning algorithms can provide useful predictions of student retention (Lin, 2012). Group of Researchers (Chacon, F., Spicer, D., & Valbuena, A., 2012) developed an analytics system that Bowie State University (BSU) has implemented and is enhancing in order to

improve the retention and success of at-risk students. Their system helps the educational institutions identify and respond to at-risk students. Their research contributes meaningfully to the EDM literature because it demonstrates a successful implementation and use of data mining. Their work is highly representative of the discipline in that it follows a strict data mining process and it is quantitative. Their research supports other work done in applying data mining to student retention issues, such as (Lin, 2012) and (Luan, 2002), all with successful results. The work goes one step further than Lin and Luan, because the researchers were able to develop and implement their solution in a production environment. BSU uses the system to aid in student retention efforts. EDM was used to assess the efficacy of a writing center in an effort to analyze student achievement and student progress to the next grade (Yeats, R., Reddy, P. J., Wheeler, A., Senior, C., & Murray, J., 2010).Their work demonstrated the ability to assess a specific educational support process, i.e., the writing center, in an effort to improve institutional effectiveness. Their research approach used a combination of quantitative work and case study analysis. The mixed methods approach to data mining was helpful in understanding much more about the ways data mining can be used in an actual implementation. Their research results were not surprising in that it found students who attend writing centers tend to do better in their classes. This research took a different approach to analyze student achievement in that it made the connection between writing center attendance and student grades. It did not mention student retention issues, but new studies and researches could examine the relationship between these three concepts: writing center attendance, student grades, and retention. Arizona State University professors used three different data mining tools ( Classification trees, multivariate adaptive regression splines and neural networks ) to determine predictors of student retention (Yu, C. H., DiGangi, S., Jannasch-Pennell, A., & Kaprolet, C., 2010).Data mining procedures identified transferred hours, residency, and ethnicity as crucial factors to retention. Through this research, they also discovered that east coast students tend to stay enrolled longer than their west coast counterparts do. Another research team used data mining to classify students into three groups as early as they could in the academic year (Vandamme, J. P., Meskens, N., & Superby, J. F., 2007). This research is a good example for predicting academic performance and student success by using data mining tools. The three groups included low risk, medium risk, and high risk students. The researchers used several data mining techniques. The student in the high risk group had a high probability of failing or dropping out of school. These types of studies are important in that they give faculty and staff a way to identify the at-risk students in a proactive way, because once a student decides to leave, it is hard to convince them to stay. In a related study, researchers examined whether the demographic background of students had any influence on their performance (Yorke, M., Barnett, G., Evanson, P., Haines, C., Jenkins, D., Knight, P., Woolf, H., 2005). Results from the study appeared inconclusive, potentially because of the type of analysis they did. Interestingly, the field of EDM is concerned with analytical methods, and not necessarily just data mining methods. In this study researchers used Microsoft Excel for their analysis and mining data. The problem with this approach is that they discuss mining the data without really applying data mining tools. It is clear that researchers

should exercise more caution when using the phrase data mining, especially when they are not referring to data mining tools. The drawback with the research used these phrases, but never applied them. This particular research demonstrates that researchers can still conduct data analyses by using Excel, but researchers should not mislead the reader when describing their approach. Contrary the Yorkes study, a different research team noted that demographic characteristics are not significant predictors of student satisfaction or success (Thomas, E. H., & Galambos, N., 2004).The results seem to report different findings related to student satisfaction or the prediction of student success. One can conclude there are significantly more factors that influence students success than what has been studied thus far. REFERENCES
Baker, R., & Yacef, K. (2009). The State of Educational Data mining in 2009: A Review and Future Visions. Journal of Educational Data Mining. Chacon, F., Spicer, D., & Valbuena, A. (2012, April 10). Analytics in Support of Student Retention and Success. Research Bulletin, pp. 1-9. Kiron, D., Shockley, R., Kruschwitz, N., Finch, G., & Haydock, M. (2012). Analytics: The Widening Divide. MIT Sloan Management Review, 1-22. Lenning, O. T. (1980). Retention and Attrition: Evidence for Action and Research. National Center for Higher Education Management Systems, 134. Lin, S. (2012). Data mining for student retention management. J. Comput. Sci. Coll., 92-99. Luan, J. (2002). Data Mining and Knowledge Management in Higher Education : Potential Applications. Toronto, Ontario, Canada: Paper presented at the Annual Forum for the Association,for Institutional Research. Thomas, E. H., & Galambos, N. (2004). What satisfies Students?: Mining Student-Opinion Data with Regression and Decision Tree Analysis. Research in Higher Education, 251-269. Vandamme, J. P., Meskens, N., & Superby, J. F. (2007). Predicting Academic Performance by Data Mining Methods. In Education Economics (pp. 405-419). Yeats, R., Reddy, P. J., Wheeler, A., Senior, C., & Murray, J. (2010). What a difference a writing centre makes: a small scale study. Education & Training, 499-507. Yorke, M., Barnett, G., Evanson, P., Haines, C., Jenkins, D., Knight, P., Woolf, H. (2005). Mining institutional datasets to support policy making and implementation. Journal of Higher Education Policy and Management, 285-298. Yu, C. H., DiGangi, S., Jannasch-Pennell, A., & Kaprolet, C. (2010). A Data Mining Approach for Identifying Predictors of Student Retention from Sophomore to Junior Year. Journal of Data Science , 307325.

You might also like