TRUE-EWS: TRUSTWORTHY EARLY WARNING WITH EQUALIZED CONFORMAL GUARANTEES FOR EQUITABLE STUDENT SUPPORT IN U.S. EDUCATION

Authors

  • QianYi Fang (Corresponding Author) University of Chichester, Chichester PO19 6PE, United Kingdom.

Keywords:

Trustworthy AI, Educational equity, Early warning systems, Conformal prediction, Algorithmic fairness, Probability calibration, Explainable AI, Human-in-the-loop decision making

Abstract

Early-warning systems are now routine in U.S. schools, colleges, and professional programs, yet a risk score alone does not tell a counselor how often at-risk students from each demographic group will be overlooked, or how often students who are on track will be flagged. This paper presents TRUE-EWS, a Transparent, Reliable, Uncertainty-aware, and Equitable early-warning framework. TRUE-EWS combines a gradient-boosted risk model, a strictly increasing group-wise Beta calibration map, an equalized conformal decision layer (ECC) that runs a Mondrian conformal predictor over group-by-outcome cells, and a TreeSHAP decomposition of between-group risk gaps. ECC maps each student to one of three actions (flag for support, no action, or referral to a counselor) and guarantees, in finite samples and without distributional assumptions, that the miss rate of at-risk students and the false-flag rate of students on track are bounded within every group. We also prove that strictly increasing recalibration leaves ECC decisions unchanged, so probability reliability and decision guarantees can be managed separately. Experiments with 30 stratified random splits on three U.S. cohorts that span the education pipeline, namely Tennessee STAR (early grades), High School and Beyond (college access), and the LSAC National Longitudinal Bar Passage Study (professional licensure), show that standard conformal prediction meets its 90% target overall while covering only 3.0% of at-risk White law students, and that label-conditional prediction flags 37.6% of Non-White students who go on to pass the bar. ECC holds every group-by-outcome coverage between 0.898 and 0.909, lowers the worst-group false-flag rate on LSAC from 0.371 to 0.098 relative to confidence-based deferral at the same counselor workload, and extends directly to intersectional groups.

References

[1] Bowers A J, Sprott R, Taff S A. Do we know who will drop out? A review of the predictors of dropping out of high school: Precision, sensitivity, and specificity. The High School Journal, 2013, 96(2): 77-100.

[2] Lakkaraju H, Aguiar E, Shan C, et al. A machine learning framework to identify students at risk of adverse academic outcomes. Proceedings of the 21st ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2015: 1909-1918.

[3] Yu R, Lee H, Kizilcec R F. Should college dropout prediction models include protected attributes? Proceedings of the 8th ACM Conference on Learning @ Scale, 2021: 91-100.

[4] Baker R S, Hawn A. Algorithmic bias in education. International Journal of Artificial Intelligence in Education, 2022, 32(4): 1052-1092.

[5] Kizilcec R F, Lee H. Algorithmic fairness in education. In: Holmes W, Porayska-Pomsta K, eds. The Ethics of Artificial Intelligence in Education. Routledge, 2022: 174-202.

[6] Jobin A, Ienca M, Vayena E. The global landscape of AI ethics guidelines. Nature Machine Intelligence, 2019, 1(9): 389-399.

[7] Holmes W, Porayska-Pomsta K, Holstein K, et al. Ethics of AI in education: Towards a community-wide framework. International Journal of Artificial Intelligence in Education, 2022, 32(3): 504-526.

[8] Gardner J, Brooks C, Baker R. Evaluating the fairness of predictive student models through slicing analysis. Proceedings of the 9th International Conference on Learning Analytics & Knowledge, 2019: 225-234.

[9] Jiang W, Pardos Z A. Towards equity and algorithmic fairness in student grade prediction. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 2021: 608-617.

[10] Kamiran F, Calders T. Data preprocessing techniques for classification without discrimination. Knowledge and Information Systems, 2012, 33(1): 1-33.

[11] Agarwal A, Beygelzimer A, Dudík M, et al. A reductions approach to fair classification. Proceedings of the 35th International Conference on Machine Learning, PMLR, 2018, 80: 60-69.

[12] Hardt M, Price E, Srebro N. Equality of opportunity in supervised learning. Advances in Neural Information Processing Systems, 2016, 29: 3315-3323.

[13] Kleinberg J, Mullainathan S, Raghavan M. Inherent trade-offs in the fair determination of risk scores. Proceedings of the 8th Innovations in Theoretical Computer Science Conference, 2017: 43:1-43:23.

[14] Pleiss G, Raghavan M, Wu F, et al. On fairness and calibration. Advances in Neural Information Processing Systems, 2017, 30: 5680-5689.

[15] Vovk V, Gammerman A, Shafer G. Algorithmic Learning in a Random World. New York: Springer, 2005.

[16] Vovk V. Conditional validity of inductive conformal predictors. Proceedings of the Asian Conference on Machine Learning, PMLR, 2012, 25: 475-490.

[17] Romano Y, Barber R F, Sabatti C, et al. With malice toward none: Assessing uncertainty via equalized coverage. Harvard Data Science Review, 2020, 2(2).

[18] Madras D, Pitassi T, Zemel R. Predict responsibly: Improving fairness and accuracy by learning to defer. Advances in Neural Information Processing Systems, 2018, 31.

[19] Jones E, Sagawa S, Koh P W, et al. Selective classification can magnify disparities across groups. Proceedings of the International Conference on Learning Representations, 2021.

[20] Guo C, Pleiss G, Sun Y, et al. On calibration of modern neural networks. Proceedings of the 34th International Conference on Machine Learning, PMLR, 2017, 70: 1321-1330.

[21] Zadrozny B, Elkan C. Transforming classifier scores into accurate multiclass probability estimates. Proceedings of the 8th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2002: 694-699.

[22] Kull M, Silva Filho T, Flach P. Beta calibration: A well-founded and easily implemented improvement on logistic calibration for binary classifiers. Proceedings of the 20th International Conference on Artificial Intelligence and Statistics, PMLR, 2017, 54: 623-631.

[23] Sadinle M, Lei J, Wasserman L. Least ambiguous set-valued classifiers with bounded error levels. Journal of the American Statistical Association, 2019, 114(525): 223-234.

[24] Lundberg S M, Erion G, Chen H, et al. From local explanations to global understanding with explainable AI for trees. Nature Machine Intelligence, 2020, 2(1): 56-67.

[25] Ke G, Meng Q, Finley T, et al. LightGBM: A highly efficient gradient boosting decision tree. Advances in Neural Information Processing Systems, 2017, 30: 3146-3154.

[26] Krueger A B. Experimental estimates of education production functions. Quarterly Journal of Economics, 1999, 114(2): 497-532.

[27] Kleiber C, Zeileis A. Applied Econometrics with R. New York: Springer, 2008.

[28] Rouse C E. Democratization or diversion? The effect of community colleges on educational attainment. Journal of Business & Economic Statistics, 1995, 13(2): 217-224.

[29] Wightman L F. LSAC National Longitudinal Bar Passage Study. Newtown: Law School Admission Council, LSAC Research Report Series, 1998.

[30] Le Quy T, Roy A, Iosifidis V, et al. A survey on datasets for fairness-aware machine learning. WIREs Data Mining and Knowledge Discovery, 2022, 12(3): e1452.

Downloads

Published

2024-12-11

Issue

Section

Research Article

DOI:

How to Cite

QianYi Fang. True-Ews: Trustworthy Early Warning With Equalized Conformal Guarantees For Equitable Student Support In U.s. Education. Educational Research and Human Development. 2024, 1(2): 44-54. DOI: https://doi.org/10.61784/erhd4059.