XOR-EW: AN EXPLAINABLE AND COST-AWARE ARTIFICIAL INTELLIGENCE FRAMEWORK FOR OPERATIONAL RISK EARLY WARNING IN FINANCIAL INSTITUTIONS
Keywords:
Operational risk, Early warning, Explainable artificial intelligence, Detection, Probability calibration, Cost-sensitive learning, Selective prediction, SHAPAbstract
Operational-risk early-warning models in financial institutions must do more than rank rare events: they must remain stable under temporal change, express meaningful probabilities, respect finite investigation capacity, and provide auditable reasons. This paper presents XOR-EW (eXplainable Operational-Risk Early Warning), a decision-oriented framework that treats transaction as an observable proxy for a class of process- and control-related operational loss events. XOR-EW combines strict chronological controls, stability-aware champion selection, post-hoc probability calibration, an exposure-sensitive threshold under an alert-capacity constraint, ensemble-disagreement triage, and SHAP-based reason codes with explanation-stability checks. Experiments were executed on a publicly accessible mirror of the ULB/Worldline credit-card data. After exact deduplication, 227,102 transactions were split chronologically into training, calibration, policy, and test windows. A balanced random forest was selected without consulting the test set because it achieved the highest temporal-robustness score across the calibration and policy windows. On the 45,421-transaction test window, calibration reduced the Brier score from 0.000511 to 0.000421 and ten-bin expected calibration error from 0.003736 to 0.000304. Relative to the raw champion with an F1-selected threshold, the complete policy increased recall from 0.561 to 0.737, increased F1 from 0.719 to 0.785, and reduced normalized exposure cost from 0.731 to 0.358. An additional eight disagreement cases raised total capture to 0.772 while total workload remained 0.128% of transactions. The results support XOR-EW as a reproducible governance architecture, not as a claim of universal superiority or production readiness.References
[1] Basel Committee on Banking Supervision. Principles for the Sound Management of Operational Risk. Basel, Switzerland: Bank for International Settlements, 2011.
[2] Basel Committee on Banking Supervision. Revisions to the Principles for the Sound Management of Operational Risk. Basel, Switzerland: Bank for International Settlements, 2021.
[3] Board of Governors of the Federal Reserve System, Office of the Comptroller of the Currency. Supervisory Guidance on Model Risk Management, SR Letter 11-7 / OCC Bulletin 2011-12, 2011.
[4] Tabassi E. Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. Gaithersburg, MD, USA: National Institute of Standards and Technology, 2023.
[5] Phillips P J, Hahn C, Fontana P, et al. Four Principles of Explainable Artificial Intelligence, NISTIR 8312. National Institute of Standards and Technology, 2021.
[6] Dal Pozzolo A, Caelen O, Johnson R A, et al. Calibrating probability with undersampling for unbalanced classification. In: 2015 IEEE Symposium Series on Computational Intelligence. Cape Town, South Africa, 2015: 159-166.
[7] Carcillo F, Dal Pozzolo A, Le Borgne Y A, et al. SCARFF: A scalable framework for streaming credit card detection with Spark. Information Fusion, 2018, 41: 182-194.
[8] Carcillo F, Le Borgne Y A, Caelen O, et al. Combining unsupervised and supervised learning in credit card detection. Information Sciences, 2021, 557: 317-331.
[9] Bahnsen A C, Aouada D, Stojanovic A, et al. Feature engineering strategies for credit card detection. Expert Systems with Applications, 2016, 51: 134-142.
[10] Bahnsen A C, Aouada D, Ottersten B. Example-dependent cost-sensitive decision trees. Expert Systems with Applications, 2015, 42(19): 6609-6619.
[11] Sahin Y, Bulkan S, Duman E. A cost-sensitive decision tree approach for detection. Expert Systems with Applications, 2013, 40(15): 5916-5923.
[12] Breiman L. Random forests. Machine Learning, 2001, 45: 5-32.
[13] Chen T, Guestrin C. XGBoost: A scalable tree boosting system. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. San Francisco, CA, USA, 2016: 785-794.
[14] Ke G, Meng Q, Finley T, et al. LightGBM: A highly efficient gradient boosting decision tree. In: Advances in Neural Information Processing Systems, 2017: 3146-3154.
[15] He H, Garcia E A. Learning from imbalanced data. IEEE Transactions on Knowledge and Data Engineering, 2009, 21(9): 1263-1284.
[16] Chawla N V, Bowyer K W, Hall L O, et al. SMOTE: Synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 2002, 16: 321-357.
[17] Saito T, Rehmsmeier M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLOS ONE, 2015, 10(3): e0118432.
[18] Davis J, Goadrich M. The relationship between precision-recall and ROC curves. In: Proceedings of the 23rd International Conference on Machine Learning. Pittsburgh, PA, USA, 2006: 233-240.
[19] Fawcett T. An introduction to ROC analysis. Pattern Recognition Letters, 2006, 27(8): 861-874.
[20] Niculescu-Mizil A, Caruana R. Predicting good probabilities with supervised learning. In: Proceedings of the 22nd International Conference on Machine Learning. Bonn, Germany, 2005: 625-632.
[21] Guo C, Pleiss G, Sun Y, et al. On calibration of modern neural networks. In: Proceedings of the 34th International Conference on Machine Learning, 2017: 1321-1330.
[22] Platt J C. Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. In: Smola A J, Bartlett P, Schölkopf B, et al., eds. Advances in Large Margin Classifiers. Cambridge, MA, USA: MIT Press, 1999: 61-74.
[23] Lundberg S M, Lee S I. A unified approach to interpreting model predictions. In: Advances in Neural Information Processing Systems 30, 2017: 4765-4774.
[24] Lundberg S M, Erion G, Chen H, et al. From local explanations to global understanding with explainable AI for trees. Nature Machine Intelligence, 2020, 2: 56-67.
[25] Ribeiro M T, Singh S, Guestrin C. “Why should I trust you?”: Explaining the predictions of any classifier. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. San Francisco, CA, USA, 2016: 1135-1144.
[26] Barredo Arrieta A, Díaz-Rodríguez N, Del Ser J, et al. Explainable artificial intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Information Fusion, 2020, 58: 82-115.
[27] Rudin C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence, 2019, 1: 206-215.
[28] Webb G I, Hyde R, Cao H, et al. Characterizing concept drift. Data Mining and Knowledge Discovery, 2016, 30(4): 964-994.
[29] Alvarez-Melis D, Jaakkola T S. On the robustness of interpretability methods. arXiv:1806.08049, 2018.
[30] Elkan C. The foundations of cost-sensitive learning. In: Proceedings of the 17th International Joint Conference on Artificial Intelligence. Seattle, WA, USA, 2001: 973-978.