FLEX-TALENT: AN EXPLAINABLE AND FAIR MACHINE LEARNING FRAMEWORK FOR HIGH-STAKES LEADERSHIP-PIPELINE EVALUATION

Authors

  • ZiMeng Wang (Corresponding Author) Brandeis University, Waltham 02453, MA, USA.
  • ZiJian Shen Carnegie Mellon University, Pittsburgh, PA 15213, USA

Keywords:

Algorithmic HRM, Executive talent evaluation, Explainable AI, Fairness, Promotion prediction, CatBoost, SHAP, Responsible AI

Abstract

Machine learning can prioritize candidates for leadership pipelines, but historical promotion records encode organizational policy rather than objective executive potential. This distinction is critical in high-stakes talent evaluation, where a technically accurate model may reproduce label bias, obscure unequal error burdens, or invite automation bias. This paper presents FLEX-Talent, a construct-aware, budget-constrained framework that combines feature governance, training-only fairness reweighting, predeclared validation gates, group calibration audits, intersectional testing, and explanation-stability checks. The framework is evaluated on a real public HR promotion dataset containing 54,808 employees and 4,668 historical promotions. Gender is never used at inference; it is retained only for auditing and reweighting. Five candidate models are compared under a fixed 10% shortlist budget. The validation rule selected a reweighted CatBoost model, which achieved 0.9134 ROC-AUC and 0.6116 average precision on an untouched test set. However, its absolute gender equal-opportunity difference increased from 0.0138 on validation to 0.0578 on test, slightly exceeding the predeclared 0.05 gate; a 95% bootstrap interval of [0.0060, 0.1497] reveals substantial fairness uncertainty. Rather than treating this as a hidden failure, FLEX-Talent makes fairness generalization a first-class release decision. Global and local SHAP evidence is accompanied by bootstrap stability analysis and explicit non-causal warnings. The resulting contribution is not an automated promotion system, but a reproducible governance architecture for human review, appeal, and post-deployment monitoring.

References

[1] Tambe P, Cappelli P, Yakubovich V. Artificial intelligence in human resources management: Challenges and a path forward. California Management Review, 2019, 61(4): 15-42. DOI: 10.1177/0008125619867910.

[2] Leicht-Deobald U. The challenges of algorithm-based HR decision-making for personal integrity. Journal of Business Ethics, 2019, 160(2): 377-392. DOI: 10.1007/s10551-019-04204-w.

[3] Köchling A, Wehner M C. Discriminated by an algorithm: A systematic review of discrimination and fairness by algorithmic decision-making in the context of HR recruitment and HR development. Business Research, 2020, 13: 795-848. DOI: 10.1007/s40685-020-00134-w.

[4] Raghavan M, Barocas S, Kleinberg J, et al. Mitigating bias in algorithmic hiring: Evaluating claims and practices. In: Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAT*), 2020: 469-481. DOI: 10.1145/3351095.3372828.

[5] Langer M, Landers R N. The future of artificial intelligence at work: A review on effects of decision automation and augmentation on workers targeted by algorithms and third-party observers. Computers in Human Behavior, 2021, 123: Art. 106878. DOI: 10.1016/j.chb.2021.106878.

[6] Rodgers W, Murray J M, Stefanidis A, et al. An artificial intelligence algorithmic approach to ethical decision-making in human resource management processes. Human Resource Management Review, 2023, 33(1): Art. 100925. DOI: 10.1016/j.hrmr.2022.100925.

[7] Hardt M, Price E, Srebro N. Equality of opportunity in supervised learning. In: Advances in Neural Information Processing Systems 29, 2016.

[8] Chouldechova A. Fair prediction with disparate impact: A study of bias in recidivism prediction instruments. Big Data, 2017, 5(2): 153-163. DOI: 10.1089/big.2016.0047.

[9] Kleinberg J, Mullainathan S, Raghavan M. Inherent trade-offs in the fair determination of risk scores. In: Proceedings of the Innovations in Theoretical Computer Science Conference (ITCS), LIPIcs, 2017, 67: Art. 43. DOI: 10.4230/LIPIcs.ITCS.2017.43.

[10] Feldman M. Certifying and removing disparate impact. In: Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), 2015: 259-268. DOI: 10.1145/2783258.2783311.

[11] Kamiran F, Calders T. Data preprocessing techniques for classification without discrimination. Knowledge and Information Systems, 2012, 33: 1-33. DOI: 10.1007/s10115-011-0463-8.

[12] Agarwal A, Beygelzimer A, Dudík M, et al. A reductions approach to fair classification. In: Proceedings of the International Conference on Machine Learning (ICML), PMLR, 2018, 80: 60-69.

[13] Pleiss G, Raghavan M, Wu F, et al. On fairness and calibration. In: Advances in Neural Information Processing Systems 30, 2017: 5684-5693.

[14] Binns R. It’s reducing a human being to a percentage: Perceptions of justice in algorithmic decisions. In: Proceedings of the ACM Conference on Human Factors in Computing Systems (CHI), 2018: Art. 377. DOI: 10.1145/3173574.3173951.

[15] Mitchell M. Model cards for model reporting. In: Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAT*), 2019: 220-229. DOI: 10.1145/3287560.3287596.

[16] Gebru T. Datasheets for datasets. Communications of the ACM, 2021, 64(12): 86-92. DOI: 10.1145/3458723.

[17] Tabassi E. Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1. National Institute of Standards and Technology, 2023. DOI: 10.6028/NIST.AI.100-1.

[18] Suresh H, Guttag J. A framework for understanding sources of harm throughout the machine learning life cycle. In: Proceedings of the ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization (EAAMO), 2021: Art. 17. DOI: 10.1145/3465416.3483305.

[19] Selbst A D. Fairness and abstraction in sociotechnical systems. In: Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAT*), 2019: 59-68. DOI: 10.1145/3287560.3287598.

[20] Jacobs A Z, Wallach H. Measurement and fairness. In: Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT), 2021: 375-385. DOI: 10.1145/3442188.3445901.

[21] Bellamy R K E. AI Fairness 360: An extensible toolkit for detecting and mitigating algorithmic bias. IBM Journal of Research and Development, 2019, 63(4/5): 4:1-4:15. DOI: 10.1147/JRD.2019.2942287.

[22] Ribeiro M T, Singh S, Guestrin C. Why should I trust you?: Explaining the predictions of any classifier. In: Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), 2016: 1135-1144. DOI: 10.1145/2939672.2939778.

[23] Lundberg S M, Lee S-I. A unified approach to interpreting model predictions. In: Advances in Neural Information Processing Systems 30, 2017: 4765-4774.

[24] Lundberg S M. From local explanations to global understanding with explainable AI for trees. Nature Machine Intelligence, 2020, 2: 56-67. DOI: 10.1038/s42256-019-0138-9.

[25] Rudin C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence, 2019, 1: 206-215. DOI: 10.1038/s42256-019-0048-x.

[26] Guidotti R,. A survey of methods for explaining black box models. ACM Computing Surveys, 2018, 51(5): Art. 93. DOI: 10.1145/3236009.

[27] Wachter S, Mittelstadt B, Russell C. Counterfactual explanations without opening the black box: Automated decisions and the GDPR. Harvard Journal of Law & Technology, 2018, 31(2): 841-887.

[28] Slack D. Fooling LIME and SHAP: Adversarial attacks on post hoc explanation methods. In: Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (AIES), 2020: 180-186. DOI: 10.1145/3375627.3375830.

[29] Adebayo J. Sanity checks for saliency maps. In: Advances in Neural Information Processing Systems 31, 2018: 9505-9515.

[30] Miller T. Explanation in artificial intelligence: Insights from the social sciences. Artificial Intelligence, 2019, 267: 1-38. DOI: 10.1016/j.artint.2018.07.007.

[31] He H, Garcia E A. Learning from imbalanced data. IEEE Transactions on Knowledge and Data Engineering, 2009, 21(9): 1263-1284. DOI: 10.1109/TKDE.2008.239.

[32] Saito T, Rehmsmeier M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS ONE, 2015, 10(3): Art. e0118432. DOI: 10.1371/journal.pone.0118432.

[33] Prokhorenkova L. CatBoost: Unbiased boosting with categorical features. In: Advances in Neural Information Processing Systems 31, 2018: 6638-6648.

Downloads

Published

2024-01-01

Issue

Section

Research Article

DOI:

How to Cite

ZiMeng Wang, ZiJian Shen. Flex-Talent: An Explainable And Fair Machine Learning Framework For High-Stakes Leadership-Pipeline Evaluation. World Journal of Management Science. 2024, 2(1): 73-82. DOI: https://doi.org/10.61784/wms4121.