BIAS-AWARE LARGE LANGUAGE MODELS FOR EQUITABLE EDUCATIONAL FEEDBACK AND ASSESSMENT

Authors

  • QianYi Fang (Corresponding Author) University of Chichester, Chichester PO19 6PE, United Kingdom.

Keywords:

Automated essay scoring, Large language models, Algorithmic fairness, Counterfactual fairness, Dialect bias, English learners, Formative feedback

Abstract

Pre-trained large language models (LLMs) are increasingly used to score student writing and to produce trait-level feedback, yet their sensitivity to construct-irrelevant linguistic variation can disadvantage students who write in African American English (AAE) or who are learning English as a second language (L2). This paper presents a bias-aware scoring framework that audits and mitigates such variation without access to student demographics. Counterfactual versions of real student essays are generated by an AAE morphosyntactic rule system and by an L2 error generator whose error-category rates are estimated from the JFLEG learner corpus and whose misspellings are drawn from TOEFL-Spell. Paired representations from a frozen RoBERTa encoder define an invariance operator, and Construct-Aligned Invariance Regularization (CAIR) applies it only to score dimensions that the rubric defines as independent of language conventions, which yields a closed-form estimator. On 5,569 ASAP essays from five prompts and five random splits, CAIR reduces the mean standardized counterfactual shift of holistic content scores by 85% and the counterfactual score-flip rate by about 40%, with no significant change in quadratic weighted kappa (0.732 vs. 0.733). For trait feedback, CAIR removes the spill-over of learner errors into Ideas, Organization and Style scores while preserving the legitimate sensitivity of the Conventions score, which task-agnostic debiasing methods erase. Pairs for 2% of the training essays already halve the shift, and invariance learned from one variety transfers to the other.

References

[1] Shermis M D. State-of-the-art automated essay scoring: Competition, results, and future directions from a United States demonstration. Assessing Writing, 2014, 20: 53-76.

[2] Taghipour K, Ng H T. A neural approach to automated essay scoring. In: Proceedings of Conference Empirical Methods in Natural Language Processing (EMNLP), 2016: 1882-1891.

[3] Dong F, Zhang Y, Yang J. Attention-based recurrent convolutional neural network for automatic essay scoring. In: Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL), 2017: 153-162.

[4] Devlin J, Chang M-W, Lee K, et al. BERT: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of Conference North American Chapter of the ACL: Human Language Technologies (NAACL-HLT), 2019: 4171-4186.

[5] Liu Y, Ott M, Goyal N, et al. RoBERTa: A robustly optimized BERT pretraining approach. arXiv:1907.11692, 2019.

[6] Brown T, Mann B, Ryder N, et al. Language models are few-shot learners. In: Advances in Neural Information Processing Systems, 2020, 33: 1877-1901.

[7] Mayfield E, Black A W. Should you fine-tune BERT for automated essay scoring? In: Proceedings of the 15th Workshop on Innovative Use of NLP for Building Educational Applications, 2020: 151-162.

[8] Yang R, Cao J, Wen Z, et al. Enhancing automated essay scoring performance via fine-tuning pre-trained language models with combination of regression and ranking. In: Findings of the Association for Computational Linguistics: EMNLP 2020, 2020: 1560-1569.

[9] Kumar R, Mathias S, Saha S, et al. Many hands make light work: Using essay traits to automatically score essays. In: Proceedings of NAACL-HLT, 2022: 1485-1495.

[10] Williamson D M, Xi X, Breyer F J. A framework for evaluation and use of automated scoring. Educational Measurement: Issues and Practice, 2012, 31(1): 2-13.

[11] Bridgeman B, Trapani C, Attali Y. Comparison of human and machine scoring of essays: Differences by gender, ethnicity, and country. Applied Measurement in Education, 2012, 25(1): 27-40.

[12] Madnani N, Loukina A, von Davier A, et al. Building better open-source tools to support fairness in automated scoring. In: Proceedings of First ACL Workshop on Ethics in Natural Language Processing, 2017: 41-52.

[13] Loukina A, Madnani N, Zechner K. The many dimensions of algorithmic fairness in educational applications. In: Proceedings of the 14th Workshop on Innovative Use of NLP for Building Educational Applications, 2019: 1-10.

[14] Mayfield E, Madaio M, Prabhumoye S, et al. Equity beyond bias in language technologies for education. In: Proceedings of the 14th Workshop on Innovative Use of NLP for Building Educational Applications, 2019: 444-460.

[15] Baker R S, Hawn A. Algorithmic bias in education. International Journal of Artificial Intelligence in Education, 2022, 32(4): 1052-1092.

[16] Kwako A, Wan Y, Zhao J, et al. Using item response theory to measure gender and racial bias of a BERT-based automated English speech assessment system. In: Proceedings of the 17th Workshop on Innovative Use of NLP for Building Educational Applications, 2022: 1-7.

[17] Blodgett S L, Barocas S, Daumé III H, et al. Language (technology) is power: A critical survey of ‘bias’ in NLP. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), 2020: 5454-5476.

[18] Blodgett S L, Green L, O’Connor B. Demographic dialectal variation in social media: A case study of African-American English. In: Proceedings of EMNLP, 2016: 1119-1130.

[19] Ziems C, Chen J, Harris C, et al. VALUE: Understanding dialect disparity in NLU. In: Proceedings of the 60th Annual Meeting of the ACL (Volume 1: Long Papers), 2022: 3701-3720.

[20] Ding Y, Riordan B, Horbach A, et al. Don’t take ‘nswvtnvakgxpm’ for an answer – The surprising vulnerability of automatic content scoring systems to adversarial input. In: Proceedings of the 28th International Conference Computational Linguistics (COLING), 2020: 882-892.

[21] Napoles C, Sakaguchi K, Tetreault J. JFLEG: A fluency corpus and benchmark for grammatical error correction. In: Proceedings of the 15th Conference European Chapter of the ACL (EACL), 2017, 2: 229-234.

[22] Flor M, Fried M, Rozovskaya A. A benchmark corpus of English misspellings and a minimally-supervised model for spelling correction. In: Proceedings of the 14th Workshop on Innovative Use of NLP for Building Educational Applications, 2019: 76-86.

[23] Felice M, Yuan Z. Generating artificial errors for grammatical error correction. In: Proceedings of Student Research Workshop at the 14th Conference European Chapter of the ACL, 2014: 116-126.

[24] Kusner M J, Loftus J, Russell C, et al. Counterfactual fairness. In: Advances in Neural Information Processing Systems, 2017, 30.

[25] Garg S, Perot V, Limtiaco N, et al. Counterfactual fairness in text classification through robustness. In: Proceedings of AAAI/ACM Conference AI, Ethics, and Society (AIES), 2019: 219-226.

[26] Zmigrod R, Mielke S J, Wallach H, et al. Counterfactual data augmentation for mitigating gender stereotypes in languages with rich morphology. In: Proceedings of the 57th Annual Meeting of the ACL, 2019: 1651-1661.

[27] Elazar Y, Goldberg Y. Adversarial removal of demographic attributes from text data. In: Proceedings of EMNLP, 2018: 11-21.

[28] Ravfogel S, Elazar Y, Gonen H, et al. Null it out: Guarding protected attributes by iterative nullspace projection. In: Proceedings of the 58th Annual Meeting of the ACL, 2020: 7237-7256.

[29] Liang P P, Li I M, Zheng E, et al. Towards debiasing sentence representations. In: Proceedings of the 58th Annual Meeting of the ACL, 2020: 5502-5515.

[30] Liu N F, Gardner M, Belinkov Y, et al. Linguistic knowledge and transferability of contextual representations. In: Proceedings of NAACL-HLT, 2019: 1073-1094.

Downloads

Published

2024-02-02

How to Cite

QianYi Fang. Bias-Aware Large Language Models For Equitable Educational Feedback And Assessment. Trends in Social Sciences and Humanities Research. 2024, 2(1): 41-50. DOI: https://doi.org/10.61784/tsshr4253.