BEYOND MODEL SCALE: A TASK-AWARE RELIABILITY BENCHMARK FOR LARGE LANGUAGE MODELS IN CLINICAL DECISION SUPPORT

Authors

  • JiTong Zou Case Western Reserve University, Cleveland, Ohio, United States.
  • YuXuan Qin (Corresponding Author) Northeastern University, Boston, MA 02115, United States.
  • MinJae Rhee University of Illinois Urbana-Champaign, Champaign, IL, United States.

Keywords:

Clinical decision support, Large language models, Reliability benchmarking, Hallucination, Task-aware evaluation, Model scale

Abstract

Model size is often treated as a proxy for trustworthiness when large language models (LLMs) are considered for clinical decision support, yet a single aggregate score can hide task-specific failure modes. This paper introduces TARB-CDS, a task-aware reliability audit that transforms a model-by-task accuracy matrix into deployment-relevant reliability profiles. The framework combines macro accuracy with lower-tail conditional value at risk over the two weakest tasks (CVaR-2), worst-task accuracy, Pareto dominance, policy-weighted scores, contextual rank stability under uncertain task mixtures, output-format reliability, and matched base-versus-instruction-tuned comparisons. We conduct a reproducible secondary analysis of 12 LLM checkpoints evaluated on seven Med-HALT reasoning and biomedical-memory tasks. The public reasoning files contain 39,590 rows across False Confidence, Fake Question, and None-of-the-Above tests; published model accuracies and format-error rates are analyzed without claiming fresh model inference. Falcon-40B has the highest seven-task macro accuracy (42.67%) and CVaR-2 (9.36%), but the best worst-task accuracy among all models is only 1.38%. Under 100,000 uniformly sampled task mixtures, Falcon-40B ranks first in 77.60% of contexts, leaving substantial rank instability. Among open checkpoints, parameter count has only a modest descriptive association with macro accuracy (Spearman ρ=0.366), while all five matched instruction/chat-tuning pairs reduce macro accuracy by an average of 9.30 percentage points. These results show that scale alone is insufficient for reliability claims and motivate task-conditional model cards and minimum-performance gates before clinical evaluation.

References

[1] Singhal K. Large language models encode clinical knowledge. Nature, 2023, 620: 172-180. DOI: 10.1038/s41586-023-06291-2.

[2] Pal A, Umapathi L K, Sankarasubbu M. MedMCQA: A large-scale multi-subject multi-choice dataset for medical domain question answering. In: Proceedings of the Conference on Health, Inference, and Learning, PMLR, 2022, 174: 248-260.

[3] Jin Q, Dhingra B, Liu Z, et al. PubMedQA: A dataset for biomedical research question answering. In: Proceedings of the Conference on Empirical Methods in Natural Language Processing and the International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 2019: 2567-2577. DOI: 10.18653/v1/D19-1259.

[4] Vilares D, Gómez-Rodríguez C. HEAD-QA: A healthcare dataset for complex reasoning. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019: 960-966. DOI: 10.18653/v1/P19-1092.

[5] Jin D. What disease does this patient have? A large-scale open domain question answering dataset from medical exams. Applied Sciences, 2021, 11(14): 6421. DOI: 10.3390/app11146421.

[6] Kung T H. Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models. PLOS Digital Health, 2023, 2(2): e0000198. DOI: 10.1371/journal.pdig.0000198.

[7] Gilson A. How does ChatGPT perform on the United States Medical Licensing Examination? The implications of large language models for medical education and knowledge assessment. JMIR Medical Education, 2023, 9: e45312. DOI: 10.2196/45312.

[8] Nori H, King N, McKinney S M, et al. Capabilities of GPT-4 on medical challenge problems. arXiv:2303.13375, 2023.

[9] Hendrycks D. Measuring massive multitask language understanding. In: Proceedings of the International Conference on Learning Representations (ICLR), 2021.

[10] He Z. MedEval: A multi-level, multi-task, and multi-domain medical benchmark for language model evaluation. In: Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2023: 8725-8744. DOI: 10.18653/v1/2023.emnlp-main.540.

[11] Li J, Cheng X, Zhao X, et al. HaluEval: A large-scale hallucination evaluation benchmark for large language models. In: Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2023: 6449-6464. DOI: 10.18653/v1/2023.emnlp-main.397.

[12] Ribeiro M T, Wu T, Guestrin C, et al. Beyond accuracy: Behavioral testing of NLP models with CheckList. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020: 4902-4912. DOI: 10.18653/v1/2020.acl-main.442.

[13] Liang P. Holistic evaluation of language models. arXiv:2211.09110, 2022.

[14] Guo C, Pleiss G, Sun Y, et al. On calibration of modern neural networks. In: Proceedings of the International Conference on Machine Learning (ICML), PMLR, 2017, 70: 1321-1330.

[15] Ovadia Y. Can you trust your model’s uncertainty? Evaluating predictive uncertainty under dataset shift. In: Advances in Neural Information Processing Systems, 2019, 32.

[16] Jiang Z, Araki J, Ding H, et al. How can we know when language models know? On the calibration of language models for question answering. Transactions of the Association for Computational Linguistics, 2021, 9: 962-977. DOI: 10.1162/tacl_a_00407.

[17] Desai S, Durrett G. Calibration of pre-trained transformers. In: Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020: 295-302. DOI: 10.18653/v1/2020.emnlp-main.21.

[18] Geifman Y, El-Yaniv R. Selective classification for deep neural networks. In: Advances in Neural Information Processing Systems, 2017, 30.

[19] Begoli E, Bhattacharya T, Kusnezov D. The need for uncertainty quantification in machine-assisted medical decision making. Nature Machine Intelligence, 2019, 1: 20-23. DOI: 10.1038/s42256-018-0004-1.

[20] Van Calster B. Calibration: The Achilles heel of predictive analytics. BMC Medicine, 2019, 17: Art. 230. DOI: 10.1186/s12916-019-1466-7.

[21] Pal A, Umapathi L K, Sankarasubbu M. Med-HALT: Medical Domain Hallucination Test for Large Language Models. In: Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL), 2023: 314-334. DOI: 10.18653/v1/2023.conll-1.21.

[22] Bommasani R. On the opportunities and risks of foundation models. arXiv:2108.07258, 2021.

[23] Brown T B. Language models are few-shot learners. In: Advances in Neural Information Processing Systems, 2020, 33: 1877-1901.

[24] Wang X. Self-consistency improves chain of thought reasoning in language models. In: Proceedings of the International Conference on Learning Representations (ICLR), 2023.

[25] Kojima T, Gu S S, Reid M, et al. Large language models are zero-shot reasoners. In: Advances in Neural Information Processing Systems, 2022, 35: 22199-22213.

[26] Chung H W. Scaling instruction-finetuned language models. arXiv:2210.11416, 2022.

[27] Wei J. Finetuned language models are zero-shot learners. In: Proceedings of the International Conference on Learning Representations (ICLR), 2022.

[28] Ouyang L. Training language models to follow instructions with human feedback. In: Advances in Neural Information Processing Systems, 2022, 35: 27730-27744.

[29] Touvron H. Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288, 2023.

Downloads

Published

2024-01-13

Issue

Section

Research Article

DOI:

How to Cite

JiTong Zou, YuXuan Qin, MinJae Rhee. Beyond Model Scale: A Task-Aware Reliability Benchmark For Large Language Models In Clinical Decision Support. Innovation and Technology Studies. 2024, 1(1): 39-47. DOI: https://doi.org/10.61784/its4026.