CLARITY: A CLOSED-LOOP FRAMEWORK FOR DUPLICATE-AWARE FAILURE PREDICTION AND BUDGETED AUTOMATED REMEDIATION IN CLOUD INFRASTRUCTURE

Authors

  • Minjae Rhee University of Illinois Urbana-Champaign, Champaign, IL 61820, Illinois, USA.
  • JiTong Zou Case Western Reserve University, Cleveland, OH 44106, Ohio, USA.
  • YuXuan Qin (Corresponding Author) Northeastern University, Boston, MA 02115, Massachusetts, USA.

Keywords:

Cloud infrastructure, Failure prediction, Log anomaly detection, Automated remediation, Probability calibration, Closed-loop AIOps

Abstract

Failure predictors for cloud infrastructure are often evaluated as isolated classifiers, although operators need calibrated risk, bounded actions, and feedback after remediation. This paper presents CLARITY, a closed-loop framework that couples duplicate-aware early prediction, probability calibration, a budget-constrained remediation policy, safety guardrails, and post-action feedback. The evaluation uses 575,061 labeled block-level event sequences from the public Loghub HDFS_v1 corpus. Exact duplicate sequences are grouped before splitting, and ten contradictory sequence patterns (46 occurrences) are removed rather than silently assigned. A prefix model combines normalized event frequencies and ordered transitions; XGBoost provides ranking and Platt scaling converts scores into operational probabilities. Five-fold grouped evaluation shows that observing 16 events yields an occurrence-weighted PR-AUC of 0.433±0.066 and a pattern-level PR-AUC of 0.751±0.139. Calibration reduces occurrence-weighted Brier score from 0.467 to 0.019 without changing ranking. In a trace-driven counterfactual policy, a 2% action budget captures 49.26% of anomalous remaining work and produces 11.10% normalized net savings under an explicitly stated 0.8 action efficacy and three-event action cost. Under controlled template-ID churn, feedback raises post-drift pattern PR-AUC from 0.190 to 0.470. The remediation results are policy simulations, not production intervention measurements. The study provides an implementable framework and a reproducible boundary between measured evidence and counterfactual assumptions.

References

[1] Verma A, Pedrosa L, Korupolu M R, et al. Large-scale cluster management at Google with Borg. Proceedings of ACM EuroSys, 2015: 18, 1-17. DOI: 10.1145/2741948.2741964.

[2] Cortez E, Bonde A, Muzio A, et al. Resource Central: Understanding and predicting workloads for improved resource management in large cloud platforms. Proceedings of ACM SOSP, 2017: 153-167. DOI: 10.1145/3132747.3132772.

[3] Dean J, Barroso L A. The tail at scale. Communications of the ACM, 2013, 56(2): 74-80. DOI: 10.1145/2408776.2408794.

[4] Xu W, Huang L, Fox A, et al. Detecting large-scale system problems by mining console logs. Proceedings of ACM SOSP, 2009: 117-132. DOI: 10.1145/1629575.1629587.

[5] Zhu J, He S, He P, et al. Loghub: A large collection of system log datasets for AI-driven log analytics. Proceedings of IEEE ISSRE, 2023: 355-366. DOI: 10.1109/ISSRE59848.2023.00071.

[6] He S, Zhu J, He P, et al. Experience report: System log analysis for anomaly detection. Proceedings of IEEE ISSRE, 2016: 207-218. DOI: 10.1109/ISSRE.2016.21.

[7] He P, Zhu J, Zheng Z, et al. Drain: An online log parsing approach with fixed depth tree. Proceedings of IEEE ICWS, 2017: 33-40. DOI: 10.1109/ICWS.2017.13.

[8] He S, He P, Chen Z, et al. A survey on automated log analysis for reliability engineering. ACM Computing Surveys, 2021, 54(6): 130. DOI: 10.1145/3460345.

[9] Oliner A J, Ganapathi A, Xu W. Advances and challenges in log analysis. Communications of the ACM, 2012, 55(2): 55-61. DOI: 10.1145/2076450.2076466.

[10] Du M, Li F, Zheng G, et al. DeepLog: Anomaly detection and diagnosis from system logs through deep learning. Proceedings of ACM CCS, 2017: 1285-1298. DOI: 10.1145/3133956.3134015.

[11] Meng W. LogAnomaly: Unsupervised detection of sequential and quantitative anomalies in unstructured logs. Proceedings of IJCAI, 2019: 4739-4745. DOI: 10.24963/ijcai.2019/658.

[12] Zhang X. Robust log-based anomaly detection on unstable log data. Proceedings of ACM ESEC/FSE, 2019: 807-817. DOI: 10.1145/3338906.3338931.

[13] Yang L, Chen J, Wang Z, et al. PLELog: Semi-supervised log-based anomaly detection via probabilistic label estimation. Proceedings of IEEE/ACM ICSE Companion, 2021: 230-231. DOI: 10.1109/ICSE-Companion52605.2021.00106.

[14] Guo H, Yuan S, Wu X. LogBERT: Log anomaly detection via BERT. Proceedings of IJCNN, 2021: 1-8. DOI: 10.1109/IJCNN52387.2021.9534113.

[15] Le V H, Zhang H. Log-based anomaly detection with deep learning: How far are we? Proceedings of ACM/IEEE ICSE, 2022: 1356-1367. DOI: 10.1145/3510003.3510155.

[16] Lin Q. Predicting node failure in cloud service systems. Proceedings of ACM ESEC/FSE, 2018: 480-490. DOI: 10.1145/3236024.3236060.

[17] Dang Y, Lin Q, Huang P. AIOps: Real-world challenges and research innovations. Proceedings of IEEE/ACM ICSE Companion, 2019: 4-5. DOI: 10.1109/ICSE-Companion.2019.00023.

[18] Cheng Y. Analyzing Alibaba’s co-located datacenter workloads. Proceedings of IEEE Big Data, 2018: 292-297. DOI: 10.1109/BigData.2018.8622518.

[19] Reiss C, Wilkes J, Hellerstein J L. Heterogeneity and dynamicity of clouds at scale: Google trace analysis. Proceedings of ACM SoCC, 2012: 7, 1-13. DOI: 10.1145/2391229.2391236.

[20] Gan Y. Seer: Leveraging big data to navigate the complexity of performance debugging in cloud microservices. Proceedings of ACM ASPLOS, 2019: 19-33. DOI: 10.1145/3293882.3330558.

[21] Chen T, Guestrin C. XGBoost: A scalable tree boosting system. Proceedings of ACM KDD, 2016: 785-794. DOI: 10.1145/2939672.2939785.

[22] Guo C, Pleiss G, Sun Y, et al. On calibration of modern neural networks. Proceedings of ICML, PMLR, 2017, 70: 1321-1330.

[23] Niculescu-Mizil A, Caruana R. Predicting good probabilities with supervised learning. Proceedings of ICML, 2005: 625-632. DOI: 10.1145/1102351.1102430.

[24] Saito T, Rehmsmeier M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS One, 2015, 10(3): e0118432. DOI: 10.1371/journal.pone.0118432.

[25] Sculley D. Hidden technical debt in machine learning systems. Advances in Neural Information Processing Systems, 2015: 2503-2511.

[26] Breck E, Cai S, Nielsen E, et al. The ML test score: A rubric for ML production readiness and technical debt reduction. Proceedings of IEEE Big Data, 2017: 1123-1132. DOI: 10.1109/BigData.2017.8258038.

Downloads

Published

2024-12-30

Issue

Section

Research Article

DOI:

How to Cite

Minjae Rhee, JiTong Zou, YuXuan Qin. Clarity: A Closed-Loop Framework For Duplicate-Aware Failure Prediction And Budgeted Automated Remediation In Cloud Infrastructure. AI and Data Science Journal. 2024, 1(1): 76-85. DOI: https://doi.org/10.61784/adsj4039.