HMCO-AT: A Hierarchical Memory - Compute Orchestrated Adaptive Training Framework for Resource-Efficient and Adversarially Robust Large Models in Cloud–Edge Systems
DOI:
https://doi.org/10.63313/JCSFT.9085Keywords:
Cloud-Edge Collaboration, Adversarial Training, Memory-Compute Optimization, Distributed Deep Learning, Large Models, Training SchedulerAbstract
Adversarial training has emerged as the most effective paradigm for improving the robustness of deep neural networks; however, it imposes substantial computational and memory demands, particularly during the adversarial example generation process. These challenges become more critical in cloud-edge collaborative learning scenarios, where edge devices have limited resources and communication bandwidth is highly variable. To address this problem, we propose HMCO-AT, a Hierarchical Memory-Compute Orchestrated Adaptive Training framework designed specifically for resource-efficient adversarial training of large-scale models. HMCO-AT explicitly models the memory footprint and computation peaks induced by adversarial sample generation and dynamically schedules model layers, training stages, and communication operations across cloud and edge nodes. Our framework integrates hierarchical parameter partitioning, memory-aware activation recomputation, quantized gradient transmission, and an RL-based adaptive scheduler that jointly optimizes memory consumption, computational load, and robustness performance. HMCO-AT is validated on large-scale language-modeling corpora including OpenWebText, WikiCorpus, and filtered C4 under realistic cloud–edge conditions. Experiments show that HMCO-AT reduces peak edge memory by 64.8% and lowers computation cost and latency by 35.8% and 26.7%, respectively. It also improves adversarial robustness, achieving 7.2% higher FGSM and 9.3% higher PGD-10 accuracy than strong baselines. These results confirm HMCO-AT as an efficient and robust framework for large-model cloud–edge training.
References
[1] Shoeybi M, Patwary M, Puri R, et al. Megatron-lm: Training multi-billion parameter language models using model parallelism[J]. arXiv preprint arXiv:1909.08053, 2019.
[2] Huang Y, Cheng Y, Bapna A, et al. Gpipe: Efficient training of giant neural networks using pipeline parallelism[J]. Advances in neural information processing systems, 2019, 32.
[3] Harlap A, Narayanan D, Phanishayee A, et al. PipeDream: Pipeline parallelism for DNN training[C]//Proceedings of the 1st Conference on Systems and Machine Learning (SysML). 2018.
[4] Narayanan D, Phanishayee A, Shi K, et al. Memory-efficient pipeline-parallel dnn training[C]//International Conference on Machine Learning. PMLR, 2021: 7937-7947.
[5] Rajbhandari S, Rasley J, Ruwase O, et al. Zero: Memory optimizations toward training trillion parameter models[C]//SC20: international conference for high performance computing, networking, storage and analysis. IEEE, 2020: 1-16.
[6] Cai G, Tian R, Yang L, et al. Efficient Inference for Edge Large Language Models: A Survey[J]. Tsinghua Science and Technology, 2026, 31(3): 1365-1380.
[7] Madry A, Makelov A, Schmidt L, et al. Towards deep learning models resistant to adversarial attacks[J]. arXiv preprint arXiv:1706.06083, 2017.
[8] Zhang H, Yu Y, Jiao J, et al. Theoretically principled trade-off between robustness and accuracy[C]//International conference on machine learning. PMLR, 2019: 7472-7482.
[9] Shafahi A, Najibi M, Ghiasi M A, et al. Adversarial training for free![J]. Advances in neural information processing systems, 2019, 32.
[10] Wong E, Rice L, Kolter J Z. Fast is better than free: Revisiting adversarial training[J]. arXiv preprint arXiv:2001.03994, 2020.
[11] Andriushchenko M, Flammarion N. Understanding and improving fast adversarial training[J]. Advances in Neural Information Processing Systems, 2020, 33: 16048-16059.
[12] Wu D, Xia S T, Wang Y. Adversarial weight perturbation helps robust generalization[J]. Advances in neural information processing systems, 2020, 33: 2958-2969.
[13] Yu C, Han B, Gong M, et al. Robust weight perturbation for adversarial training[J]. arXiv preprint arXiv:2205.14826, 2022.
[14] Grathwohl W, Wang K C, Jacobsen J H, et al. Your classifier is secretly an energy based model and you should treat it like one[J]. arXiv preprint arXiv:1912.03263, 2019.
[15] Gupta O, Raskar R. Distributed learning of deep neural network over multiple agents[J]. Journal of Network and Computer Applications, 2018, 116: 1-8.
[16] Hu Y, Imes C, Zhao X, et al. Pipeedge: Pipeline parallelism for large-scale model inference on heterogeneous edge devices[C]//2022 25th Euromicro Conference on Digital System Design (DSD). IEEE, 2022: 298-307.
[17] Yoon J Y, Byeon Y, Kim J, et al. Edgepipe: Tailoring pipeline parallelism with deep neural networks for volatile wireless edge devices[J]. IEEE Internet of Things Journal, 2021, 9(14): 11633-11647.
[18] Huang C C, Jin G, Li J. Swapadvisor: Pushing deep learning beyond the gpu memory limit via smart swapping[C]//Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems. 2020: 1341-1355.
[19] Mohan J, Phanishayee A, Chidambaram V. {CheckFreq}: Frequent,{Fine-Grained}{DNN} checkpointing[C]//19th USENIX Conference on File and Storage Technologies (FAST 21). 2021: 203-216.
[20] Meng Y, Iyer M A, Prasanna V K. An acceleration framework for deep reinforcement learning using heterogeneous systems[J]. IEEE Transactions on Parallel and Distributed Systems, 2025, 36(7): 1401-1415.
[21] Tang L, Wang Y, Willke T L, et al. Scheduling computation graphs of deep learning models on manycore cpus[J]. arXiv preprint arXiv:1807.09667, 2018.
[22] Lin, Ziyu, and Biliang Wang. "Adaptive load balancing algorithms for cloud computing distributed systems." IET Conference Proceedings CP952. Vol. 2025. No. 39. Stevenage, UK: The Institution of Engineering and Technology, 2025.
[23] Li J, Zeng P, Luo P. CANAO: A Cloud-Aware Native Agentic AI Framework for Adaptive Task Orchestration in Cloud-Native Environments[J]. Frontiers in Artificial Intelligence Research, 2026, 3(1): 187-198.
[24] Wang Y. Low-power design of advanced image processing algorithms under fpga in real-time applications[C]//2024 IEEE 4th International Conference on Power, Electronics and Computer Applications (ICPECA). IEEE, 2024: 1080-1084.
[25] Xu S, Jiang L, Gu B. Design and Validation of a Smart Neuromorphic System Architecture for Algorithmic Trading[C]//Proceedings of the 2nd International Symposium on Integrated Circuit Design and Integrated Systems. 2025: 127-136.
[26] Li Z, Hao Y, Zeng P. CLASNet: A Cognitive Load–Aware CNN-LSTM-Attention Framework for Supply Chain Demand Forecasting and Adaptive Human–Computer Interaction[J]. Journal of Computer Science and Frontier Technologies, 2026, 3(2): 44-57.
[27] Sun Q, Zhao X, Lin X. Design of a Hardware-Software Co-designed Real-Time Machine Learning System for Big Data Streams[C]//Proceedings of the 2nd International Symposium on Integrated Circuit Design and Integrated Systems. 2025: 265-271.
[28] Yao H, Ding J, Wang Z. CloudPayGuard: Hardware-Aware Real-Time Fraud Detection for Cloud-Native Credit Systems with OoO CPU Microarchitecture Optimization[J]. Journal of Computer Science and Frontier Technologies, 2026, 3(2): 154-166.
[29] Yao Y, Zhang W, Li M. Cloud-Edge Federated Incremental Learning Framework for PMSM Efficiency Optimization with Lightweight CNN-LSTM Models and OTA Differential Deployment[J]. Journal of Computer Science and Frontier Technologies, 2026, 3(1): 98-111.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 by author(s) and Erytis Publishing Limited

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.













