PRISM: Proximal policy optimization with deep Reinforcement learning for Intelligent Scheduling in Multipath QUIC under heterogeneous and hybrid 5G/B5G–satellite networks
2026 (English)In: Computer Communications, ISSN 0140-3664, E-ISSN 1873-703X, Vol. 252, article id 108523Article in journal (Refereed) Published
Abstract [en]
The increasing reliance on and integration of heterogeneous access networks such as 5G, Beyond 5G (B5G), and satellite communication (Satcom) systems have not been adequately leveraged in existing multipath transport schedulers. Although protocols like Multipath TCP (MPTCP) and Multipath QUIC (MPQUIC) enable bandwidth aggregation across multiple paths, current schedulers largely rely on reactive decisions and struggle to cope with rapidly varying hybrid network conditions involving bandwidth fluctuations, intermittent outages, and heterogeneous link characteristics. This work presents PRISM (Proximal policy optimization-based Reinforcement learning for Intelligent Scheduling of Multipaths), a deep reinforcement learning (DRL)-based scheduler designed for MPQUIC in hybrid 5G and Satcom environments. PRISM follows a hybrid-adaptive approach by integrating LSTM-enhanced actor–critic learning with an adaptive exploration strategy, allowing the scheduler to leverage temporal network behavior and make proactive scheduling decisions. The design remains lightweight, enabling efficient operation with low CPU and memory overhead, which is critical for real-time deployment. We evaluate PRISM under diverse and dynamically changing network conditions, including bandwidth variability and link disruptions, using real-world traces from Lumos5G and Starlink Satcom datasets. The evaluation compares PRISM against state-of-the-art learning-based schedulers (Peekaboo and DEAR) as well as widely used rule-based schedulers (RR, ECF, BLEST, and MinRTT). Experimental results show that PRISM consistently achieves superior performance, providing improvements of up to 25.64%–31.25% over other learning-based schedulers and substantially higher gains over rule-based approaches across heterogeneous network scenarios.
Place, publisher, year, edition, pages
Elsevier, 2026. Vol. 252, article id 108523
Keywords [en]
5G/B5G networks, DRL, MPQUIC, MPTCP, Multipath networking, Multipath schedulers, Satcom networks, Bandwidth, Deep learning, Deep reinforcement learning, Heterogeneous networks, Internet protocols, Multipath propagation, Reinforcement learning, Satellite communication systems, 5g/beyond 5g network, Multipath, Multipath QUIC, Multipath scheduler, Multipath TCP, Reinforcement learnings, Satellite communication networks, Prisms
National Category
Computer Sciences
Research subject
Computer Science
Identifiers
URN: urn:nbn:se:kau:diva-109771DOI: 10.1016/j.comcom.2026.108523ISI: 001742091200001Scopus ID: 2-s2.0-105035004529OAI: oai:DiVA.org:kau-109771DiVA, id: diva2:2054224
2026-04-202026-04-202026-04-27Bibliographically approved