Skip to main navigation Skip to search Skip to main content

Serving Long-Context LLMs at the Mobile Edge: Test-Time Reinforcement Learning-Based Model Caching and Inference Offloading

Research output: Contribution to journalArticlepeer-review

Abstract

Large Language Models (LLMs) can perform zero-shot learning on unseen tasks and few-shot learning on complex reasoning tasks. However, resource-limited mobile edge networks struggle to support long-context LLM serving for LLM agents during multi-round interactions with users. Unlike stateless computation offloading and static service offloading in edge computing, optimizing LLM serving at edge servers is challenging because LLMs continuously learn from context which raises accuracy, latency, and resource consumption dynamics. In this paper, we propose a joint model caching and inference offloading framework that utilizes test-time deep reinforcement learning (T2DRL) to optimize deployment and execution strategies for long-context LLM serving. In this framework, we analyze the performance convergence and design an optimization problem considering the utilization of context windows in LLMs. Furthermore, the T2DRL algorithm can learn in both the training phase and the testing phase to proactively manage cached models and service requests and adapt to context changes and usage patterns during execution. To further enhance resource allocation efficiency, we propose a double Dutch auction (DDA) mechanism, which dynamically aligns the marginal value of an additional reasoning path with the marginal cost of reasoning services. Finally, experimental results demonstrate that the T2DRL algorithm can reduce system costs by at least 30% compared to baselines while guaranteeing the performance of LLM agents in real-world perception and reasoning tasks.

Original languageEnglish (US)
Pages (from-to)3808-3823
Number of pages16
JournalIEEE/ACM Transactions on Networking
Volume34
DOIs
StatePublished - 2026
Externally publishedYes

All Science Journal Classification (ASJC) codes

  • Computer Networks and Communications
  • Computer Science Applications
  • Software
  • Electrical and Electronic Engineering

Keywords

  • Mobile edge networks
  • auction theory
  • deep reinforcement learning (DRL)
  • large language models (LLMs)

Fingerprint

Dive into the research topics of 'Serving Long-Context LLMs at the Mobile Edge: Test-Time Reinforcement Learning-Based Model Caching and Inference Offloading'. Together they form a unique fingerprint.

Cite this