Skip to content

E<inline-formula><tex-math notation="LaTeX">$^{2}$</tex-math><alternatives><mml:math><mml:msup><mml:mrow/><mml:mn>2</mml:mn></mml:msup></mml:math><inline-graphic xlink:href="feng-ieq1-3676689.gif"/></alternatives></inline-formula>LLM: Structure-Guided Efficient Inference for LLMs in Distributed Edge

Sep 2026 · IEEE Transactions on Mobile Computing · Vol 25, pp. 14239-14253 · 0 citations · 46 references

Abstract

Large language models (LLMs) are increasingly deployed in edge computing environments to reduce latency and preserve privacy. However, their inference process presents fundamental challenges for resource-constrained IoT devices. LLM inference involves computationally asymmetric stages: parallelizable prompt processing and sequential token decoding. This asymmetry creates deployment bottlenecks where IoT devices lack capacity for prompt processing while edge nodes suffer from inefficient sequential decoding. This paper presents <italic>E<inline-formula><tex-math notation="LaTeX">$^{2}$</tex-math><alternatives><mml:math><mml:msup><mml:mrow/><mml:mn>2</mml:mn></mml:msup></mml:math><inline-graphic xlink:href="feng-ieq3-3676689.gif"/></alternatives></inline-formula>LLM</italic>, an efficient distributed inference framework for large language models in heterogeneous edge-IoT environments. <italic>E<inline-formula><tex-math notation="LaTeX">$^{2}$</tex-math><alternatives><mml:math><mml:msup><mml:mrow/><mml:mn>2</mml:mn></mml:msup></mml:math><inline-graphic xlink:href="feng-ieq4-3676689.gif"/></alternatives></inline-formula>LLM</italic> leverages high-capacity edge devices for structural planning and introduces auxiliary lightweight models to generate segment-specific key-value (KV) caches. These minimal inference artifacts enable collaborative parallel decoding across IoT devices without requiring full model instantiation. The framework employs static-dynamic KV cache separation to minimize communication overhead while maintaining semantic coherence through structure-guided coordination. Extensive evaluation on realistic edge testbeds demonstrates significant performance improvements. Under diverse deployment settings, <italic>E<inline-formula><tex-math notation="LaTeX">$^{2}$</tex-math><alternatives><mml:math><mml:msup><mml:mrow/><mml:mn>2</mml:mn></mml:msup></mml:math><inline-graphic xlink:href="feng-ieq5-3676689.gif"/></alternatives></inline-formula>LLM</italic> achieves 74% –87.7% end-to-end latency reduction compared with several state-of-the-art baselines, while maintaining comparable generation quality; meanwhile, it also delivers a 34.6% –72.2% reduction in communication overhead, improves 9-12 × in energy efficiency. The framework exhibits strong scalability under bandwidth-limited conditions, enabling efficient LLM deployment across heterogeneous edge-IoT environments.

View source

Similar papers

Jun 2025

<inline-formula><tex-math notation="LaTeX">${\sf BandPilot}$</tex-math><alternatives><mml:math><mml:mi mathvariant="sans-serif">BandPilot</mml:mi></mml:math><inline-graphic xlink:href="tang-ieq1-3716225.gif"/></alternatives></inline-formula>: Toward Performance- and Contention-Aware GPU Dispatching

Modern multi-tenant AI clusters are increasingly communication-bound, driven by high-volume and multi-round GPU-to-GPU collective communication. Consequently, the GPU dispatcher’s choice of a physical GPU subset for each tenant largely determines the job’s effective collective bandwidth and thus its performance ceiling. Existing dispatchers predominantly rely on static, topology-aware heuristics that prioritize GPU resource compactness, assuming that minimizing physical distance maximizes communication bandwidth. However, we reveal that this assumption often fails due to complex system-level bottlenecks, such as non-linear NIC saturation and inter-node link heterogeneity. This paper presents <inline-formula><tex-math notation="LaTeX">${\sf BandPilot}$</tex-math><alternatives><mml:math><mml:mi mathvariant="sans-serif">BandPilot</mml:mi></mml:math><inline-graphic xlink:href="tang-ieq2-3716225.gif"/></alternatives></inline-formula>, a performance- and contention-aware GPU dispatching primitive that optimizes effective collective bandwidth for multi-tenant AI clusters. Specifically, <inline-formula><tex-math notation="LaTeX">${\sf BandPilot}$</tex-math><alternatives><mml:math><mml:mi mathvariant="sans-serif">BandPilot</mml:mi></mml:math><inline-graphic xlink:href="tang-ieq3-3716225.gif"/></alternatives></inline-formula> learns a data-efficient bandwidth model from sparse NCCL measurements via a hierarchical design. Guided by the model, <inline-formula><tex-math notation="LaTeX">${\sf BandPilot}$</tex-math><alternatives><mml:math><mml:mi mathvariant="sans-serif">BandPilot</mml:mi></mml:math><inline-graphic xlink:href="tang-ieq4-3716225.gif"/></alternatives></inline-formula> uses an equilibrium-driven heuristic as a fast front end, and invokes a pruned elimination search when a controller predicts that further refinement is worthwhile. To account for multi-tenant interference, <inline-formula><tex-math notation="LaTeX">${\sf BandPilot}$</tex-math><alternatives><mml:math><mml:mi mathvariant="sans-serif">BandPilot</mml:mi></mml:math><inline-graphic xlink:href="tang-ieq5-3716225.gif"/></alternatives></inline-formula> virtually merges a candidate allocation with co-located cross-host jobs to conservatively estimate shared bottleneck capacity and predict contention-degraded bandwidth. Across a 32-GPU H100 cluster and heterogeneous simulations, <inline-formula><tex-math notation="LaTeX">${\sf BandPilot}$</tex-math><alternatives><mml:math><mml:mi mathvariant="sans-serif">BandPilot</mml:mi></mml:math><inline-graphic xlink:href="tang-ieq6-3716225.gif"/></alternatives></inline-formula> achieves 90 – 97% bandwidth efficiency relative to the best-found reference, improving average efficiency by 20–30% over topology-compactness heuristics.

Kunming Zhang, Hanlong Liao, Junyu Xue et al. · 0 citations
Open access Aug 2026

The p-Modular Green Correspondence for $${\text {SL}_2(\mathbb {F}_p)}$$

<jats:p> Let <jats:italic>p</jats:italic> be an odd prime. Let <jats:inline-formula> <jats:alternatives> <jats:tex-math>$${G = SL_2(\mathbb {F}_p)}$$</jats:tex-math> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML"> <mml:mrow> <mml:mi>G</mml:mi> <mml:mo>=</mml:mo> <mml:mi>S</mml:mi> <mml:msub> <mml:mi>L</mml:mi> <mml:mn>2</mml:mn> </mml:msub> <mml:mrow> <mml:mo>(</mml:mo> <mml:msub> <mml:mi>F</mml:mi> <mml:mi>p</mml:mi> </mml:msub> <mml:mo>)</mml:mo> </mml:mrow> </mml:mrow> </mml:math> </jats:alternatives> </jats:inline-formula> and let <jats:italic>B</jats:italic> denote the subgroup of upper triangular matrices of <jats:italic>G</jats:italic> . Finally, let <jats:inline-formula> <jats:alternatives> <jats:tex-math>$${\mathbb {F}}$$</jats:tex-math> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML"> <mml:mi>F</mml:mi> </mml:math> </jats:alternatives> </jats:inline-formula> be an algebraically closed field of characteristic <jats:italic>p</jats:italic> . The Green correspondence gives a bijection between the non-projective indecomposable <jats:inline-formula> <jats:alternatives> <jats:tex-math>$${\mathbb {F}[G]}$$</jats:tex-math> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML"> <mml:mrow> <mml:mi>F</mml:mi> <mml:mo>[</mml:mo> <mml:mi>G</mml:mi> <mml:mo>]</mml:mo> </mml:mrow> </mml:math> </jats:alternatives> </jats:inline-formula> modules and non-projective indecomposable <jats:inline-formula> <jats:alternatives> <jats:tex-math>$${\mathbb {F}[B]}$$</jats:tex-math> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML"> <mml:mrow> <mml:mi>F</mml:mi> <mml:mo>[</mml:mo> <mml:mi>B</mml:mi> <mml:mo>]</mml:mo> </mml:mrow> </mml:math> </jats:alternatives> </jats:inline-formula> modules, realised by restriction and induction. In this paper, after recalling a suitable description of the non-projective indecomposable modules for these group algebras, we explicitly describe the Green correspondence bijection. We do this by pinpointing the modules’ position on the Stable Auslanden-Reiten quivers. Finally, we obtain two corollaries in terms of this description: formula for lifting the <jats:inline-formula> <jats:alternatives> <jats:tex-math>$${\mathbb {F}[B]}$$</jats:tex-math> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML"> <mml:mrow> <mml:mi>F</mml:mi> <mml:mo>[</mml:mo> <mml:mi>B</mml:mi> <mml:mo>]</mml:mo> </mml:mrow> </mml:math> </jats:alternatives> </jats:inline-formula> module decomposition of an <jats:inline-formula> <jats:alternatives> <jats:tex-math>$${\mathbb {F}[G]}$$</jats:tex-math> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML"> <mml:mrow> <mml:mi>F</mml:mi> <mml:mo>[</mml:mo> <mml:mi>G</mml:mi> <mml:mo>]</mml:mo> </mml:mrow> </mml:math> </jats:alternatives> </jats:inline-formula> module, and a complete description of <jats:inline-formula> <jats:alternatives> <jats:tex-math>$${\text { Ind}_B^G}$$</jats:tex-math> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML"> <mml:mrow> <mml:mspace/> <mml:msubsup> <mml:mtext>Ind</mml:mtext> <mml:mi>B</mml:mi> <mml:mi>G</mml:mi> </mml:msubsup> </mml:mrow> </mml:math> </jats:alternatives> </jats:inline-formula> and <jats:inline-formula> <jats:alternatives> <jats:tex-math>$${\text { Res}^G_B}$$</jats:tex-math> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML"> <mml:mrow> <mml:mspace/> <mml:msubsup> <mml:mtext>Res</mml:mtext> <mml:mi>B</mml:mi> <mml:mi>G</mml:mi> </mml:msubsup> </mml:mrow> </mml:math> </jats:alternatives> </jats:inline-formula> . </jats:p>

Denver-James Logan Marchment · 0 citations
Open access Jul 2026

Condition for <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" display="inline"> <mml:mrow> <mml:mn>1</mml:mn> <mml:mi>/</mml:mi> <mml:mi>f</mml:mi> </mml:m

While $1/f$ noise is ubiquitous and has been found in various systems, its physics remains uncertain. From an analytical study of an ordinary diffusion equation, we find an additional example of the $1/f$ noise. The formula for this example, together with existing knowledge about scaling in fluid turbulence, implies a necessary and sufficient condition for the occurrence of any stationary $1/f$ noise. That is, the noise needs to be characterized by two constant frequencies of $f_{\rm low} \ll f_{\rm high}$. For a frequency range from $f = f_{\rm low}$ to $f_{\rm high}$, it is further needed that, except for the mean amplitude of the noise, there is no other constant parameter. Then, at $f_{\rm low} \ll f \ll f_{\rm high}$, the noise scales asymptotically as $1/f$. Being statistical and simple, our condition applies to any system and hence explains the ubiquity of the $1/f$ noise. It is also applicable to some systems with noise of $\alpha \ne 1.0$ for $1/f^{\alpha}$, via intermittency analogous to that of the turbulence.

H. Mouri · 0 citations
Open access Aug 2026

Core breaking at low spin in <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" display="inline"> <mml:msup> <mml:mi/> <mml:mn>68</mml:mn> </mml:msup> </mml:math>

Low-spin excited states in $^{68}$Zn have been studied at the High Intensity Gamma-Ray Source (HI$\gamma$S) from the ground state up to the particle emission threshold using the nuclear resonance fluorescence technique (NRF) and the newly developed Clover Array. Low-spin levels were excited by linearly-polarized, $2.90 - 9.79$ MeV photon beams. Spin-parity quantum numbers as well as associated $M1$ and $E1$ decay strengths were determined for a large fraction of the 158 states observed. In addition, long-duration coincidence measurements at 9.46 and 9.79 MeV enabled the investigation of the level scheme near the ground state. The results have been interpreted with shell-model calculations using two different model spaces and several effective interactions often used to describe nuclei in this mass region. While the structure near the ground state can be understood in terms of excitations involving solely valence nucleons, core breaking is required to account for the evolution of the total $M1$ strength at excitation energies above $\sim5$ MeV.

S. R. Johnson, R. V. Janssens, B. A. Brown et al. · 0 citations
Open access Mar 2024

Arithmetic of Critical 𝑝-Adic 𝐿-Functions

<p> Our objective in the present work is to develop a fairly complete arithmetic theory of critical <inline-formula content-type="math/mathml"> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" alttext="p"> <mml:semantics> <mml:mi>p</mml:mi> <mml:annotation encoding="application/x-tex">p</mml:annotation> </mml:semantics> </mml:math> </inline-formula> -adic <inline-formula content-type="math/mathml"> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" alttext="upper L"> <mml:semantics> <mml:mi>L</mml:mi> <mml:annotation encoding="application/x-tex">L</mml:annotation> </mml:semantics> </mml:math> </inline-formula> -functions on the eigencurve. To this end, we carry out the following tasks: </p> <p> We give an “étale” construction of Bellaïche’s <inline-formula content-type="math/mathml"> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" alttext="p"> <mml:semantics> <mml:mi>p</mml:mi> <mml:annotation encoding="application/x-tex">p</mml:annotation> </mml:semantics> </mml:math> </inline-formula> -adic <inline-formula content-type="math/mathml"> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" alttext="upper L"> <mml:semantics> <mml:mi>L</mml:mi> <mml:annotation encoding="application/x-tex">L</mml:annotation> </mml:semantics> </mml:math> </inline-formula> -functions at a <inline-formula content-type="math/mathml"> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" alttext="theta"> <mml:semantics> <mml:mi> θ </mml:mi> <mml:annotation encoding="application/x-tex">\theta</mml:annotation> </mml:semantics> </mml:math> </inline-formula> -critical point on the cuspidal Coleman–Mazur–Buzzard eigencurve. </p> <p> We introduce the algebraic counterparts of these objects (which arise as appropriately defined Selmer complexes) and develop Iwasawa theory in this context, including a definition of an Iwasawa theoretic <inline-formula content-type="math/mathml"> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" alttext="script upper L"> <mml:semantics> <mml:mrow class="MJX-TeXAtom-ORD"> <mml:mi mathvariant="script">L</mml:mi> </mml:mrow> <mml:annotation encoding="application/x-tex">\mathscr L</mml:annotation> </mml:semantics> </mml:math> </inline-formula> -invariant <inline-formula content-type="math/mathml"> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" alttext="script upper L Subscript upper I w Superscript c r"> <mml:semantics> <mml:msubsup> <mml:mrow class="MJX-TeXAtom-ORD"> <mml:mi mathvariant="script">L</mml:mi> </mml:mrow> <mml:mrow class="MJX-TeXAtom-ORD"> <mml:mi>I</mml:mi> <mml:mi>w</mml:mi> </mml:mrow> <mml:mrow class="MJX-TeXAtom-ORD"> <mml:mi>c</mml:mi> <mml:mi>r</mml:mi> </mml:mrow> </mml:msubsup> <mml:annotation encoding="application/x-tex">\mathscr {L}^{cr}_{Iw}</mml:annotation> </mml:semantics> </mml:math> </inline-formula> . </p> <p>We formulate the (punctual) critical main conjectures, and study its relationship with its slope-zero counterpart. Along the way, we also develop descent theory (paralleling Perrin-Riou’s work).</p> <p> We introduce what we call <italic>thick</italic> (Iwasawa theoretic) fundamental line and the <italic>thick</italic> Selmer complex to counter Bellaïche’s secondary <inline-formula content-type="math/mathml"> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" alttext="p"> <mml:semantics> <mml:mi>p</mml:mi> <mml:annotation encoding="application/x-tex">p</mml:annotation> </mml:semantics> </mml:math> </inline-formula> -adic <inline-formula content-type="math/mathml"> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" alttext="upper L"> <mml:semantics> <mml:mi>L</mml:mi> <mml:annotation encoding="application/x-tex">L</mml:annotation> </mml:semantics> </mml:math> </inline-formula> -functions. This allows us to formulate an infinitesimal thickening of the Iwasawa main conjecture, and we observe that it implies both slope-zero and punctual critical main conjectures, but it seems stronger than both. </p> <p> We establish an <inline-formula content-type="math/mathml"> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" alttext="script upper O Subscript script upper X"> <mml:semantics> <mml:msub> <mml:mrow class="MJX-TeXAtom-ORD"> <mml:mi class="MJX-tex-caligraphic" mathvariant="script">O</mml:mi> </mml:mrow> <mml:mrow class="MJX-TeXAtom-ORD"> <mml:mi class="MJX-tex-caligraphic" mathvariant="script">X</mml:mi> </mml:mrow> </mml:msub> <mml:annotation encoding="application/x-tex">\mathcal {O}_\mathcal {X}</mml:annotation> </mml:semantics> </mml:math> </inline-formula> -adic leading term formula for the two-variable <inline-formula content-type="math/mathml"> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" alttext="p"> <mml:semantics> <mml:mi>p</mml:mi> <mml:annotation encoding="application/x-tex">p</mml:annotation> </mml:semantics> </mml:math> </inline-formula> -adic <inline-formula content-type="math/mathml"> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" alttext="upper L"> <mml:semantics> <mml:mi>L</mml:mi> <mml:annotation encoding="application/x-tex">L</mml:annotation> </mml:semantics> </mml:math> </inline-formula> -function over the affinoid neighborhood <inline-formula content-type="math/mathml"> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" alttext="script upper X equals upper S p m left-parenthesis script upper O Subscript script upper X Baseline right-parenthesis"> <mml:semantics> <mml:mrow> <mml:mrow class="MJX-TeXAtom-ORD"> <mml:mi class="MJX-tex-caligraphic" mathvariant="script">X</mml:mi> </mml:mrow> <mml:mo>=</mml:mo> <mml:mrow class="MJX-TeXAtom-ORD"> <mml:mi>S</mml:mi> <mml:mi>p</mml:mi> <mml:mi>m</mml:mi> </mml:mrow> <mml:mo stretchy="false">(</mml:mo> <mml:msub> <mml:mrow class="MJX-TeXAtom-ORD"> <mml:mi class="MJX-tex-caligraphic" mathvariant="script">O</mml:mi> </mml:mrow> <mml:mrow class="MJX-TeXAtom-ORD"> <mml:mi class="MJX-tex-caligraphic" mathvariant="script">X</mml:mi> </mml:mrow>

D. Benois, Kazim Buyukboduk · 3 citations · ⚡1

Related blog posts