In this paper, we investigate the angle-of-arrival (AoA) estimation problem for wireless sensing systems equipped with movable antennas (MA). To achieve high estimation performance and accuracy, we formulate a joint optimization problem integrating the sidelobes of steering vector correlation (SVC) and the Cram\'er-Rao bound (CRB). We first mathematically transform the SVC and the CRB into tractable objective functions. Specifically, we introduce a proxy variable and apply a discrete grid search strategy to overcome the intractability of optimizing the SVC with unknown target angles. Concurrently, we derive a generalized lower bound for the CRB, which yields a scalar function of the MA positions. Guided by the transformed objective, we propose a successive convex approximation-based position optimization algorithm. The proposed algorithm handles the non-convex terms by employing first order Taylor expansions within a defined trust region, which allows the MA positions to be updated incrementally in each iteration. Simulation results demonstrate that the proposed algorithm achieves superior AoA estimation performance.
Chengzhi Ye, Ruoyu Zhang, Shichao Wei et al.· 1 citation
Diffusion multimodal large language models (dMLLMs) have recently emerged as a new decoding paradigm for multimodal generation. Starting from a fully masked sequence, dMLLMs progressively decode the sequence by unmasking a subset of the remaining masked positions at each step. Since the selected tokens serve as the prediction context for subsequent steps, deciding which tokens to decode is crucial to the quality of the final output. The most common strategy prioritizes tokens based on a certainty measure that tends to favor tokens frequently observed in the training data. Recent approaches instead order tokens according to their influence on subsequent predictions, but do not explicitly account for the input image. We propose the Visual Information-Guided Sampler (VIG-Sampler), which prioritizes tokens based on their attention to image tokens. We further impose a constraint that penalizes candidate tokens whose image-attention distributions are similar to those of previously selected tokens, thereby increasing the information gain of the decoded subset. Extensive experiments on 7 captioning and VQA benchmarks with 3 open-source dMLLMs demonstrate the effectiveness of VIG-Sampler, which outperforms the Info-Gain Sampler by an average of 19.3 CIDEr points across the captioning benchmarks and surpasses it on COCO Caption while using only half as many decoding steps.
Insu Lee, Woo-Soon Park, Wonseok Shin et al.· 1 citation
DARD is proposed, a training-free framework that separates tokens into masked, candidate, and unmasked states and adaptively regulates their influence on subsequent decoding, and consistently improves the speed-quality Pareto frontier over recent revocable decoding methods.
Woo-Soon Park, Insu Lee, Minyoung Noh et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.