Skip to content

Author

Haoxu Wang

We have 4 of 51 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Sep 2026

TWZA-MVTFSF: A novel short-term photovoltaic power forecasting model

Photovoltaic (PV) power generation is affected by meteorological and geographical factors, displaying volatility, nonlinearity, and characteristics across multiple timescales. Accurate short-term forecasting of PV power is essential for enhancing PV resource utilization in the power grid and maintaining the quality of renewable energy integration. This paper introduces a Multi-Variable Time-Frequency Synergistic Fusion network model, enhanced by the Triangular Wave Zebra Algorithm (TWZA), for short-term PV power forecasting. To tackle the challenges of disjointed time-frequency feature fusion and limited generalization in complex fluctuation scenarios, a dual-parallel branch structure is developed to synergistically model temporal and spectral features in both time and frequency domains. In the time-domain branch, a kLSTM encoder is employed, which integrates block embedding and a Multi-Variable Correlation Attention mechanism to uniformly model variable coupling relationships and long-term dependencies. The frequency-domain branch incorporates the frequency decomposition architecture of the Multi-order Kolmogorov–Arnold Network (KAN), decomposing the power signal into multi-scale frequency components. By combining multi-order KAN representation learning with deep separable convolutions, it extracts nonlinear spectral features. A cross-domain attention-based time-frequency fusion mechanism is developed for adaptive and collaborative integration of time-domain and frequency-domain features. Second, to address the inability of time-frequency models to adaptively tune hyperparameters, we introduce TWZA to adaptively optimize time-frequency fusion weights, network hyperparameters, and convolution kernel sizes. This enhances the model's global search and local exploration capabilities while reducing training loss. Extensive experimental validation demonstrates that the proposed model outperforms all baseline models in terms of R2, root mean square error, and mean absolute error, achieving improvements of at least 18.2%, 27.3%, and 14.9%, respectively.

Wan-Nian Wei, Zhiwen Wang, Haoxu Wang et al. · 0 citations
Open access Jul 2026

Short-Term Electrical Load Forecasting Based on IMBKA-BiGRU-Attention Model

This paper proposes a short-term electrical load forecasting method based on a BiGRU-Attention network optimized by an improved multi-strategy black-winged kite algorithm (IMBKA), which achieves favorable forecasting performance among the compared models.

Binglin Liang, Zhiwen Wang, Bo Tian et al. · 0 citations
Preprint Aug 2026

Beyond Reconstruction: Full-Context Generative DiT for Music Generation

FullDiT is introduced, a conditional DiT that fuses eight frame-aligned RVQ streams with independently encoded captions and lyrics and uses non-causal self-attention over the complete acoustic latent sequence and outperforms five commercial systems on 15 of 18 automatic metrics.

Yun-Jia Li, Meng-Li Wu, Jun-Yu Dai et al. · 0 citations
Preprint Jul 2026

Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm

In this report, we present Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system that jointly advances content consistency, speaker similarity, prosodic naturalness, audio quality, controllability, multilingual coverage, efficiency, and robustness. It combines a 12.5~Hz low-frame-rate speech tokenizer for reduced inference latency with a five-stage progressive training paradigm for coordinated language model (LM) and flow-matching model (FM) optimization. The model provides production-level control through free-style natural-language instructions and fine-grained inline tags, while supporting 16 languages, 20 Chinese dialect regions, one-pass long-form synthesis up to 3 minutes, and robust generation from noisy, reverberant, or unclear reference speech. Across SEED-TTS-Eval, CV3-Eval, instruction-following, long-form, and acoustic-robustness evaluations, Qwen-Audio-3.0-TTS achieves state-of-the-art performance on many reported dimensions or the strongest aggregate results. It also ranks first on the independent Artificial Analysis Text-to-Speech Leaderboard. These results establish Qwen-Audio-3.0-TTS as a strong foundation for production-level speech synthesis.

Bajian Xiang, Cheng Wen, Han Zhao et al. · 3 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.