Welcome to Journal of Beijing Institute of Technology
Yun Wu, Xinting Li, Zhimin Cao. Beyond Baseline MP-SENet: A Synergistic Optimization Framework for High-Fidelity Monaural Speech EnhancementJ. JOURNAL OF BEIJING INSTITUTE OF TECHNOLOGY, 2026, 35(4): 485-498. DOI: 10.15918/j.jbit1004-0579.2025.052
Citation: Yun Wu, Xinting Li, Zhimin Cao. Beyond Baseline MP-SENet: A Synergistic Optimization Framework for High-Fidelity Monaural Speech EnhancementJ. JOURNAL OF BEIJING INSTITUTE OF TECHNOLOGY, 2026, 35(4): 485-498. DOI: 10.15918/j.jbit1004-0579.2025.052

Beyond Baseline MP-SENet: A Synergistic Optimization Framework for High-Fidelity Monaural Speech Enhancement

  • In distributed speech front-end frameworks, multi-channel enhancement typically relies on inter-channel time synchronization and is therefore susceptible to delay mismatch, which may lead to performance degradation. Based on this, this work regards independent single-channel enhancement as a complementary solution: when time synchronization information becomes unreliable, it can still maintain relatively stable performance. In single-channel research, noisy scenarios hinder accurate retrieval of non-structured, envelope-bound phase; without explicit constraints, amplitude-phase imbalance disrupts harmonics and continuity, making explicit modeling pivotal. MP-SENet (mask phase aware speech enhancement network) significantly advanced parallel amplitude-phase modeling for simultaneous optimization but struggles with residual noise and phase adaptation in low-SNR environments. To address this, an enhanced MP-SENet is proposed: FRI (filter recycle interguide) in TF-Transformer disentangles global-local speech-noise features; a lightweight SR post-processing module refines the amplitude–phase estimation, curbing noise and artifacts; parallel phase and metric discriminators enable joint optimization to enhance quality and intelligibility. It achieves state-of-the-art results on VoiceBank+DEMAND (PESQ: 3.67, CSIG: 4.96, STOI: 0.97) and robust generalization on real-world WHISPER_SET_1 at 0, –5, –10 dB, which makes it promising for distributed front-end systems.
  • loading

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return