I am a Ph.D. candidate in Speech Signal Processing at the University of New South Wales (UNSW), Sydney. My research focuses on binaural and spatial hearing for machine listening, with an emphasis on computational auditory modelling and deep neural networks that emulate human auditory processing.
My research interests include:
- Binaural and Spatial Hearing for Machine Listening — spatial cues, binaural modelling, target speaker extraction, sound source localisation, and selective auditory attention.
- Computational Auditory Modelling — biologically inspired audio front-ends that emulate human auditory processing (e.g., adaptive gain control, dynamic range compression).
- Room-Acoustic and Spatial Audio Modelling — estimation of reverberation, DRR, clarity, room volume, and sound orientation from ambisonic recordings.
- Music Information Retrieval — symbolic music understanding, automatic transcription, and generative AI for music.
Education
Ph.D. in Speech Signal Processing
University of New South Wales (UNSW), Sydney
June 2022 – June 2026
Research project: Biologically inspired selective machine hearing for binaural modelling.
Supervisors: Prof. Eliathamby Ambikairajah and A/Prof. Vidhyasaharan Sethu.
Bachelor of Engineering (Honours) in Telecommunications
University of New South Wales (UNSW), Sydney
July 2018 – May 2022
First Class Honours.
Honours thesis: Singing Voice Separation by Embedding Accompaniment Repetitive Features into the Deep Neural Network.
WAM: 85.77/100.
Working Experience
Research Scientist Intern, Incoming
Meta
Cambridge, United Kingdom
June 2026 – December 2026
R&D Engineer Intern
Sony Corporation
Osaki, Tokyo, Japan
June 2025 – July 2025
Mentors: Tomohiro Sawada and Keiichi Taguchi.
Designed a customizable interface for music style transfer to support Sony’s Sound AI business.
Research Intern
Sony Computer Science Laboratories
Tokyo, Japan
January 2025 – June 2025
Mentor: Taketo Akama.
Developed an instrument sound synthesis model using diffusion-based generative models.
Research Intern
Dolby Laboratories
Sydney, Australia
May 2024 – September 2024
Mentors: Jeroen Breebaart and Jeremy Stoddard.
Worked on acoustic context estimation from ambisonic speech recordings, including DRR, T60, C50, room volume, and sound orientation. This work led to an ICASSP 2025 paper and an extended EURASIP journal paper.
Research Assistant
Institute of Software, Chinese Academy of Sciences
Beijing, China
February 2021 – July 2021
Worked on parallel computing and deep learning-based image segmentation for industrial vehicle fault detection.
Publications
Journal Paper
A unified deep learning framework for estimating acoustic context parameters from first order ambisonic speech recordings
H. Meng, J. Breebaart, J. Stoddard, V. Sethu, and E. Ambikairajah
EURASIP Journal on Audio, Speech, and Music Processing, 2026.
DOI: 10.1186/s13636-025-00443-0
Conference Papers
Auto-MatchCut: An audio-visual retrieval framework for seamless match cutting
H. Chen, H. Meng, G. Bhattacharya, L. Lu, J. Kimball, and R. Rossi
IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2026.
Adaptive per-channel energy normalization front-end for robust audio signal processing
H. Meng, V. Sethu, E. Ambikairajah, Q. Zhang, and H. Li
IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2026.
Joint estimation of piano dynamics and metrical structure with a multitask multi-scale CNN
Z. He, H. Meng, D. Huang, and R. Togneri
IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2026.
Blind estimation of sub-band acoustic parameters from ambisonics recordings using spectro-spatial covariance features
H. Meng, J. Breebaart, J. Stoddard, V. Sethu, and E. Ambikairajah
IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2025.
DOI: 10.1109/ICASSP49660.2025.10887842
Binaural selective attention model for target speaker extraction
H. Meng, Q. Zhang, X. Zhang, V. Sethu, and E. Ambikairajah
Proc. Interspeech, 2024, pp. 4323–4327.
DOI: 10.21437/Interspeech.2024-683
Speaking in wavelet domain: A simple and efficient approach to speed up speech diffusion model
X. Zhang, D. Liu, H. Liu, Q. Zhang, H. Meng, L. P. Garcia, E. S. Chng, and L. Yao
Empirical Methods in Natural Language Processing (EMNLP), 2024.
What is learnt by the LEArnable Front-end (LEAF)? Adapting Per-Channel Energy Normalisation (PCEN) to noisy conditions
H. Meng, V. Sethu, and E. Ambikairajah
Proc. Interspeech, 2023, pp. 2898–2902.
DOI: 10.21437/Interspeech.2023-1617
Awards
- The Australasian Speech Science and Technology Association (ASSTA) Conference Travel Awards (2026).
- Asian Deans’ Forum 2025, The Rising Stars Women in Engineering Workshop (November 2025).
- IEEE Signal Processing Society Travel Grant for ICASSP 2025, 2025.
- UNSW 3 Minute Thesis (3MT) Runner-up Prize, School of Electrical Engineering and Telecommunications, 2024.
- University International Postgraduate Award (UIPA), UNSW, 2022.
- Taste of Research Scholarship, UNSW, 2021.
- Dean’s Honour List, UNSW, 2021.
- Dean’s Honour List, UNSW, 2019.
Teaching Experience
Lab Demonstrator
School of Electrical Engineering and Telecommunications, UNSW
February 2021 – December 2024
Demonstrated signal processing courses from first-year to fourth-year undergraduate levels, including ELEC2134, TELE3113, ELEC4622, ELEC3104, and ELEC9123.
Marker
School of Electrical Engineering and Telecommunications, UNSW
May 2022 – December 2024
Marked exams, lab reports, and project deliverables for TELE3113, ELEC3104, and ELEC1111.
Reviewer
Conferences
- ICASSP
- Interspeech
Journals
- IEEE Transactions on Affective Computing (TAC)
CV
You can download my CV here:
Contact
Email: hanyu.meng@unsw.edu.au
Google Scholar: Hanyu Meng
GitHub: Hanyu-Meng
LinkedIn: Hanyu Meng
