I am a Ph.D. candidate in Speech Signal Processing at the University of New South Wales (UNSW), Sydney. My research focuses on binaural and spatial hearing for machine listening, with an emphasis on computational auditory modelling and deep neural networks that emulate human auditory processing.

My research interests include:

  • Binaural and Spatial Hearing for Machine Listening — spatial cues, binaural modelling, target speaker extraction, sound source localisation, and selective auditory attention.
  • Computational Auditory Modelling — biologically inspired audio front-ends that emulate human auditory processing (e.g., adaptive gain control, dynamic range compression).
  • Room-Acoustic and Spatial Audio Modelling — estimation of reverberation, DRR, clarity, room volume, and sound orientation from ambisonic recordings.
  • Music Information Retrieval — symbolic music understanding, automatic transcription, and generative AI for music.

Education

Ph.D. in Speech Signal Processing

University of New South Wales (UNSW), Sydney
June 2022 – June 2026

Research project: Biologically inspired selective machine hearing for binaural modelling.
Supervisors: Prof. Eliathamby Ambikairajah and A/Prof. Vidhyasaharan Sethu.

Bachelor of Engineering (Honours) in Telecommunications

University of New South Wales (UNSW), Sydney
July 2018 – May 2022

First Class Honours.
Honours thesis: Singing Voice Separation by Embedding Accompaniment Repetitive Features into the Deep Neural Network.
WAM: 85.77/100.


Working Experience

Research Scientist Intern, Incoming

Meta
Cambridge, United Kingdom
June 2026 – December 2026

R&D Engineer Intern

Sony Corporation
Osaki, Tokyo, Japan
June 2025 – July 2025

Mentors: Tomohiro Sawada and Keiichi Taguchi.
Designed a customizable interface for music style transfer to support Sony’s Sound AI business.

Research Intern

Sony Computer Science Laboratories
Tokyo, Japan
January 2025 – June 2025

Mentor: Taketo Akama.
Developed an instrument sound synthesis model using diffusion-based generative models.

Research Intern

Dolby Laboratories
Sydney, Australia
May 2024 – September 2024

Mentors: Jeroen Breebaart and Jeremy Stoddard.
Worked on acoustic context estimation from ambisonic speech recordings, including DRR, T60, C50, room volume, and sound orientation. This work led to an ICASSP 2025 paper and an extended EURASIP journal paper.

Research Assistant

Institute of Software, Chinese Academy of Sciences
Beijing, China
February 2021 – July 2021

Worked on parallel computing and deep learning-based image segmentation for industrial vehicle fault detection.


Publications

Journal Paper

A unified deep learning framework for estimating acoustic context parameters from first order ambisonic speech recordings
H. Meng, J. Breebaart, J. Stoddard, V. Sethu, and E. Ambikairajah
EURASIP Journal on Audio, Speech, and Music Processing, 2026.
DOI: 10.1186/s13636-025-00443-0

Conference Papers

Auto-MatchCut: An audio-visual retrieval framework for seamless match cutting
H. Chen, H. Meng, G. Bhattacharya, L. Lu, J. Kimball, and R. Rossi
IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2026.

Adaptive per-channel energy normalization front-end for robust audio signal processing
H. Meng, V. Sethu, E. Ambikairajah, Q. Zhang, and H. Li
IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2026.

Joint estimation of piano dynamics and metrical structure with a multitask multi-scale CNN
Z. He, H. Meng, D. Huang, and R. Togneri
IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2026.

Blind estimation of sub-band acoustic parameters from ambisonics recordings using spectro-spatial covariance features
H. Meng, J. Breebaart, J. Stoddard, V. Sethu, and E. Ambikairajah
IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2025.
DOI: 10.1109/ICASSP49660.2025.10887842

Binaural selective attention model for target speaker extraction
H. Meng, Q. Zhang, X. Zhang, V. Sethu, and E. Ambikairajah
Proc. Interspeech, 2024, pp. 4323–4327.
DOI: 10.21437/Interspeech.2024-683

Speaking in wavelet domain: A simple and efficient approach to speed up speech diffusion model
X. Zhang, D. Liu, H. Liu, Q. Zhang, H. Meng, L. P. Garcia, E. S. Chng, and L. Yao
Empirical Methods in Natural Language Processing (EMNLP), 2024.

What is learnt by the LEArnable Front-end (LEAF)? Adapting Per-Channel Energy Normalisation (PCEN) to noisy conditions
H. Meng, V. Sethu, and E. Ambikairajah
Proc. Interspeech, 2023, pp. 2898–2902.
DOI: 10.21437/Interspeech.2023-1617


Awards

  • The Australasian Speech Science and Technology Association (ASSTA) Conference Travel Awards (2026).
  • Asian Deans’ Forum 2025, The Rising Stars Women in Engineering Workshop (November 2025).
  • IEEE Signal Processing Society Travel Grant for ICASSP 2025, 2025.
  • UNSW 3 Minute Thesis (3MT) Runner-up Prize, School of Electrical Engineering and Telecommunications, 2024.
  • University International Postgraduate Award (UIPA), UNSW, 2022.
  • Taste of Research Scholarship, UNSW, 2021.
  • Dean’s Honour List, UNSW, 2021.
  • Dean’s Honour List, UNSW, 2019.

Teaching Experience

Lab Demonstrator

School of Electrical Engineering and Telecommunications, UNSW
February 2021 – December 2024

Demonstrated signal processing courses from first-year to fourth-year undergraduate levels, including ELEC2134, TELE3113, ELEC4622, ELEC3104, and ELEC9123.

Marker

School of Electrical Engineering and Telecommunications, UNSW
May 2022 – December 2024

Marked exams, lab reports, and project deliverables for TELE3113, ELEC3104, and ELEC1111.


Reviewer

Conferences

  • ICASSP
  • Interspeech

Journals

  • IEEE Transactions on Affective Computing (TAC)

CV

You can download my CV here:

Download CV


Contact

Email: hanyu.meng@unsw.edu.au

Google Scholar: Hanyu Meng

GitHub: Hanyu-Meng

LinkedIn: Hanyu Meng