Research project

PawCT: End-to-End Choral Transcription into Separate SATB Note Tracks

Hanyu Meng1, Zhanhong He2, Zixun Guo3, Yaolong Ju4

1UNSW · 2UWA · 3Queen Mary University of London · 4Great Bay University

Most choral transcription systems return one merged stream of notes. PawCT identifies which sections are singing and writes a separate MIDI track for soprano, alto, tenor, and bass.

From one choir recording to four usable parts

A mixed choir is transcribed either into one merged note track or into separate soprano, alto, tenor, and bass tracks.
Conventional choral AMT merges every voice. PawCT preserves the SATB structure needed for rehearsal, education, and score reconstruction.

Performance at a glance

0.382 single-track note F1 +61.18% over the previous state of the art
0.225 average SATB note F1 +36.36% over the adapted choral baseline
0.313 zero-shot note F1 on Cantoría 2.72× the two-stage assignment baseline
Single-track transcriptionNote F1 · YouChorale
Previous SOTA0.237
PagCT0.382
Separate SATB transcriptionAverage note F1 · YouChorale
Adapted choral model0.165
Post-hoc assignment0.175
PawCT-OC0.225

Bold: best. Underlined: second best.

Full paper results

Table 1 · Part-agnostic transcription on YouChorale

Model #Params Frame Onset (50 ms) Onset (100 ms)
PRF1 PRF1 PRF1
Onsets and Frames6.79M0.8060.3260.4280.4500.1780.2420.6880.2480.344
MT344.7M0.5900.2430.3440.1170.1480.1270.2000.2550.217
YourMT3+45.8M0.5290.4230.4570.1450.1190.1270.2520.2070.221
MuScriptor-medium307M0.7200.7450.7300.0860.0830.0840.2360.2280.231
Yu et al.4.23M0.6780.6190.6470.2510.2320.2370.3660.3410.347
PagCT w/o Aug.14.84M0.6900.8330.7510.3460.3290.3330.5390.5060.516
PagCT14.84M0.6260.9090.7380.3980.3780.3820.5970.5690.573

Table 2 · Part-aware SATB transcription on YouChorale · Note F1 @ 50 ms

Model SopranoAltoTenorBassAverageVA Rate (%)
FrameNoteFrameNoteFrameNoteFrameNoteFrameNoteFrameNote
YourMT3+★0.3710.0830.2970.0680.2940.0730.3710.0930.3330.07972.8762.20
Yu et al.★0.5280.1830.4510.1510.4360.1380.5600.1860.4940.16576.3569.62
PagCT + Post-VA0.5010.1800.4390.1450.4470.1560.5680.2190.4890.17566.2645.81
PawCT w/o union-loss0.4170.1920.3380.1690.3430.1780.4190.2220.3790.19051.3649.74
PawCT0.5300.2280.4300.1910.4290.1970.5520.2520.4580.21762.0656.81
PawCT-RP0.5420.2100.4530.1850.4580.1940.5850.2470.5100.20969.1154.71
PawCT-RP-OC0.5350.2310.4470.1970.4510.2080.5780.2640.5030.22568.1658.90

Table 3 · Cross-dataset part-aware note F1 @ 50 ms

DatasetModelSopranoAltoTenorBassAverage
CSD
n = 3
PagCT + Post-VA0.0730.1220.1310.1430.118
PawCT0.1400.1750.1830.1450.172
PawCT-RP-OC0.1450.2230.1990.1300.174
Cantoría
n = 14
PagCT + Post-VA0.1790.1150.0850.0840.115
PawCT0.2400.2390.2750.3280.271
PawCT-RP-OC0.3290.2730.3000.3480.313

Model Overview

PawCT detects notes and assigns them to vocal parts in one model. A global training signal keeps the complete pitch content intact, while pitch range and melodic continuity help maintain consistent SATB lines.

PawCT predicts separate SATB note events directly, while the comparison systems produce one merged track or assign parts after transcription.
PawCT predicts SATB notes directly from the mixed recording. The two-stage pipeline is shown only for comparison: it first transcribes a merged track and assigns parts afterwards, and is not part of PawCT.

Listen to the transcription

This 35-second excerpt of Exsultate Deo lets you compare the part-agnostic ground truth with merged predictions from our PagCT model, MuScriptor-medium, and Yu & Duan (2024). PawCT additionally returns four independently playable vocal parts.

Original recording Palestrina · Exsultate Deo · 00:10–00:45
35-second on-page excerpt
Preparing waveform…
0:00 / 0:35
The audio stays on this page—click the waveform to seek. Recording source and attribution

Loading the reference and model outputs…

1

Merged transcription

All annotated or detected notes share one track.

Ground truth Reference · merged note track
F1@50 ms
—
F1@100 ms
—
PagCT Ours Our part-agnostic model · single note track
F1@50 ms
0.433
F1@100 ms
0.668
MuScriptor-medium [1] Rouard et al., 2026
F1@50 ms
0.037
F1@100 ms
0.240
Yu & Duan (2024) [2] ISMIR 2024
F1@50 ms
0.017
F1@100 ms
0.030
2

Separate vocal parts

Play each PawCT output on its own, or compare it with post-hoc assignment.

SopranoS
AltoA
TenorT
BassB
View the full piano-roll comparison
Ground truth, PawCT, PagCT, and PagCT plus Post-VA piano rolls for Exsultate Deo.
Ground truth and model outputs for the 10–100 second passage. Color denotes the assigned SATB part.

References

  1. S. Rouard, M. Krause, A. Roebel, C.-J. Simon-Gabriel, and A. Défossez, “MuScriptor: An Open Model for Multi-Instrument Music Transcription,” arXiv:2607.08168 [cs.SD], 2026. Paper · Code · Model
  2. H. Yu and Z. Duan, “Note-Level Transcription of Choral Music,” in Proceedings of the 25th International Society for Music Information Retrieval Conference (ISMIR), pp. 182–188, 2024. Paper · Code

Error analysis

From pitch activity to note events

MuScriptor-medium and PagCT achieve similar frame F1 (0.730 vs. 0.738), but PagCT obtains substantially higher 50-ms note F1 (0.382 vs. 0.084). This suggests that the main gap lies in localizing note events precisely rather than detecting sustained pitch activity.

Comparison of frame F1 and note F1 for all baselines, highlighting MuScriptor's large frame-to-event drop.
Strong frame-level pitch tracking does not necessarily translate into accurately localized note events.