The notes were all there. They were in the wrong place.
StaveWave turns sheet music into notation you can edit and play back. The field is called optical music recognition, or OMR (1).
This post is about the easiest input we see: scores that were typeset in notation software and exported straight to PDF. The staff lines are perfectly straight, the symbols are crisp, and every note is drawn exactly where the composer's software put it. If a reader is going to be near-perfect anywhere, it should be on these pages. We still make mistakes on them, and we wanted to know which kind.
How we score a transcription
We grade our output with an edit distance. The idea comes from text comparison (4): count the smallest number of edits that would turn one sequence into another. For sheet music, the question becomes: how many notes, rests, bars and markings would someone have to add or remove to turn our transcription into a correct reference transcription of the same score? Divide that count by the size of the score and you get a number from 0 (identical) to 1 (nothing in common).
The metric we use is OMR-NED (optical music recognition normalized edit distance). It was proposed in 2025 as a standard way to compare OMR systems (2), and it is computed with musicdiff, an open-source tool that lines up two scores bar by bar and lists every difference between them (3, 5). Using a shared, published metric means our numbers mean the same thing as everyone else's.
One property of this metric turned out to matter a great deal. In text edit distance, fixing a typo counts as one "substitution": swap the wrong letter for the right one. This metric has no substitution for notes. A note that is wrong in any way, whether its pitch, its length or the moment it starts, is counted as two edits: delete the wrong note, insert the right one. That is exactly what a note we missed plus a note we invented costs. So from the headline number alone, "we didn't see this note" and "we saw it and wrote it in the wrong place" look the same, even though they need completely different fixes.
What we expected, and what we found
Before measuring, we assumed most of the remaining error was missing music (notes, bars or whole lines we failed to pick up): at least 40% of it.
It was about a quarter. The largest single share, about 30%, was notes that were right in every way a musician would check first. They had the correct pitch, the correct length, and they were in the correct bar. They just started at the wrong beat inside that bar.
We never dropped a whole staff. Whole missing bars were only about 7% of the error, and wrong note lengths barely registered.
How we separated "missed" from "misplaced"
Each time the scorer charges an edit, our scoring tool records which note was deleted or inserted, and in which bar. Afterwards it pairs each deleted note with an inserted note from the same bar:
- Same pitch and length, different start: misplaced
- Same start and length, different pitch: wrong pitch
- Same pitch and start, different length: wrong length
- No partner in the bar: missing or extra
Then it checks its own work. For every score, the edits across all the classes must add up exactly to the scorer's own total, with nothing left over and nothing counted twice. If a single score fails that check, the whole run is thrown out. Every score passed.
Why the notes landed late
A staff often carries more than one voice at once, such as a melody over a held bass note. In MusicXML they appear one after another. After the first voice, a backup instruction rewinds the clock to the start of the bar so the next voice can begin on beat 1. There is also a forward instruction, which moves the clock ahead through silent, invisible time, like a rest nobody sees (6).
Our output was splitting music into more voices than the score really has, and it was starting those extra voices with invisible padding instead of rewinding to beat 1. Every note in a padded voice arrived late by the same amount. In one small worked example, all of a bar's notes were correct, but one group came in 1¼ beats late and the next group 2½ beats late. A constant, stacking delay like that is the fingerprint of padding, not of misreading. About seven in ten of our output files used this padding. None of the reference transcriptions did.
Two traps we nearly fell into
Trap 1: pairing notes without their bar. Inside the scorer, a note's position is counted from the start of its own bar, so "beat 2" exists in every bar of the piece. Our first analysis paired deleted and inserted notes without checking which bar they came from, so it would happily match a note in bar 3 with one in bar 17. The result looked perfectly reasonable: misplaced about 26%, wrong pitch about 22%. Adding the bar number moved those to 30% and 12%. The first answer was wrong in both directions, and nothing about it looked wrong. We only caught it by looking at what the position number actually meant.
Trap 2: measuring under settings production doesn't use. Afterwards we checked how much of the misplaced error a candidate correction step could even reach. The correction can only run on pages that meet certain conditions, so we counted the pages that qualified. The first count said almost none, a fraction of a percent of the problem, and we set the idea aside. A later check found that the count had run with one configuration setting different from production. Rerun with production's settings, the reachable share was an order of magnitude larger. It was still small, and the decision to hold off didn't change, but the reason we had written down was wrong. Every count like this now runs with production's exact configuration, and records it.
What changes for us
The headline score pointed us at recognition: get better at seeing notes. The breakdown points somewhere else. On clean, typeset scores, the biggest single lever is how we write voices out once the notes are already found. That's a different part of the system, and a different kind of work.
We're also keeping the tool. Whenever the headline number moves from now on, we can say which kind of error moved with it.
References
- J. Calvo-Zaragoza, J. Hajič Jr. and A. Pacha. "Understanding Optical Music Recognition." ACM Computing Surveys 53(4), 2020. doi:10.1145/3397499
- J. C. Martinez-Sevilla, J. Cerveto-Serrano, N. Luna, G. Chapman, C. Sapp, D. Rizo and J. Calvo-Zaragoza. "Sheet Music Benchmark: Standardized Optical Music Recognition Evaluation." ISMIR 2025.
- F. Foscarin, F. Jacquemard and R. Fournier-S'niehotta. "A diff procedure for music score files." Digital Libraries for Musicology (DLfM) 2019.
- V. I. Levenshtein. "Binary codes capable of correcting deletions, insertions, and reversals." Soviet Physics Doklady 10(8), 1966. (overview)
- G. Chapman. musicdiff, an open-source music notation diff tool.
- W3C Music Notation Community Group. MusicXML 4.0 reference: forward and backup.