Skip to content

Jun 19, 20261 min read

What 12.9M audio samples taught me about data quality

The model was never the bottleneck. The labels were.

A grid of audio waveforms

We kept tuning the model. The wins came from fixing the data.

Where the errors hid

Most of the label noise clustered in a few recording conditions. Fixing those beat any architecture change we tried that quarter.

Read this next