Machine learning and signal processing for music. Currently focused on detecting and quantifying AI-generated audio in realistic production and broadcast settings.
🎓 Google Scholar · 🆔 ORCID
Abstract
Binary detectors hit >99% accuracy separating fully-AI from fully-human tracks, but producers mix an AI drum loop with a synthetic bassline and a human vocal. We reframe detection as regression over a continuous AI energy ratio: the fraction of a track’s acoustic energy coming from AI stems.
Binary detectors break down on such mixtures; detectability varies sharply by instrument (drums and guitars leak strong neural-codec artifacts, vocals and bass far less); and a regression CNN recovers the ratio at 0.076 MAE, R² = 0.85.
Abstract
Detectors that excel on clean audio degrade sharply in real broadcast conditions. We introduce BAMM, a 40-hour dataset of AI-generated and human-made music sourced from actual television recordings, and evaluate CNN detectors across three tiers of realism: clean foreground music, simulated TV broadcast, and genuine broadcast.
Broadcast-aware training improves robustness, but a substantial gap between laboratory results and practical deployment remains.
Abstract
A deep-learning instrument that generates guitar sounds from vocal commands, extracting timbral descriptors from spoken input to steer a latent space that is otherwise high-dimensional and semantically opaque.
Built within The Sound of AI Open Source Research initiative (200+ contributors), where I served as Project Manager and contributed to the synthesizer implementation.
