Speaker: Miguel Fernández Lara.

Abstract: This work presents the development and implementation of a noise-robust end-to-end Automatic Music Transcription (AMT) system for polyphonic piano audio under realistic acoustic conditions. The proposed system employs a sequence-to-sequence Transformer with a convolutional front-end to map mel-spectrograms to tokens representing note onsets, offsets, and velocities, and is trained using an augmentation pipeline that simulates mobile recording conditions through room reverberation, environmental noise, and signal-processing effects. Trained on the MAESTRO dataset, the model achieves significant improvements in transcription quality under various acoustic degradation conditions while maintaining comparable performance on clean audio. Additionally, a client-server mobile application is developed to allow users to record piano performances and automatically obtain MIDI transcriptions and Synthesia-style piano-roll videos, demonstrating the feasibility of deploying robust and practical AMT systems for real-world use.