EMTDAF: An Explainable Multi-Channel Temporal-Domain Attention Framework for Automated Piano Skill Assessment and Personalized Feedback Generation

Authors

  • Lin Zhang College of Music, Mongolian National University of Arts and Culture, Ulaanbaata,14180, Mongolia & College of Education, Sias University , Zhengzhou, Henan, 450000, China.
  • Zultsetseg Erdenebeleg College of Music, Mongolian National University of Arts and Culture, Ulaanbaata,14180, Mongolia.
  • Tsevegsuren Tserenbaljir College of Music, Mongolian National University of Arts and Culture, Ulaanbaata,14180, Mongolia.

DOI:

https://doi.org/10.56979/1102/2026/1631

Keywords:

Automated Music Performance Assessment, Deep Learning, Attention Mechanism, Explainable Artificial Intelligence, Mel-Spectrogram, MFCC, Intelligent Tutoring Systems

Abstract

Automated assessment of musical performance has long been a challenging problem in music information retrieval (MIR) and intelligent tutoring systems. Existing approaches predominantly rely on hand-crafted acoustic features or unimodal deep learning models that lack interpretability and fail to provide actionable pedagogical feedback. In this paper, we present EMTDAF the Explainable Multi-Channel Temporal-Domain Attention Framework a novel end-to-end deep learning system for automated piano skill assessment that simultaneously addresses accuracy, ordinal consistency, and explainability. EMTDAF fuses three complementary audio representations Mel-spectrograms, Mel-Frequency Cepstral Coefficients (MFCCs), and Chroma features through a dual-branch Frequency-Temporal Attention (FTA) module, and optimises a combined ordinal cross-entropy loss to preserve the natural ordering of skill levels. The framework is evaluated on the PISA (Piano Skills Assessment) benchmark dataset comprising 65 recordings across 10 player levels. EMTDAF achieves 80.0% exact-match accuracy and a Mean Absolute Error (MAE) of 0.50 levels, surpassing the best published PISA baseline (PISA-MMDL Uniform, 74.6%) by 5.4 percentage points. Beyond classification, EMTDAF integrates a three-tier XAI pipeline comprising GradCAM, SHAP, and attention map visualisation, together with a skill-gap radar chart system that delivers per-student diagnostic feedback across six musical dimensions: Tempo Control, Dynamic Range, Note Clarity, Rhythmic Accuracy, Harmonic Complexity, and Technical Fluency. An extensive ablation study confirms the contribution of each architectural component, with the Frequency-Temporal Attention module providing the largest single improvement of 6.5 percentage points. These results establish EMTDAF as a state-of-the-art, interpretable, and educationally actionable framework for automated music performance assessment.

Downloads

Published

2026-09-01

How to Cite

Lin Zhang, Zultsetseg Erdenebeleg, & Tsevegsuren Tserenbaljir. (2026). EMTDAF: An Explainable Multi-Channel Temporal-Domain Attention Framework for Automated Piano Skill Assessment and Personalized Feedback Generation. Journal of Computing & Biomedical Informatics, 11(02). https://doi.org/10.56979/1102/2026/1631