CLASS-CONDITIONAL INTERMEDIATE-DOMAIN ADVERSARIAL ADAPTATION FOR CROSS-DEVICE ACOUSTIC SCENE CLASSIFICATION
Keywords:
Acoustic scene classification, Unsupervised domain adaptation, Cross-device generalization, Intermediate domains, Conditional adversarial learning, Maximum mean discrepancyAbstract
Cross-device acoustic scene classification (ASC) is difficult because microphone frequency response, noise floor, nonlinear compression, and device–scene interactions can move recordings across class boundaries. Global adversarial alignment is especially fragile when the domain gap is large: it can reduce marginal discrepancy while mixing semantically different modes. This paper proposes class-conditional intermediate-domain adversarial adaptation (CC-IDA), a lightweight unsupervised adaptation framework that decomposes a large device shift into a physically interpretable source-to-bridge-to-target path. Two label-preserving bridge devices are generated from source recordings through progressively stronger transfer functions. A shared convolutional encoder is trained with source and bridge classification, a global maximum mean discrepancy anchor, a four-domain conditional discriminator acting on the multilinear feature–prediction representation, a class-centroid bridge chain, confident target prototype alignment, and target entropy regularization. To ensure complete reproducibility and avoid unsupported benchmark claims, experiments use a controlled procedural proxy benchmark rather than reporting results on TAU/DCASE: six acoustic-scene waveforms are rendered through two source, two intermediate, and one unseen target device; target latent recordings are independent of source/bridge recordings. Across three deterministic seeds, CC-IDA achieves 89.78% ± 8.95% target accuracy and 87.18% ± 11.32% macro-F1, compared with 42.44% ± 15.64% and 31.78% ± 16.04% for source-only training. It also improves over unconditional intermediate-domain adaptation by 1.78 percentage points in accuracy and 2.87 points in macro-F1, with no additional inference parameters. The results support gradual, class-aware alignment while also exposing limitations: only three seeds and a controlled proxy are evaluated, so real-corpus validation remains necessary.References
[1] Stowell D, Giannoulis D, Benetos E, et al. Detection and Classification of Acoustic Scenes and Events. IEEE Transactions on Multimedia, 2015, 17(10): 1733-1746. DOI: 10.1109/TMM.2015.2428998.
[2] Piczak K J. ESC: Dataset for Environmental Sound Classification. Proceedings of the 23rd ACM International Conference on Multimedia, 2015: 1015-1018. DOI: 10.1145/2733373.2806390.
[3] Salamon J, Bello J P. Deep Convolutional Neural Networks and Data Augmentation for Environmental Sound Classification. IEEE Signal Processing Letters, 2017, 24(3): 279-283. DOI: 10.1109/LSP.2017.2657381.
[4] Kong Q, Cao Y, Iqbal T, et al. PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2020, 28: 2880-2894. DOI: 10.1109/TASLP.2020.3030497.
[5] Gemmeke J F, Ellis D P W, Freedman D, et al. Audio Set: An Ontology and Human-Labeled Dataset for Audio Events. Proceedings of IEEE ICASSP, 2017: 776-780. DOI: 10.1109/ICASSP.2017.7952261.
[6] Gong Y, Chung Y-A, Glass J. AST: Audio Spectrogram Transformer. Proceedings of Interspeech, 2021: 571-575.
[7] Mesaros A, Heittola T, Virtanen T. A Multi-Device Dataset for Urban Acoustic Scene Classification. Proceedings of the DCASE 2018 Workshop, 2018.
[8] Heittola T, Mesaros A, Virtanen T. Acoustic Scene Classification in DCASE 2020 Challenge: Generalization Across Devices and Low Complexity Solutions. Proceedings of the DCASE 2020 Workshop, 2020: 56-60.
[9] Mesaros A, Heittola T, Virtanen T. TAU Urban Acoustic Scenes 2020 Mobile, Development Dataset. Zenodo, Dataset record, 2020. DOI: 10.5281/zenodo.3670167.
[10] Martín-Morató I, Heittola T, Mesaros A, et al. Low-Complexity Acoustic Scene Classification for Multi-Device Audio: Analysis of DCASE 2021 Challenge Systems. Proceedings of the DCASE 2021 Workshop, 2021.
[11] Martín-Morató I, Paissan F, Ancilotto A, et al. Low-Complexity Acoustic Scene Classification in DCASE 2022 Challenge. Proceedings of the DCASE 2022 Workshop, 2022: 111-115. DOI: 10.48550/arXiv.2206.03835.
[12] Park D S, Chan W, Zhang Y, et al. SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition. Proceedings of Interspeech, 2019: 2613-2617. DOI: 10.21437/Interspeech.2019-2680.
[13] Ben-David S, Blitzer J, Crammer K, et al. A Theory of Learning from Different Domains. Machine Learning, 2010, 79: 151-175. DOI: 10.1007/s10994-009-5152-4.
[14] Gretton A, Borgwardt K M, Rasch M J, et al. A Kernel Two-Sample Test. Journal of Machine Learning Research, 2012, 13(25): 723-773.
[15] Long M, Cao Y, Wang J, et al. Learning Transferable Features with Deep Adaptation Networks. Proceedings of ICML, PMLR, 2015: 97-105.
[16] Ganin Y, Ustinova E, Ajakan H, et al. Domain-Adversarial Training of Neural Networks. Journal of Machine Learning Research, 2016, 17(59): 1-35.
[17] Sun B, Saenko K. Deep CORAL: Correlation Alignment for Deep Domain Adaptation. ECCV Workshops, 2016: 443-450. DOI: 10.1007/978-3-319-49409-8_35.
[18] Long M, Cao Z, Wang J, et al. Conditional Adversarial Domain Adaptation. Advances in Neural Information Processing Systems, 2018: 1640-1650.
[19] Long M, Zhu H, Wang J, et al. Deep Transfer Learning with Joint Adaptation Networks. Proceedings of ICML, PMLR, 2017: 2208-2217.
[20] Tzeng E, Hoffman J, Saenko K, et al. Adversarial Discriminative Domain Adaptation. Proceedings of IEEE CVPR, 2017: 7167-7176.
[21] Saito K, Watanabe K, Ushiku Y, et al. Maximum Classifier Discrepancy for Unsupervised Domain Adaptation. Proceedings of IEEE CVPR, 2018: 3723-3732.
[22] Zhang Y, Liu T, Long M, et al. Bridging Theory and Algorithm for Domain Adaptation. Proceedings of ICML, PMLR, 2019: 7404-7413.
[23] Kumar A, Sattigeri P, Wadhawan K, et al. Co-Regularized Alignment for Unsupervised Domain Adaptation. Advances in Neural Information Processing Systems, 2018: 9345-9356.
[24] Chen X, Wang S, Long M, et al. Transferability vs. Discriminability: Batch Spectral Penalization for Adversarial Domain Adaptation. Proceedings of ICML, PMLR, 2019: 1081-1090.
[25] Kumar A, Ma T, Liang P. Understanding Self-Training for Gradual Domain Adaptation. Proceedings of ICML, PMLR, 2020: 5468-5479.
[26] Wang H, Li B, Zhao H. Understanding Gradual Domain Adaptation: Improved Analysis, Optimal Path and Beyond. Proceedings of ICML, PMLR, 2022: 22784-22801.
[27] Chen L, Wei Z, Jin X, et al. Deliberated Domain Bridging for Domain Adaptive Semantic Segmentation. Advances in Neural Information Processing Systems, 2022: 15105-15118.
[28] Ioffe S, Szegedy C. Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. Proceedings of ICML, PMLR 37, 2015: 448-456.
[29] Kingma D P, Ba J. Adam: A Method for Stochastic Optimization. International Conference on Learning Representations, 2015.