Patch-Mix Contrastive Learning with Audio Spectrogram Transformer on Respiratory Sound Classification (INTERSPEECH 2023)
-
Updated
Mar 11, 2025 - Python
Patch-Mix Contrastive Learning with Audio Spectrogram Transformer on Respiratory Sound Classification (INTERSPEECH 2023)
This is the official repository of the papers "Parameter-Efficient Transfer Learning of Audio Spectrogram Transformers" [IEEE MLSP 2024] and "Efficient Fine-tuning of Audio Spectrogram Transformers via Soft Mixture of Adapters" [Interspeech 2024].
Pytorch implementation of INTEGRATED PARAMETER-EFFICIENT TUNING FOR GENERAL-PURPOSE AUDIO MODELS
Three in one. capture audio with mobile mic and then play, play songs from internal download directly, displaying spectogram (frequency) of an audio in Flutter.
Two-branch audio authenticity app that detects AI-generated speech and environmental audio using WavLM, AST, FastAPI, and React.
GenreAST is a highly accurate music genre classification model relying on the AST model
A full-stack multimodal music agent for audio analysis, playlist planning, and LLM-based music recommendation.
AI-powered deepfake audio detection tool built with Kotlin to analyze and verify voice authenticity in real time.
Audio Spectrogram Transformer with LoRA adapter.
Code accompanying ESANN 2025 submission "Exploring Model Architectures for Real-Time Lung Sound Event Detection". Dataset used was ICBHI 2017.
High-performance multimodal pipeline synchronizing Meta's SAM 2 (Vision) and MIT's AST (Audio). Features O(1) resource management, temporal video slicing, and automated JSON metadata orchestration for edge-device AI.
汎用音楽分析 CLI: BPM・拍子・コード・調性・音楽理論・ジャンル (madmom + librosa + music21 + AST/AudioSet)
Add a description, image, and links to the audio-spectrogram-transformer topic page so that developers can more easily learn about it.
To associate your repository with the audio-spectrogram-transformer topic, visit your repo's landing page and select "manage topics."