AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

From Masking to Merging: Rethinking SpecAugment for Efficient Audio Spectrogram Transformer

arXiv · AI, language, vision and robotics · article · Sep 6, 2026 · UTC

This paper proposes SpecAugment-Patch Merging, a simple yet effective method to accelerate Audio Spectrogram Transformer (AST) training. We first apply SpecAugment to mask input spectrograms at the patch level, and after positional embeddings are added, the method selects r pairs of masked patches and merges them, reducing the number of tokens processed by the Transformer. Increasing the number of merged pairs r from 0 to 100 keeps mAP on AudioSet nearly unchanged (34.07 to 34.08) while throughput increases from 43.3 to 49.3 samples/sec, which is a relatively 13.9% improvement. Similar pattern

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T21:12:06.801Z. This is not the publication date.