Nearby in the stack

Improving Query-by-Vocal Imitation with Contrastive Learning and Audio Pretraining · arXivDesk