Nearby in the stack

Language as the Medium: Multimodal Video Classification through text only · arXivDesk