Nearby in the stack

Large Language Models and Multimodal Retrieval for Visual Word Sense Disambiguation · arXivDesk