Nearby in the stack

Cross-modal Attention Congruence Regularization for Vision-Language Relation Alignment · arXivDesk