Nearby in the stack

Compositional Context Fine-Tuning Vision-Language Model for Complex Assembly Action Understanding from Videos · arXivDesk