Nearby in the stack

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds · arXivDesk