VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning Paper • 2608.26105 • Published 7 days ago • 268
Mage Collection A family of lightweight multimodal models, including understanding and generation. • 8 items • Updated Jul 26 • 29
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published Jul 27 • 37
view article Article Welcome NVIDIA Cosmos 3: The First Open Omni-model for Physical AI Reasoning and Action nvidia • Jun 1 • 89
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Paper • 2506.05414 • Published Jun 4, 2025 • 4