站内快照 · 国内可打开。外网原文可能无法访问。
- 论文公开站arXiv
VISTA: A Visual Harness for Reasoning in an Interactive World
We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vis