- 论文公开站arXiv
WorldSonus: Bringing Sound to Worlds
Recent advances in world models have enabled increasingly realistic visual synthesis. However, these generated environments remain largely silent. Bringing sound to world models poses three core challenges: real-time gen…
- 论文公开站arXiv
DepthWorld: 3D World Model for Robot Manipulation
World models offer a data-driven alternative to traditional simulators for robotics, with applications spanning policy evaluation, improvement, and planning. All of these uses depend on faithful 3D geometry, yet current …
- 论文公开站arXiv
EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras
Inspired by human vision, we introduce a framework using active gaze to enable fine-grained bimanual manipulation with only a single stereo camera. EyeRobot 2.0 physically attends to a 3D fixation point in the scene by s…
- 论文公开站arXiv
Less Decoder is More Encoder: Geometric Representation Learning from Novel View Synthesis
This paper examines the role of Novel View Synthesis (NVS) in geometric representation learning. In principle, NVS should reason about 3D scene structure, thereby enabling transferable multi-view geometric representation…
- 论文公开站arXiv
ViTeX-Bench:高保真视频场景文本编辑基准
ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing
视频生成日益逼真可控,但视频编辑仍欠成熟,尤其是需保持原始场景动态的精确局部编辑。视频场景文本编辑需替换店面招牌等场景表面上的文字。
- 论文公开站arXiv
DexRoam:从第一人称全身人类演示学习移动双臂灵巧操作
DexRoam: Learning Mobile Bimanual Dexterous Manipulation from Egocentric Whole-Body Human Demonstrations
移动双臂灵巧操作需在单条轨迹中协调运动、全身动作与手指级灵巧性,造成严重的机器人演示瓶颈;第一人称人类演示提供了可扩展的替代方案。
- 论文公开站arXiv
适应AI:小学教师如何调整教学实践以融入AI课程
Adapting for AI: How elementary teachers adjust their practices for an AI-integrated curriculum
摘要显示,该研究追踪三名小学教师在为期13天、三周夏令营中实施围绕ToyTalk会话AI玩具开发平台构建的AI素养与英语语言艺术课程。基于每日个人反思、小组反思及营后访谈,研究发现教师的修复、差异化、翻译与平衡等适应性实践处于技术、学习者与教学三重张力交汇处,且教师对AI及自身角色的理解随营期经历演变。研究据此提出在小学课堂部署会话AI的设计启示与考量。
意义:为教育AI产品设计者与课程开发者提供一线教师适应性实践的实证依据,有助于让会话AI工具更贴合小学课堂真实教学情境。
- 论文公开站arXiv
OC-GS:面向不规则转台拍摄的高斯泼溅重建
OC-GS: Gaussian Splatting for Irregular Turntable Capture
摘要显示,针对转台拍摄中旋转不均与丢帧导致等角度假设失效的问题,OC-GS 提出以物体为中心的高斯泼溅方法,在共享相机、旋转轴与枢轴的前提下细化每张图像的拍摄角度,联合优化图像几何与角度。在 12、8、6 个不规则视角下,前景 PSNR 分别为 21.26、19.36、15.83dB,均优于四个无位姿高斯泼溅基线;共享训练器下细化角度比固定估计提升 7.88dB,真实拍摄提升 0.70dB。
意义:为稀疏、不规则转台拍摄提供无需精确位姿的三维重建思路,对低成本物体扫描与高斯泼溅应用有参考价值。
- 论文公开站arXiv
WorldCrafter:具有隐式3D感知记忆的一致视频世界模型
WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory
摘要显示,WorldCrafter 是一种视频世界模型,学习可被相机查询的隐式3D感知记忆,让请求视角决定多视角证据如何压缩进视频生成器有限的 token 预算。记忆编码器与姿态条件读出模块同视频生成器联合训练,在去噪前将历史观测整合为固定数量的目标视角专属 token,无需显式深度对应。结合近期时序上下文与少步蒸馏,该模型支持从单张图像或文本提示进行流式场景探索,在静态与动态场景中提升了长时程一致性与相机控制精度。
意义:为长时程、跨视角一致的交互式视频世界模型提供隐式3D记忆方案,利于场景探索与相机控制类应用。