- 论文公开站arXiv
InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation
We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills, improving from its own attempts, and retaining what it learns, withou…
- 论文公开站arXiv
FERPO: Forward Entropy-Regularized Policy Optimization
Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurat…
- 论文公开站arXiv
TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning
Full-parameter fine-tuning of large language models (LLMs) incurs substantial optimizer state memory overhead, limiting the model sizes that fit on modern GPUs. Existing approaches either compress optimizer state, abando…
- 论文公开站arXiv
VISTA: A Visual Harness for Reasoning in an Interactive World
We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness th…
- 论文公开站arXiv
SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation
High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict active structure and then synthesize local geometry. While effective, this design fragments continuous surfa…
- 论文公开站arXiv
ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of research, a new problem…
- 论文公开站arXiv
Embedding Prediction Helps Image Generation
In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embeddin…
- 论文公开站arXiv
Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents
Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real…
- 论文公开站arXiv
KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agen…
- 论文公开站arXiv
One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars
3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bottleneck and show that the animation of pretrained avatar models can be…
- 论文公开站arXiv
Phase-Field Modeling of Liquid Phases with Ordering
A strong negative enthalpy of mixing can lead to ordering in liquid solutions. To describe ordering in liquids, different thermodynamic descriptions of liquid phases have been developed in the CALPHAD approach, including…
- 论文公开站arXiv
Compression Footprints as Security Signals for Model-Poisoning Defense in Federated Learning
Lossy compression is widely used in Federated Learning (FL) but is generally treated as an error source, while conventional poisoning defenses inspect update geometry. In this work, we instead treat the compressor's resp…
- 论文公开站arXiv
Scaling Laws for Looped Mixture of Experts
Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed parameters, while MoE sparsity expands total capacity at fixed active co…
- 论文公开站arXiv
MatLoom: Layered Text-to-Material Generation in a Compact Program Space
Material generation should produce not only an appearance, but also the rules that construct it. We introduce MatLoom, a compact, layer-oriented language for text-to-material generation with pretrained language models. E…
- 论文公开站arXiv
Cogentic:面向自动证明发现的多智能体编排
Cogentic: Multi-Agent Orchestration for Automated Proof Discovery
提出 Cogentic,一个用于开放研究问题自动证明发现的多智能体框架。前沿语言模型虽能单次生成优质数学想法,但面对需多路径探索的开放问题,单次生成往往不足。
- 论文公开站arXiv
WorldAuditBench:多模态智能体的交互式3D世界审计
WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents
随着交互式3D世界被广泛用于研究智能行为,识别其中异常(如悬浮物体、可穿墙、与环境不一致的物体)的高效流程变得重要,本文提出相应基准。
- 论文公开站arXiv
Turbo Harness:实例自适应框架优化
Turbo Harness: Instance-Adaptive Harness Optimization
自动化搜索有效框架是实现智能体递归自我改进的重要一步。现有优化通常产出统一应用于所有任务的单一全局框架,但适用于某类任务的框架未必通用。
- 论文公开站arXiv
DP-SGD 下权重绑定对纯解码器 LLM 是否仍有益?
Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?
DP-SGD 是大模型隐私保护微调的主流方法。许多纯解码器 LLM 采用输入输出嵌入权重绑定,本文在差分隐私设定下重新审视这一设计是否仍然有利。
- 论文公开站arXiv
Ego4WAM:扩展第一人称人类数据用于机器人学习的关键因素
Ego4WAM: What Matters When Scaling Egocentric Human Data for Robot Learning?
第一人称人类数据为机器人学习提供可扩展经验,但在人机对齐、行为覆盖和监督信号上差异显著。已有工作显示数据越多效果越好,但哪些数据属性真正关键仍不明确。
- 论文公开站arXiv
图像分类器是高效的自监督视频表征学习器
Image Classifiers are Efficient Self-Supervised Video Representation Learners
提出 VideoMSN,一种用于视频高效自监督时空表征学习的掩码孪生网络框架。它不依赖重型3D架构或重建式自编码器,而是复用标准图像分类器处理无标注数据。
- 论文公开站arXiv
AssemblyWorld:用通用智能体重思3D装配
AssemblyWorld: Rethinking 3D Assembly with General-Purpose Agents
3D装配需将零件及其关系的理解转化为精确空间排布。本文探究预训练通用智能体能否仅通过视觉交互完成物体装配,而无需针对装配任务额外微调。
- 论文公开站arXiv
ViTeX-Bench:高保真视频场景文本编辑基准
ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing
视频生成日益逼真可控,但视频编辑仍欠成熟,尤其是需保持原始场景动态的精确局部编辑。视频场景文本编辑需替换店面招牌等场景表面上的文字。
- 论文公开站arXiv
去除时间捷径可改进非侵入式脑机文本解码
Removing Timing Shortcuts Improves Non-Invasive Brain-to-Text
研究发现,多项报道的非侵入式脑信号词解码性能提升,在完全不用脑数据的情况下也能复现。在 d'Ascoli 等人(2025)的工作中,受试者感知连续语音的脑活动时间序列被分段处理。
- 论文公开站arXiv
半事实信用增强策略优化
Semifactual Credit-Augmented Policy Optimization
可验证奖励强化学习(RLVR)提升了 LLM 推理能力,但其预测仍对任务无关的提示特征敏感。本文通过半事实提示干预研究这种敏感性。
- 论文公开站arXiv
面向多模态临床诊断的排序感知提示优化
Ranking-Aware Prompt Optimization for Multimodal Clinical Diagnosis
多模态大模型正快速推进临床诊断,但其适配流程仍以准确率为目标。临床数据类别严重不平衡,恒定多数类预测器准确率可超90%却无临床价值。