- 论文公开站arXiv
AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model
Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, be…
- 论文公开站arXiv
DepthWorld: 3D World Model for Robot Manipulation
World models offer a data-driven alternative to traditional simulators for robotics, with applications spanning policy evaluation, improvement, and planning. All of these uses depend on faithful 3D geometry, yet current …
- 论文公开站arXiv
IdeaAnchor: Teaching LLMs to Turn Literature into Research Ideas
Scientific research often begins by synthesizing ideas from a set of related papers to identify gaps and formulate new directions. However, training language models to perform this form of literature-grounded ideation re…
- 论文公开站arXiv
QF3: Fast Flow RL with Filtered Q-Gradients
Flow policies have become a standard policy class for learning robot behaviors from demonstrations, but reinforcement learning is still critical for improving pre-trained flow policies or learning them from scratch throu…
- 论文公开站arXiv
Towards Looped Models Done Right, Part II: Rethinking at Fixed Points
Every recurrence of a looped language model adds cost in training, decoding, prefill, and reinforcement learning (RL). The closer recurrent states get to fixed points, the less the path to them matters. This enables trun…
- 论文公开站arXiv
Learning to Read the Contextual Tokens in Diffusion Transformers
Multimodal Diffusion Transformers (MM-DiTs) jointly process visual and textual representations throughout generation. These models repeatedly update the text tokens through multimodal attention, forming dynamic contextua…
- 论文公开站arXiv
Planning to Learn
Policy-gradient methods are central to modern reinforcement learning, including LLM post-training. When they struggle, the usual suspects are exploration, credit assignment and action-sampling noise. Classification has n…
- 论文公开站arXiv
Simulation-Free Learning of Population Dynamics with Wasserstein Lagrangian Residuals
The dynamics of cells, organisms, and fluids are often modeled as probability distributions evolving over time. Reconstructing and extrapolating this evolution from unpaired snapshots requires assumptions about the under…
- 论文公开站arXiv
Risk-Aware Input-Constrained Safe Intercept Guidance Against Multiple Moving Defenders
This paper develops a risk-aware guidance law for an attacker to intercept a stationary target in the presence of multiple moving defenders while respecting the control input limits. We represent defender threats via eng…
- 论文公开站arXiv
LESSER: Post-Training Data Selection with Output-Layer Gradients
The choice of post-training data for large language models substantially affects downstream performance. Gradient-based data selection is a popular approach that ranks training data by how well their gradients align with…
- 论文公开站arXiv
Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features
Intrinsically disordered protein regions (IDRs) play central roles in cellular processes such as transcriptional regulation, signal transduction, and subcellular localization, yet their functional design remains challeng…
- 论文公开站arXiv
Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning
Step-size selection remains a central challenge in large-scale neural network optimization; conservative steps slow convergence, while aggressive steps can destabilize it. We combine \textbf{Z}ero-and-\textbf{F}irst-\tex…
- 论文公开站arXiv
Cost-augmented Schrödinger bridges on graphs are exactly solvable: a Feynman-Kac tilt replaces learned control
The generalized Schrödinger bridge on a graph moves mass between two distributions while charging a cost for the states visited. It has been approached by learning the rates of a controlled continuous-time Markov chain, …
- 论文公开站arXiv
FERPO: Forward Entropy-Regularized Policy Optimization
Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurat…
- 论文公开站arXiv
TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning
Full-parameter fine-tuning of large language models (LLMs) incurs substantial optimizer state memory overhead, limiting the model sizes that fit on modern GPUs. Existing approaches either compress optimizer state, abando…
- 论文公开站arXiv
WorldAuditBench:多模态智能体的交互式3D世界审计
WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents
随着交互式3D世界被广泛用于研究智能行为,识别其中异常(如悬浮物体、可穿墙、与环境不一致的物体)的高效流程变得重要,本文提出相应基准。
- 论文公开站arXiv
DP-SGD 下权重绑定对纯解码器 LLM 是否仍有益?
Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?
DP-SGD 是大模型隐私保护微调的主流方法。许多纯解码器 LLM 采用输入输出嵌入权重绑定,本文在差分隐私设定下重新审视这一设计是否仍然有利。
- 论文公开站arXiv
打破同质化陷阱:通过SplitMoE扩展视频扩散模型
Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE
受大语言模型启发的混合专家(MoE)是扩展视觉生成模型的有前景范式,但传统token级MoE在同类专家池中独立路由token并趋向均匀化,限制了性能。本文提出SplitMoE方法。
- 论文公开站arXiv
重新思考世界-动作建模的表征
Rethinking Representations for World-Action Modeling
世界-动作模型联合学习机器人策略并预测未来观测,使表征空间成为控制与预测的接口。本文通过受控比较研究该空间的设计,发现重建保真度和预训练均非关键。
- 论文公开站arXiv
LeapQuant:具有精确循环状态量化的高效线性注意力
LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization
近期大模型越来越多采用线性注意力替代标准注意力,如Gated DeltaNet和Kimi Delta Attention。这些方法将上下文压缩为固定大小循环状态,大幅降低长上下文成本。本文提出LeapQuant。
- 论文公开站arXiv
简化上下文机器人学习:面向操作任务的民主化方案
In-context Robot Learning Made Simple: A Democratized Recipe for Manipulation Tasks
本文研究机器人上下文学习(ICL),这一新兴范式使机器人能从视觉演示中推断并执行任务。尽管前景广阔,但问题本身仍定义不清:视觉演示同时传达动作轨迹和物体信息。
- 论文公开站arXiv
MINT:建模生成式AI对网络流量的影响
MINT: Modeling GenAI Impact on Network Traffic
生成式AI正成为主流网络负载,但数据包级模拟器缺乏基于实测的GenAI流量模型。目前研究人员只能用文件传输和视频流等传统来源近似GenAI服务,限制了真实性。
- 论文公开站arXiv
FinAutoRubric:专家引导的金融研究智能体自动评分标准生成
FinAutoRubric: Expert-Guided Automatic Rubric Generation for Evaluating Financial Research Agents
评估金融研究智能体需要反映专家标准并固定信息截止时点取值的评分标准。经专家评审的金融基准依赖固定的逐项评分标准,扩展成本高,且无法编码各机构自身的标准。
- 论文公开站arXiv
TokenCast:预测LLM智能体执行中的Token消耗
TokenCast: Forecasting Token Consumption During LLM Agent Execution
同一任务下LLM智能体各次运行的Token消耗可相差一个数量级以上,因智能体依据工具反馈与中间结果选择下一步,且上下文不断膨胀推高输入规模。
- 论文公开站arXiv
DexRoam:从第一人称全身人类演示学习移动双臂灵巧操作
DexRoam: Learning Mobile Bimanual Dexterous Manipulation from Egocentric Whole-Body Human Demonstrations
移动双臂灵巧操作需在单条轨迹中协调运动、全身动作与手指级灵巧性,造成严重的机器人演示瓶颈;第一人称人类演示提供了可扩展的替代方案。