- 论文公开站arXiv
WorldSonus: Bringing Sound to Worlds
Recent advances in world models have enabled increasingly realistic visual synthesis. However, these generated environments remain largely silent. Bringing sound to world models poses three core challenges: real-time gen…
- 论文公开站arXiv
VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning
Self-improving policies continually expose new failure patterns, changing what their judges must be able to verify. However, current fixed judges constrain both optimization feedback and the discovery of useful training …
- 论文公开站arXiv
DepthWorld: 3D World Model for Robot Manipulation
World models offer a data-driven alternative to traditional simulators for robotics, with applications spanning policy evaluation, improvement, and planning. All of these uses depend on faithful 3D geometry, yet current …
- 论文公开站arXiv
Conformal Prediction Sets Quantify Information Gain: A Theoretical Perspective
Conformal prediction is a popular tool for uncertainty quantification that outputs prediction sets with finite-sample coverage guarantees. While prediction set size is commonly used as a heuristic measure of uncertainty,…
- 论文公开站arXiv
TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts
We introduce TasteVal, a benchmark to evaluate the experimental research taste of frontier models. We define research taste as the ability to pick interesting problems to solve, design experiments, and interpret experime…
- 论文公开站arXiv
CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling
Open-source web agents are now strong enough to execute realistic browser tasks, but training them with reinforcement learning still depends on weak supervision: binary task success is too sparse for credit assignment, w…
- 论文公开站arXiv
MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents
Memory has become integral to the LLM agent ecosystem, supporting information retention and reuse across interactions. However, most existing agent memory systems construct memory in a query-agnostic manner, which can in…
- 论文公开站arXiv
Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models
Masked diffusion language models (dLMs) offer a promising parallel alternative to autoregressive models for complex reasoning. However, they face a distinct credit-assignment challenge, since a few commitments during den…
- 论文公开站arXiv
Risk-Aware Input-Constrained Safe Intercept Guidance Against Multiple Moving Defenders
This paper develops a risk-aware guidance law for an attacker to intercept a stationary target in the presence of multiple moving defenders while respecting the control input limits. We represent defender threats via eng…
- 论文公开站arXiv
Transcriptome-informed multi-modal AI for predicting neoadjuvant therapy response from breast cancer biopsies
Scarcity of labeled data limits development of deep learning biomarkers in oncology. We develop a two-stage AI model predicting pathological complete response (pCR) to neoadjuvant therapy in breast cancer. The first stag…
- 论文公开站arXiv
From Mixing to Tearing: Graph Decomposition in Decentralized Optimization via Message Passing
We study the minimization of sums of smooth strongly convex functions over undirected graphs, with each function held by one agent and communication restricted to neighbors in the graph. Existing decentralized methods, w…
- 论文公开站arXiv
EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras
Inspired by human vision, we introduce a framework using active gaze to enable fine-grained bimanual manipulation with only a single stereo camera. EyeRobot 2.0 physically attends to a 3D fixation point in the scene by s…
- 论文公开站arXiv
RNADyn: A Benchmark for Generating and Understanding RNA Dynamics
Ribonucleic acid (RNA) functions through conformational changes that are not fully captured by static structures. However, large-scale standardized RNA dynamics data remain limited, and existing approaches typically trea…
- 论文公开站arXiv
What Should World Models Forget? Stratified Retention for Continual Adaptation
Continual learning treats degradation on previously seen data as evidence of failure, a convention inherited from settings with a stationary prediction target, where a correct label remains correct indefinitely. World mo…
- 论文公开站arXiv
4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes
We introduce 4DCodeBench, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs. To accomplish this, agents must translate visu…
- 论文公开站arXiv
Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features
Intrinsically disordered protein regions (IDRs) play central roles in cellular processes such as transcriptional regulation, signal transduction, and subcellular localization, yet their functional design remains challeng…
- 论文公开站arXiv
Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning
Step-size selection remains a central challenge in large-scale neural network optimization; conservative steps slow convergence, while aggressive steps can destabilize it. We combine \textbf{Z}ero-and-\textbf{F}irst-\tex…
- 论文公开站arXiv
Hierarchical Continuous Diffusion Language Models
Discrete diffusion language models offer a compelling alternative to autoregressive generation for tasks demanding bidirectional reasoning and global constraint satisfaction. Yet they share a structural bottleneck: when …
- 论文公开站arXiv
ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of research, a new problem…
- 论文公开站arXiv
Embedding Prediction Helps Image Generation
In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embeddin…
- 论文公开站arXiv
Compression Footprints as Security Signals for Model-Poisoning Defense in Federated Learning
Lossy compression is widely used in Federated Learning (FL) but is generally treated as an error source, while conventional poisoning defenses inspect update geometry. In this work, we instead treat the compressor's resp…
- 论文公开站arXiv
Cogentic:面向自动证明发现的多智能体编排
Cogentic: Multi-Agent Orchestration for Automated Proof Discovery
提出 Cogentic,一个用于开放研究问题自动证明发现的多智能体框架。前沿语言模型虽能单次生成优质数学想法,但面对需多路径探索的开放问题,单次生成往往不足。
- 论文公开站arXiv
去除时间捷径可改进非侵入式脑机文本解码
Removing Timing Shortcuts Improves Non-Invasive Brain-to-Text
研究发现,多项报道的非侵入式脑信号词解码性能提升,在完全不用脑数据的情况下也能复现。在 d'Ascoli 等人(2025)的工作中,受试者感知连续语音的脑活动时间序列被分段处理。
- 论文公开站arXiv
大模型图重建失真的谱理论:紧界与实证刻画
A Spectral Theory of Distortion in LLM Graph Reconstruction: Sharp Bounds and Empirical Characterization
语言模型图重建评估通常报告原始图与重建图之间的单一聚合距离。本文证明对于拉普拉斯谱间的Wasserstein距离,该汇总受两个边计数约束。
- 论文公开站arXiv
FinAutoRubric:专家引导的金融研究智能体自动评分标准生成
FinAutoRubric: Expert-Guided Automatic Rubric Generation for Evaluating Financial Research Agents
评估金融研究智能体需要反映专家标准并固定信息截止时点取值的评分标准。经专家评审的金融基准依赖固定的逐项评分标准,扩展成本高,且无法编码各机构自身的标准。