- 论文公开站arXiv
WorldSonus: Bringing Sound to Worlds
Recent advances in world models have enabled increasingly realistic visual synthesis. However, these generated environments remain largely silent. Bringing sound to world models poses three core challenges: real-time gen…
- 论文公开站arXiv
Small Distortions, Big Polarization: Tetragonal BaTiO3 Nanoparticles for High-Performance Piezoelectric Nanogenerators
Lead-free ferroelectric BaTiO3 (BTO) nanoparticles with stabilized tetragonal distortion were synthesized via a ligand-assisted sol-gel method to serve as the active piezoelectric phase in high-performance hybrid piezoel…
- 论文公开站arXiv
Sherpa: Teaching LLMs to Teach Adaptively
Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing approaches to training LLMs as teachers rely on demonst…
- 论文公开站arXiv
Conformal Prediction Sets Quantify Information Gain: A Theoretical Perspective
Conformal prediction is a popular tool for uncertainty quantification that outputs prediction sets with finite-sample coverage guarantees. While prediction set size is commonly used as a heuristic measure of uncertainty,…
- 论文公开站arXiv
TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts
We introduce TasteVal, a benchmark to evaluate the experimental research taste of frontier models. We define research taste as the ability to pick interesting problems to solve, design experiments, and interpret experime…
- 论文公开站arXiv
Towards Looped Models Done Right, Part II: Rethinking at Fixed Points
Every recurrence of a looped language model adds cost in training, decoding, prefill, and reinforcement learning (RL). The closer recurrent states get to fixed points, the less the path to them matters. This enables trun…
- 论文公开站arXiv
InterMimicGen: Scaling Humanoid Loco-Manipulation through Self-Evolving Motion Imitation
Captured human-object interactions provide rich supervision for humanoid loco-manipulation, but they are sparse, heterogeneous, and not directly executable by robots. We introduce InterMimicGen, a self-evolving motion-im…
- 论文公开站arXiv
Revisiting Input Time-frequency Representations in Multi-pitch Estimation for Vocal Ensembles
Multi-pitch estimation in vocal ensembles is challenging because singers occupy overlapping pitch ranges and often sing at closely spaced fundamental frequencies, causing their harmonics to overlap in time-frequency repr…
- 论文公开站arXiv
FrugalEvo: Towards Cost-Aware LLM-Guided Program Evolution
LLM-guided evolutionary methods, such as AlphaEvolve, have emerged as powerful approaches for challenging computational optimization problems, such as circle packing. However, prior work typically optimizes performance g…
- 论文公开站arXiv
Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features
Intrinsically disordered protein regions (IDRs) play central roles in cellular processes such as transcriptional regulation, signal transduction, and subcellular localization, yet their functional design remains challeng…
- 论文公开站arXiv
The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models
While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions. …
- 论文公开站arXiv
Cost-augmented Schrödinger bridges on graphs are exactly solvable: a Feynman-Kac tilt replaces learned control
The generalized Schrödinger bridge on a graph moves mass between two distributions while charging a cost for the states visited. It has been approached by learning the rates of a controlled continuous-time Markov chain, …
- 论文公开站arXiv
FERPO: Forward Entropy-Regularized Policy Optimization
Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurat…
- 论文公开站arXiv
SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation
High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict active structure and then synthesize local geometry. While effective, this design fragments continuous surfa…
- 论文公开站arXiv
MatLoom: Layered Text-to-Material Generation in a Compact Program Space
Material generation should produce not only an appearance, but also the rules that construct it. We introduce MatLoom, a compact, layer-oriented language for text-to-material generation with pretrained language models. E…
- 论文公开站arXiv
Turbo Harness:实例自适应框架优化
Turbo Harness: Instance-Adaptive Harness Optimization
自动化搜索有效框架是实现智能体递归自我改进的重要一步。现有优化通常产出统一应用于所有任务的单一全局框架,但适用于某类任务的框架未必通用。
- 论文公开站arXiv
Ego4WAM:扩展第一人称人类数据用于机器人学习的关键因素
Ego4WAM: What Matters When Scaling Egocentric Human Data for Robot Learning?
第一人称人类数据为机器人学习提供可扩展经验,但在人机对齐、行为覆盖和监督信号上差异显著。已有工作显示数据越多效果越好,但哪些数据属性真正关键仍不明确。
- 论文公开站arXiv
图像分类器是高效的自监督视频表征学习器
Image Classifiers are Efficient Self-Supervised Video Representation Learners
提出 VideoMSN,一种用于视频高效自监督时空表征学习的掩码孪生网络框架。它不依赖重型3D架构或重建式自编码器,而是复用标准图像分类器处理无标注数据。
- 论文公开站arXiv
去除时间捷径可改进非侵入式脑机文本解码
Removing Timing Shortcuts Improves Non-Invasive Brain-to-Text
研究发现,多项报道的非侵入式脑信号词解码性能提升,在完全不用脑数据的情况下也能复现。在 d'Ascoli 等人(2025)的工作中,受试者感知连续语音的脑活动时间序列被分段处理。
- 论文公开站arXiv
面向多模态临床诊断的排序感知提示优化
Ranking-Aware Prompt Optimization for Multimodal Clinical Diagnosis
多模态大模型正快速推进临床诊断,但其适配流程仍以准确率为目标。临床数据类别严重不平衡,恒定多数类预测器准确率可超90%却无临床价值。
- 论文公开站arXiv
打破同质化陷阱:通过SplitMoE扩展视频扩散模型
Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE
受大语言模型启发的混合专家(MoE)是扩展视觉生成模型的有前景范式,但传统token级MoE在同类专家池中独立路由token并趋向均匀化,限制了性能。本文提出SplitMoE方法。
- 论文公开站arXiv
学习元技能用于测试时AI4AI的智能体框架设计
Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI
智能体性能取决于推理能力和所处环境。本文研究测试时AI-for-AI,探讨Builder如何在两个模型权重固定的情况下为Target构建更好的执行环境。
- 论文公开站arXiv
思考之前先思考:通过元推理扩展智能体推理
Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning
随着智能体处理更长更复杂的问题,控制执行本身成为一项任务。本文提出智能体元推理,在运行中做出选择部分工作、是否重新开始或何时停止等控制决策。
- 论文公开站arXiv
重新思考世界-动作建模的表征
Rethinking Representations for World-Action Modeling
世界-动作模型联合学习机器人策略并预测未来观测,使表征空间成为控制与预测的接口。本文通过受控比较研究该空间的设计,发现重建保真度和预训练均非关键。
- 论文公开站arXiv
TokenCast:预测LLM智能体执行中的Token消耗
TokenCast: Forecasting Token Consumption During LLM Agent Execution
同一任务下LLM智能体各次运行的Token消耗可相差一个数量级以上,因智能体依据工具反馈与中间结果选择下一步,且上下文不断膨胀推高输入规模。