- 论文公开站arXiv
Small Distortions, Big Polarization: Tetragonal BaTiO3 Nanoparticles for High-Performance Piezoelectric Nanogenerators
Lead-free ferroelectric BaTiO3 (BTO) nanoparticles with stabilized tetragonal distortion were synthesized via a ligand-assisted sol-gel method to serve as the active piezoelectric phase in high-performance hybrid piezoel…
- 论文公开站arXiv
Rapid Fredholm stabilization of the Kuramoto--Sivashinsky equation with unrestricted, spatially-varying anti-diffusion
We develop the first feedback design for rapid stabilization of the Kuramoto--Sivashinsky equation with a spatially varying anti-diffusion coefficient. For constant coefficients, the single-input Fredholm design of Coron…
- 论文公开站arXiv
Paradee: Distilling Kokoro-82M into an 8M-Parameter Single-Voice Text-to-Speech Model
We distill Kokoro-82M, a widely used open text-to-speech model with 54 voices, into Paradee, an 8.07M-parameter model that speaks one of them. Paradee keeps Kokoro's architecture with much narrower layers, and each of it…
- 论文公开站arXiv
TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts
We introduce TasteVal, a benchmark to evaluate the experimental research taste of frontier models. We define research taste as the ability to pick interesting problems to solve, design experiments, and interpret experime…
- 论文公开站arXiv
Direct Intermediate Initialization for Tilted Diffusion Samplers
Some diffusion posterior samplers construct Gaussian-tilted intermediate distributions along the reverse process. We observe that these targets can be pulled back to clean-space posteriors with weaker conditioning, with …
- 论文公开站arXiv
Simulation-Free Learning of Population Dynamics with Wasserstein Lagrangian Residuals
The dynamics of cells, organisms, and fluids are often modeled as probability distributions evolving over time. Reconstructing and extrapolating this evolution from unpaired snapshots requires assumptions about the under…
- 论文公开站arXiv
LESSER: Post-Training Data Selection with Output-Layer Gradients
The choice of post-training data for large language models substantially affects downstream performance. Gradient-based data selection is a popular approach that ranks training data by how well their gradients align with…
- 论文公开站arXiv
Less Decoder is More Encoder: Geometric Representation Learning from Novel View Synthesis
This paper examines the role of Novel View Synthesis (NVS) in geometric representation learning. In principle, NVS should reason about 3D scene structure, thereby enabling transferable multi-view geometric representation…
- 论文公开站arXiv
TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning
Full-parameter fine-tuning of large language models (LLMs) incurs substantial optimizer state memory overhead, limiting the model sizes that fit on modern GPUs. Existing approaches either compress optimizer state, abando…
- 论文公开站arXiv
SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation
High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict active structure and then synthesize local geometry. While effective, this design fragments continuous surfa…
- 论文公开站arXiv
Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents
Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real…
- 论文公开站arXiv
One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars
3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bottleneck and show that the animation of pretrained avatar models can be…
- 论文公开站arXiv
MatLoom: Layered Text-to-Material Generation in a Compact Program Space
Material generation should produce not only an appearance, but also the rules that construct it. We introduce MatLoom, a compact, layer-oriented language for text-to-material generation with pretrained language models. E…
- 论文公开站arXiv
DP-SGD 下权重绑定对纯解码器 LLM 是否仍有益?
Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?
DP-SGD 是大模型隐私保护微调的主流方法。许多纯解码器 LLM 采用输入输出嵌入权重绑定,本文在差分隐私设定下重新审视这一设计是否仍然有利。
- 论文公开站arXiv
AssemblyWorld:用通用智能体重思3D装配
AssemblyWorld: Rethinking 3D Assembly with General-Purpose Agents
3D装配需将零件及其关系的理解转化为精确空间排布。本文探究预训练通用智能体能否仅通过视觉交互完成物体装配,而无需针对装配任务额外微调。
- 论文公开站arXiv
去除时间捷径可改进非侵入式脑机文本解码
Removing Timing Shortcuts Improves Non-Invasive Brain-to-Text
研究发现,多项报道的非侵入式脑信号词解码性能提升,在完全不用脑数据的情况下也能复现。在 d'Ascoli 等人(2025)的工作中,受试者感知连续语音的脑活动时间序列被分段处理。
- 论文公开站arXiv
AdviSD:通过定向多轮自蒸馏学习指导前沿大模型
AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation
小型可训练顾问可利用自然语言建议引导冻结的语言模型执行器。除任务奖励外,顾问还可利用已完成交互的反馈改进建议,但看似合理的纠正未必能改变执行结果。
- 论文公开站arXiv
思考之前先思考:通过元推理扩展智能体推理
Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning
随着智能体处理更长更复杂的问题,控制执行本身成为一项任务。本文提出智能体元推理,在运行中做出选择部分工作、是否重新开始或何时停止等控制决策。
- 论文公开站arXiv
大模型图重建失真的谱理论:紧界与实证刻画
A Spectral Theory of Distortion in LLM Graph Reconstruction: Sharp Bounds and Empirical Characterization
语言模型图重建评估通常报告原始图与重建图之间的单一聚合距离。本文证明对于拉普拉斯谱间的Wasserstein距离,该汇总受两个边计数约束。
- 论文公开站arXiv
MINT:建模生成式AI对网络流量的影响
MINT: Modeling GenAI Impact on Network Traffic
生成式AI正成为主流网络负载,但数据包级模拟器缺乏基于实测的GenAI流量模型。目前研究人员只能用文件传输和视频流等传统来源近似GenAI服务,限制了真实性。
- 论文公开站arXiv
照搬相同,蒸馏差异:线性视觉Transformer的初始化
Copy the Same, Distill the Difference: Initializing Linear Vision Transformers
线性视觉Transformer(ViT)旨在用线性复杂度注意力算子替代Softmax ViT中的注意力,以实现更高效的token路由,但需要从头预训练,且通常性能不及原版Softmax。如何初始化是关键。
- 论文公开站arXiv
伸缩式语言模型
Telescopic Language Models
一个已部署的语言模型常需服务多种算力预算,而每个预算点仍需单独训练或压缩;本文训练伸缩式语言模型(TLM)作为连续体,由随机前缀监督的嵌套容量Transformer构成。
- 论文公开站arXiv
通过发现时间解释中的重复概念揭示音频分类器中的捷径学习
Uncovering shortcut learning in audio classifiers by discovering recurring concepts in temporal explanations
机器学习数据集中事件间的相关性可能导致捷径学习,模型基于相关事件的出现来预测目标事件。当这些相关性是虚假的(源于数据收集偏差)时,模型很可能学到错误关联。
- 论文公开站arXiv
Kafila:在可信异构消费级机器集合上服务大语言模型
Kafila: Serving Large Language Models on a Trusted Set of Heterogeneous Commodity Machines
研究组成员或朋友圈各自拥有几台消费级计算机,但单台都不足以运行一个有能力的大语言模型。现有系统在开放集群中汇聚此类算力,而仅接纳可信机器的群体则无法使用。
- 论文公开站arXiv
ARCH-B:建筑表征、理解与层级基准
ARCH-B: Architectural Representation, Comprehension and Hierarchy Benchmark
多模态模型越来越多地解读视觉环境,但其在照片、平面图、立面图、剖面图和渲染图中识别同一建筑的能力仍缺乏表征。我们推出 ARCH-B,一个包含 354 道四选一题的基准。