- 论文公开站arXiv
Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?
Large language models (LLMs) can solve many narrow tasks, but querying them separately for millions of related instances can be prohibitively expensive. Can LLM agents autonomously create cheaper solutions for such workl…
- 论文公开站arXiv
Sherpa: Teaching LLMs to Teach Adaptively
Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing approaches to training LLMs as teachers rely on demonst…
- 论文公开站arXiv
IdeaAnchor: Teaching LLMs to Turn Literature into Research Ideas
Scientific research often begins by synthesizing ideas from a set of related papers to identify gaps and formulate new directions. However, training language models to perform this form of literature-grounded ideation re…
- 论文公开站arXiv
MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents
Memory has become integral to the LLM agent ecosystem, supporting information retention and reuse across interactions. However, most existing agent memory systems construct memory in a query-agnostic manner, which can in…
- 论文公开站arXiv
Recursive Video In-Context Learning for Agentic Robot
LLM agents that orchestrate frozen vision-language-action (VLA) policies improve across episodes through text memory, which records what the agent did but not how the task is done. A demonstration video shows it, but fit…
- 论文公开站arXiv
Learning to Read the Contextual Tokens in Diffusion Transformers
Multimodal Diffusion Transformers (MM-DiTs) jointly process visual and textual representations throughout generation. These models repeatedly update the text tokens through multimodal attention, forming dynamic contextua…
- 论文公开站arXiv
Planning to Learn
Policy-gradient methods are central to modern reinforcement learning, including LLM post-training. When they struggle, the usual suspects are exploration, credit assignment and action-sampling noise. Classification has n…
- 论文公开站arXiv
FrugalEvo: Towards Cost-Aware LLM-Guided Program Evolution
LLM-guided evolutionary methods, such as AlphaEvolve, have emerged as powerful approaches for challenging computational optimization problems, such as circle packing. However, prior work typically optimizes performance g…
- 论文公开站arXiv
Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning
Step-size selection remains a central challenge in large-scale neural network optimization; conservative steps slow convergence, while aggressive steps can destabilize it. We combine \textbf{Z}ero-and-\textbf{F}irst-\tex…
- 论文公开站arXiv
The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models
While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions. …
- 论文公开站arXiv
InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation
We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills, improving from its own attempts, and retaining what it learns, withou…
- 论文公开站arXiv
TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning
Full-parameter fine-tuning of large language models (LLMs) incurs substantial optimizer state memory overhead, limiting the model sizes that fit on modern GPUs. Existing approaches either compress optimizer state, abando…
- 论文公开站arXiv
Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents
Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real…
- 论文公开站arXiv
KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agen…
- 论文公开站arXiv
DP-SGD 下权重绑定对纯解码器 LLM 是否仍有益?
Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?
DP-SGD 是大模型隐私保护微调的主流方法。许多纯解码器 LLM 采用输入输出嵌入权重绑定,本文在差分隐私设定下重新审视这一设计是否仍然有利。
- 论文公开站arXiv
去除时间捷径可改进非侵入式脑机文本解码
Removing Timing Shortcuts Improves Non-Invasive Brain-to-Text
研究发现,多项报道的非侵入式脑信号词解码性能提升,在完全不用脑数据的情况下也能复现。在 d'Ascoli 等人(2025)的工作中,受试者感知连续语音的脑活动时间序列被分段处理。
- 论文公开站arXiv
半事实信用增强策略优化
Semifactual Credit-Augmented Policy Optimization
可验证奖励强化学习(RLVR)提升了 LLM 推理能力,但其预测仍对任务无关的提示特征敏感。本文通过半事实提示干预研究这种敏感性。
- 论文公开站arXiv
面向多模态临床诊断的排序感知提示优化
Ranking-Aware Prompt Optimization for Multimodal Clinical Diagnosis
多模态大模型正快速推进临床诊断,但其适配流程仍以准确率为目标。临床数据类别严重不平衡,恒定多数类预测器准确率可超90%却无临床价值。
- 论文公开站arXiv
AdviSD:通过定向多轮自蒸馏学习指导前沿大模型
AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation
小型可训练顾问可利用自然语言建议引导冻结的语言模型执行器。除任务奖励外,顾问还可利用已完成交互的反馈改进建议,但看似合理的纠正未必能改变执行结果。
- 论文公开站arXiv
大模型图重建失真的谱理论:紧界与实证刻画
A Spectral Theory of Distortion in LLM Graph Reconstruction: Sharp Bounds and Empirical Characterization
语言模型图重建评估通常报告原始图与重建图之间的单一聚合距离。本文证明对于拉普拉斯谱间的Wasserstein距离,该汇总受两个边计数约束。
- 论文公开站arXiv
LeapQuant:具有精确循环状态量化的高效线性注意力
LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization
近期大模型越来越多采用线性注意力替代标准注意力,如Gated DeltaNet和Kimi Delta Attention。这些方法将上下文压缩为固定大小循环状态,大幅降低长上下文成本。本文提出LeapQuant。
- 论文公开站arXiv
MINT:建模生成式AI对网络流量的影响
MINT: Modeling GenAI Impact on Network Traffic
生成式AI正成为主流网络负载,但数据包级模拟器缺乏基于实测的GenAI流量模型。目前研究人员只能用文件传输和视频流等传统来源近似GenAI服务,限制了真实性。
- 论文公开站arXiv
面向智能体强化学习高效压缩的KV-streams
KV-streams for Efficient Compaction in Agentic Reinforcement Learning
扩展智能体LLM的时域受限于需将越来越长的上下文轨迹装入GPU内存。上下文压缩是缓解该问题最常用的机制,可在给定轨迹下保持GPU内存恒定,但多数压缩策略存在不足。
- 论文公开站arXiv
TokenCast:预测LLM智能体执行中的Token消耗
TokenCast: Forecasting Token Consumption During LLM Agent Execution
同一任务下LLM智能体各次运行的Token消耗可相差一个数量级以上,因智能体依据工具反馈与中间结果选择下一步,且上下文不断膨胀推高输入规模。
- 论文公开站arXiv
通过信念自蒸馏进行用户模型提取
User Model Extraction via Belief Self-Distillation
摘要显示,该研究提出 Belief Self-Distillation(BSD)框架,让冻结的 LLM 充当自身教师,从自然对话中蒸馏出紧凑的用户表征,既可解码也可写回模型,从而将线性探测与因果探测统一。实验表明,BSD 能跨多个模型家族忠实恢复用户信念,且干预效果显著强于匹配的隐状态引导;改变模型推断的用户意图会改变拒绝行为,而请求本身保持不变。研究还发现不同独立训练的 LLM 在用户表征上收敛出共享几何结构。
意义:该工作把用户信念变为可读可写的内部状态,为AI安全中的拒绝行为归因与干预提供了新工具,也提示跨模型用户表征可能存在共性。