- 论文公开站arXiv
Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?
Large language models (LLMs) can solve many narrow tasks, but querying them separately for millions of related instances can be prohibitively expensive. Can LLM agents autonomously create cheaper solutions for such workl…
- 论文公开站arXiv
Sherpa: Teaching LLMs to Teach Adaptively
Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing approaches to training LLMs as teachers rely on demonst…
- 论文公开站arXiv
IdeaAnchor: Teaching LLMs to Turn Literature into Research Ideas
Scientific research often begins by synthesizing ideas from a set of related papers to identify gaps and formulate new directions. However, training language models to perform this form of literature-grounded ideation re…
- 资讯公开站rss_arxiv_cs_ai
SkillScriptBench: Benchmarking Self-Evolution of Executable Agent Skill Packages Beyond Markdown
arXiv:2610.04008v1 Announce Type: new Abstract: Executable Agent Skills combine natural-language instructions and scripts into reusable packages for LLM agents, and revising them requires fixing errors without breaking …
- 资讯公开站rss_arxiv_cs_ai
Behavioral History Outperforms Descriptions of the Person for LLM Synthetic Personas
arXiv:2610.03998v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as synthetic personas representing survey respondents. Their validity as substitutes for particular respondents depends …
- 资讯公开站Digitimes
DIGITIMES Insight: AI servers will carry nearly twice as many CPUs per accelerator by 2027
Demand for agentic AI in commercial and consumer markets was rising rapidly, and the market appeal of consumer services such as Meta Muse was expanding. In agentic AI tasks, large language model (LLM) inference was mainl…
- 资讯公开站Semiconductor Engineering
Chip Industry Technical Paper Roundup: Oct. 6
Low-contact-resistance WSe₂ transistors; backside clock meshes for 2nm GAAFETs; row-parallel processing in DRAM; HBF for high-throughput LLM serving; formal security analysis of CAN XL; abstraction and validation from ph…
- 论文公开站arXiv
MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents
Memory has become integral to the LLM agent ecosystem, supporting information retention and reuse across interactions. However, most existing agent memory systems construct memory in a query-agnostic manner, which can in…
- 论文公开站arXiv
Recursive Video In-Context Learning for Agentic Robot
LLM agents that orchestrate frozen vision-language-action (VLA) policies improve across episodes through text memory, which records what the agent did but not how the task is done. A demonstration video shows it, but fit…
- 论文公开站arXiv
Learning to Read the Contextual Tokens in Diffusion Transformers
Multimodal Diffusion Transformers (MM-DiTs) jointly process visual and textual representations throughout generation. These models repeatedly update the text tokens through multimodal attention, forming dynamic contextua…
- 资讯公开站rss_arxiv_cs_ai
DeReAct: Decomposed Reasoning and Acting for Reliable AI Agents
arXiv:2610.02351v1 Announce Type: new Abstract: ReAct-based agents typically rely on a single LLM policy to propose actions, interact with the environment, and decide when a task is complete. This coupling makes action …
- 资讯公开站rss_arxiv_cs_ai
Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents
arXiv:2610.02330v1 Announce Type: new Abstract: Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condition subsequent deci…
- 资讯公开站MIT Technology Review
The Download: AI’s popularity paradox and EmTech Future 2026
This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. People really hate AI, so why can’t they get enough? —Will Douglas Hea…
- 资讯公开站MIT Technology Review
People really hate AI, so why can’t they get enough?
Over the summer I talked to the CEO of Springboards, a startup building an LLM that’s designed to come up with a wider variety of responses than its mainstream rivals do. At the start of the call, he said something that’…
- 论文公开站arXiv
Planning to Learn
Policy-gradient methods are central to modern reinforcement learning, including LLM post-training. When they struggle, the usual suspects are exploration, credit assignment and action-sampling noise. Classification has n…
- 论文公开站arXiv
FrugalEvo: Towards Cost-Aware LLM-Guided Program Evolution
LLM-guided evolutionary methods, such as AlphaEvolve, have emerged as powerful approaches for challenging computational optimization problems, such as circle packing. However, prior work typically optimizes performance g…
- 资讯公开站MIT Technology Review
The Download: a biological de-aging contest and why LLMs don’t reason
This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. A new contest pits competitors against each other in a race to biologi…
- 资讯公开站Semiconductor Engineering
HBF for High-Throughput LLM Serving (UC Berkeley, FuriosaAI)
Researchers at the UC Berkeley and FuriosaAI published a technical paper titled “Characterizing High Bandwidth Flash for LLM Serving.” Abstract: “Large language model (LLM) serving requires substantial memory to store mo…
- 资讯公开站rss_arxiv_cs_ai
Gradient-Aligned Pair Selection for Personalized Preference Optimization
arXiv:2610.00061v1 Announce Type: new Abstract: Personalizing large language models (LLMs) requires aligning generation behavior with user-specific preferences rather than aggregate quality. While Direct Preference Opti…
- 资讯公开站Digitimes
Global annual AI server shipments, 2025-2026
Continued advances in top-tier LLM capabilities, coupled with the potential for business automation enabled by agentic AI, are prompting major North American cloud providers, Neo Clouds, and leading AI labs to accelerate…
- 资讯公开站MIT Technology Review
Don’t be fooled—LLMs don’t reason
On an afternoon in Seoul in March 2016, I watched a program I helped build put a stone on the fifth line of a Go board in what looked like a gift to its human opponent. Move 37 in game two of the five-game match looked s…
- 论文公开站arXiv
Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning
Step-size selection remains a central challenge in large-scale neural network optimization; conservative steps slow convergence, while aggressive steps can destabilize it. We combine \textbf{Z}ero-and-\textbf{F}irst-\tex…
- 论文公开站arXiv
The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models
While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions. …
- 论文公开站arXiv
InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation
We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills, improving from its own attempts, and retaining what it learns, withou…
- 论文公开站arXiv
TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning
Full-parameter fine-tuning of large language models (LLMs) incurs substantial optimizer state memory overhead, limiting the model sizes that fit on modern GPUs. Existing approaches either compress optimizer state, abando…