- 资讯公开站rss_arxiv_cs_ai
SkillScriptBench: Benchmarking Self-Evolution of Executable Agent Skill Packages Beyond Markdown
arXiv:2610.04008v1 Announce Type: new Abstract: Executable Agent Skills combine natural-language instructions and scripts into reusable packages for LLM agents, and revising them requires fixing errors without breaking …
- 资讯公开站rss_arxiv_cs_ai
Behavioral History Outperforms Descriptions of the Person for LLM Synthetic Personas
arXiv:2610.03998v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as synthetic personas representing survey respondents. Their validity as substitutes for particular respondents depends …
- 资讯公开站Digitimes
DIGITIMES Insight: AI servers will carry nearly twice as many CPUs per accelerator by 2027
Demand for agentic AI in commercial and consumer markets was rising rapidly, and the market appeal of consumer services such as Meta Muse was expanding. In agentic AI tasks, large language model (LLM) inference was mainl…
- 资讯公开站Semiconductor Engineering
Chip Industry Technical Paper Roundup: Oct. 6
Low-contact-resistance WSe₂ transistors; backside clock meshes for 2nm GAAFETs; row-parallel processing in DRAM; HBF for high-throughput LLM serving; formal security analysis of CAN XL; abstraction and validation from ph…
- 资讯公开站rss_arxiv_cs_ai
DeReAct: Decomposed Reasoning and Acting for Reliable AI Agents
arXiv:2610.02351v1 Announce Type: new Abstract: ReAct-based agents typically rely on a single LLM policy to propose actions, interact with the environment, and decide when a task is complete. This coupling makes action …
- 资讯公开站rss_arxiv_cs_ai
Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents
arXiv:2610.02330v1 Announce Type: new Abstract: Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condition subsequent deci…
- 资讯公开站MIT Technology Review
The Download: AI’s popularity paradox and EmTech Future 2026
This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. People really hate AI, so why can’t they get enough? —Will Douglas Hea…
- 资讯公开站MIT Technology Review
People really hate AI, so why can’t they get enough?
Over the summer I talked to the CEO of Springboards, a startup building an LLM that’s designed to come up with a wider variety of responses than its mainstream rivals do. At the start of the call, he said something that’…
- 资讯公开站MIT Technology Review
The Download: a biological de-aging contest and why LLMs don’t reason
This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. A new contest pits competitors against each other in a race to biologi…
- 资讯公开站Semiconductor Engineering
HBF for High-Throughput LLM Serving (UC Berkeley, FuriosaAI)
Researchers at the UC Berkeley and FuriosaAI published a technical paper titled “Characterizing High Bandwidth Flash for LLM Serving.” Abstract: “Large language model (LLM) serving requires substantial memory to store mo…
- 资讯公开站rss_arxiv_cs_ai
Gradient-Aligned Pair Selection for Personalized Preference Optimization
arXiv:2610.00061v1 Announce Type: new Abstract: Personalizing large language models (LLMs) requires aligning generation behavior with user-specific preferences rather than aggregate quality. While Direct Preference Opti…
- 资讯公开站Digitimes
Global annual AI server shipments, 2025-2026
Continued advances in top-tier LLM capabilities, coupled with the potential for business automation enabled by agentic AI, are prompting major North American cloud providers, Neo Clouds, and leading AI labs to accelerate…
- 资讯公开站MIT Technology Review
Don’t be fooled—LLMs don’t reason
On an afternoon in Seoul in March 2016, I watched a program I helped build put a stone on the fifth line of a Go board in what looked like a gift to its human opponent. Move 37 in game two of the five-game match looked s…
- 资讯公开站Digitimes
Commentary: AMD challenges Nvidia in embodied AI as Lisa Su acquires Fei-Fei Li's World Labs
Advanced Micro Devices (AMD) has agreed to acquire World Labs in an all-stock transaction valued at approximately US$8.2 billion. Following the closing, World Labs will operate as a dedicated frontier research organizati…
- 资讯公开站rss_arxiv_cs_ai
Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization Limits
arXiv:2609.38386v1 Announce Type: new Abstract: Concurrent autoregressive inference creates a fundamental interference problem: prefilling a newly arrived long prompt can delay tokens for requests that are already decod…
- 资讯公开站rss_arxiv_cs_ai
Aligned Data Can Induce Misalignment via Context Confusion
arXiv:2609.38379v1 Announce Type: new Abstract: Large language models (LLMs) are frequently updated for various use cases, where filtering out misaligned training samples is a common practice for preventing post-update …
- 资讯公开站rss_arxiv_cs_ai
CARAT: Do Materials LLMs Reason or Recite?
arXiv:2609.38340v1 Announce Type: new Abstract: When a materials LLM answers a question about crystal structure, does it reason from the structure or copy an answer already printed in its input? Accuracy cannot tell: a …
- 资讯公开站Supply Chain Dive
Walmart to build Ohio fulfillment center for oversized goods
The $300 million investment near Cincinnati will specialize in handling larger, non-sortable items, including televisions and furniture.
- 资讯公开站Semiconductor Engineering
Why LLMs Are The Best Thing To Happen To Chip Design
- 资讯公开站Semiconductor Engineering
LLM Performance And Acceleration: Part 1
- 资讯公开站rss_arxiv_cs_ai
更多程序还是更多重复?区分LLM执行框架中的覆盖度与专业化
More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses
arXiv:2609.35873v1 公告类型:新 摘要:LLM执行框架的自动生成有望通过任务专业化提升推理能力。然而,额外的答案覆盖度也可能来自同一程序的重复执行,这使得专业化难以识别。我们引入一种受控评估,将答案覆盖度、可重复的任务优势以及执行前选择带来的增益区分开来。在386个MATH-500任务上,我们比较了八个生成的执行框架以及一个基线及其九个字节完全相同的副本,每个成员执行三次。相同程序产生了2.16个百分点的重复平均ora…
- 资讯公开站rss_arxiv_cs_ai
有效的大语言模型微调是否必须依赖人类可读文本?
Is Human-Readable Text Necessary for Effective LLM Fine-Tuning?
arXiv:2609.35868v1 公告类型:新 摘要:有效微调大语言模型是否必须依赖人类可读性?我们研究模型条件化的训练表示能否在无需人类可读文本形式的情况下保持或提升适配效用。我们提出 Desired-Update-Aligned Synthetic Data(DASA),利用冻结参考模型的激活梯度反馈来指导连续合成输入嵌入的优化。受激活梯度在局部风险降低中作用的启发,DASA 针对有用的适配更新,而非源文本重建或语言流畅性。所得…
- 资讯公开站Semiconductor Engineering
芯片产业技术论文汇总:9月29日
Chip Industry Technical Paper Roundup: Sept. 29
涵盖A7 CFET与A10 NSFET对比、晶圆级亚5nm MoS₂晶体管、3D HI多千瓦供电、铜微结构与TSV残余应力、GPU Rowhammer攻击、门级RTL木马定位、LLM推理中HBM与主机内存并发访问等。
- 资讯公开站rss_arxiv_cs_ai
BioDyad:同步生物医学发现与机器学习工程
BioDyad: Synchronize Biomedical Discovery and Machine Learning Engineering
arXiv:2609.31939v1 公告类型:新 摘要:智能体生物医学机器学习(ML)依托生物医学证据获取与可执行程序搜索两方面的互补进展。现有系统连接了这些能力的部分方面,但在整个程序搜索过程中协调它们仍具挑战性。新证据必须指导候选构建,执行结果必须为后续发现与复用提供信息,验证需求必须契合搜索预算。我们提出 BioDyad,通过蒙特卡洛图搜索中的两个层级将生物医学发现与 ML 工程耦合。其科学层级将先验生物医学指导与迭代发现相结合…
- 资讯公开站rss_arxiv_cs_ai
通过嵌入式编码改进大语言模型的医学计算能力
Improving Medical Calculation of LLMs with Embedded Coding
arXiv:2609.31908v1 公告类型:新 摘要:大语言模型(LLM)在医学考试和问答基准上表现良好,但在需要精确数值输出的医学计算任务上仍不可靠。这些计算支撑着用药剂量、器官功能评估和预后评分等高风险决策,即使很小的误差也可能带来严重的临床后果。我们提出 MedCode,一个通过训练 LLM 生成嵌入式可执行代码来改进医学计算的框架。给定临床情境,模型识别相关计算器,提取其输入变量,并生成一段将算术运算委托给确定性解释器的脚本…