- 资讯公开站Digitimes
DIGITIMES Insight: AI servers will carry nearly twice as many CPUs per accelerator by 2027
Demand for agentic AI in commercial and consumer markets was rising rapidly, and the market appeal of consumer services such as Meta Muse was expanding. In agentic AI tasks, large language model (LLM) inference was mainl…
- 资讯公开站Digitimes
AMD to expand Taiwan supply chain investment as AI demand drives capacity needs
AMD CEO Lisa Su said the company plans to further increase investment in Taiwan's semiconductor supply chain as demand for CPUs, GPUs, and AI computing continues to outpace available capacity.
- 资讯公开站gnews_nvidia_supply
Calif. CEO Charged in $300M Nvidia GPU Smuggling Case [2026] - tech-insider.org
- 资讯公开站rss_arxiv_cs_ai
FlashSinkhorn 2: Block-Sparse Entropic Optimal Transport
arXiv:2610.02395v1 Announce Type: new Abstract: Streaming GPU solvers for entropic optimal transport (EOT), such as FlashSinkhorn, avoid storing the dense kernel but still evaluate all $n\times m$ point pairs in every S…
- 资讯公开站Digitimes
AM Intelligence expands India AI infrastructure push with 20,000 more Nvidia Rubin GPUs
AM Intelligence (AMI) is widening its AI infrastructure push in India and Malaysia, adding to a regional buildout that could matter well beyond Asia. The latest orders expand access to frontier computing for cloud provid…
- 资讯公开站Digitimes
Biren readies BR20X, sharpening China's AI GPU challenge to Nvidia
Biren Technology is preparing its next-generation GPU for mass production, advancing a domestic AI accelerator roadmap built around broader low-precision support and locally available manufacturing.
- 资讯公开站Digitimes
GMI Cloud secures NT$14.05 billion loan for AI factory
CTBC Bank and GMI Cloud on October 1 announced the closing of an NT$14.05 billion (about US$445 million) syndicated credit facility to fund Taiwan's first AI factory. The deal is the first in Taiwan to fully back GPU fin…
- 论文公开站arXiv
TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning
Full-parameter fine-tuning of large language models (LLMs) incurs substantial optimizer state memory overhead, limiting the model sizes that fit on modern GPUs. Existing approaches either compress optimizer state, abando…
- 资讯公开站rss_arxiv_cs_ai
Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization Limits
arXiv:2609.38386v1 Announce Type: new Abstract: Concurrent autoregressive inference creates a fundamental interference problem: prefilling a newly arrived long prompt can delay tokens for requests that are already decod…
- 资讯公开站Digitimes
Neweb Information targets late-October listing on Taipei Exchange
Neweb Information, a Taiwan-based IT systems integrator, is targeting a late-October listing on the Taipei Exchange. Long focused on enterprise core systems, the company has expanded in recent years into AI GPU computing…
- 资讯公开站Digitimes
2027年ASIC出货量超越GPU,ABF载板瓶颈转移
ASIC shipments top GPUs in 2027 as ABF substrate bottlenecks shift
2027年高端云AI加速器市场将发生结构性转变,ASIC总出货量首次超过GPU,受谷歌、亚马逊和华为产量上升推动,市场进入英伟达与ASIC双雄竞争格局。
- 资讯公开站rss_arxiv_cs_ai
PowerZooJax:面向强化学习的基于JAX的电力系统基准
PowerZooJax: A JAX-based Power System Benchmark for Reinforcement Learning
arXiv:2609.36052v1 公告类型:新 摘要:电力系统运行是一个安全关键的序贯决策问题,因此是强化学习(RL)的天然试验平台。然而,现有的电力系统RL环境往往范围狭窄,且受限于基于CPU的仿真工作流,难以进行大规模评估。我们提出PowerZooJax,一个基于JAX的电力系统运行RL基准套件。它提供五个约束马尔可夫决策过程任务,涵盖发电、输电、配电、分布式能源和数据中心微电网。通过将潮流计算、经济调度、市场出清和设备动态重写…
- 资讯公开站rss_arxiv_cs_ai
有效的大语言模型微调是否必须依赖人类可读文本?
Is Human-Readable Text Necessary for Effective LLM Fine-Tuning?
arXiv:2609.35868v1 公告类型:新 摘要:有效微调大语言模型是否必须依赖人类可读性?我们研究模型条件化的训练表示能否在无需人类可读文本形式的情况下保持或提升适配效用。我们提出 Desired-Update-Aligned Synthetic Data(DASA),利用冻结参考模型的激活梯度反馈来指导连续合成输入嵌入的优化。受激活梯度在局部风险降低中作用的启发,DASA 针对有用的适配更新,而非源文本重建或语言流畅性。所得…
- 资讯公开站rss_arxiv_cs_ai
面向资源受限边缘设备可靠推理的神经符号路由
Neurosymbolic Routing for Reliable Reasoning on Resource-Constrained Edge Devices
在边缘硬件上运行语言模型可在无网络连接的情况下提供私密、低延迟的推理,但适合此类设备的小模型在计算机本应擅长的任务(如算术、代数和形式逻辑问题)上并不可靠。我们认为这种不可靠性在很大程度上是可以避免的。许多看似需要推理的查询实际上是结构确定性的,可以用快速且精确的符号方法求解。因此,强迫概率模型去近似它们只会牺牲准确性和能耗而收益甚微。我们提出一种神经符号路由器,对每个传入查询进行分类,并将其分派给成本最低的正确求解器:将结构化任务发送…
- 资讯公开站Digitimes
大同集团目标180天部署AI数据中心
Tatung targets 180-day AI data center deployment
为展示AI基础设施竞争力,大同集团9月30日举办“Power Ready for AI”论坛,发布基于“CUBE AI Data Center”概念的模块化AI数据中心,集成电力、冷却、监控与计算。
- 资讯公开站Semiconductor Engineering
E系列GPU IP:迈向融合加速的第一步
E-Series GPU IP: The First Step Towards Converged Acceleration
在单一灵活架构和可编程软件栈上融合图形、计算与AI,每核在1GHz下提供高达32 TOPS Int8算力,并可同时运行图形和AI工作负载。
- 资讯公开站Digitimes
专栏:AI改变平衡,内存与逻辑碰撞
Column: Memory and logic collide as AI shifts the balance
随着HBM成本接近GPU-HBM CoWoS封装的一半,内存还能被视为AI计算的被动组件吗?答案要回到“内存墙”,这一挑战塑造了近半个世纪的半导体发展。
- 论文公开站arXiv
STEPQuant:Delta规则循环状态量化中误差何时何地重要
STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization
线性注意力用固定大小循环状态替代增长的KV缓存,但这些持久状态在并发服务下可能成为内存瓶颈。直接低精度量化循环状态常导致严重精度下降。本文提出STEPQuant。
- 资讯公开站Digitimes
Meta与Firmus达成协议,在东南亚获取AI算力
Meta reaches agreements with Firmus on securing AI compute in Southeast Asia
Meta与澳大利亚新型云服务商Firmus达成多项协议,后者将在其东南亚AI工厂租赁GPU算力,支持Meta在亚太地区的AI扩张。
- 资讯公开站Semiconductor Engineering
芯片产业技术论文汇总:9月29日
Chip Industry Technical Paper Roundup: Sept. 29
涵盖A7 CFET与A10 NSFET对比、晶圆级亚5nm MoS₂晶体管、3D HI多千瓦供电、铜微结构与TSV残余应力、GPU Rowhammer攻击、门级RTL木马定位、LLM推理中HBM与主机内存并发访问等。
- 资讯公开站Digitimes
韩国AI算力紧张冲击研究人员与初创企业
South Korea AI compute squeeze hits researchers and startups
韩国AI行业面临日益严重的GPU短缺,研究机构和小企业难以获得足够的算力用于模型训练和部署。随着政府支持结束,研究团队发现训练更难持续,初创企业也受到挤压。
- 论文公开站arXiv
面向智能体强化学习高效压缩的KV-streams
KV-streams for Efficient Compaction in Agentic Reinforcement Learning
扩展智能体LLM的时域受限于需将越来越长的上下文轨迹装入GPU内存。上下文压缩是缓解该问题最常用的机制,可在给定轨迹下保持GPU内存恒定,但多数压缩策略存在不足。
- 资讯公开站gnews_nvidia_supply
英伟达可能减少每颗GPU的HBM用量,但总体采购量增加
Nvidia Could Use Less HBM Per GPU—and Buy More Overall - drrobertcastellano.substack.com
- 资讯公开站gnews_nvidia_supply
英伟达 B200 价格达每小时 8.01 美元,GPU 成本走势分化
Nvidia B200 Price Hits $8.01/Hr as GPU Costs Diverge - shattered.io
- 论文公开站arXiv
伸缩式语言模型
Telescopic Language Models
一个已部署的语言模型常需服务多种算力预算,而每个预算点仍需单独训练或压缩;本文训练伸缩式语言模型(TLM)作为连续体,由随机前缀监督的嵌套容量Transformer构成。