全球科技每日监测AI 与全技术每日扫描

中文读懂 AI 与全技术今天发生了什么

邮箱轻订阅 · 免费开订每日精选技术情报:中文标题 → 要点 → 详情链。主题月卡加量 · 数据 API 可对接。

站内快照 · 国内可打开。外网原文可能无法访问。

  • 资讯公开站rss_arxiv_cs_ai

    Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization Limits

    arXiv:2609.38386v1 Announce Type: new Abstract: Concurrent autoregressive inference creates a fundamental interference problem: prefilling a newly arrived long prompt can delay tokens for requests that are already decoding. Fixed prefill chunks reduce this interference, but the best chunk size depends on the model, hardware, load, and latency objective. We introduce Decode-Latency Feedback Prefill (DLFP), a model-free controller that changes only prefill work that overlaps active decodes. After a guarded scheduling cycle, DLFP uses the observed interval as proportional feedback to resize the next prefill chunk; isolated prefills remain unrestricted. We implement DLFP in vLLM and evaluate it with open-loop Poisson arrivals, exact token accounting, raw request traces, and NVIDIA telemetry

    未知 tech_breakthrough source_collector