- 论文公开站arXiv
WorldSonus: Bringing Sound to Worlds
Recent advances in world models have enabled increasingly realistic visual synthesis. However, these generated environments remain largely silent. Bringing sound to world models poses three core challenges: real-time gen…
- 论文公开站arXiv
LBA-CBF: Rapidly Adaptive Safety Filters via Parallel Dynamics Inference
Control barrier functions (CBFs) certify commands through an assumed dynamics model, so an abrupt, unmeasured regime change can undermine the certificate exactly when safety matters most. We present Look-Back Adaptive Co…
- 论文公开站arXiv
QF3: Fast Flow RL with Filtered Q-Gradients
Flow policies have become a standard policy class for learning robot behaviors from demonstrations, but reinforcement learning is still critical for improving pre-trained flow policies or learning them from scratch throu…
- 资讯公开站rss_arxiv_cs_ai
SkillScriptBench: Benchmarking Self-Evolution of Executable Agent Skill Packages Beyond Markdown
arXiv:2610.04008v1 Announce Type: new Abstract: Executable Agent Skills combine natural-language instructions and scripts into reusable packages for LLM agents, and revising them requires fixing errors without breaking …
- 论文公开站arXiv
Paradee: Distilling Kokoro-82M into an 8M-Parameter Single-Voice Text-to-Speech Model
We distill Kokoro-82M, a widely used open text-to-speech model with 54 voices, into Paradee, an 8.07M-parameter model that speaks one of them. Paradee keeps Kokoro's architecture with much narrower layers, and each of it…
- 资讯公开站rss_arxiv_cs_ai
FlashSinkhorn 2: Block-Sparse Entropic Optimal Transport
arXiv:2610.02395v1 Announce Type: new Abstract: Streaming GPU solvers for entropic optimal transport (EOT), such as FlashSinkhorn, avoid storing the dense kernel but still evaluate all $n\times m$ point pairs in every S…
- 论文公开站arXiv
Simulation-Free Learning of Population Dynamics with Wasserstein Lagrangian Residuals
The dynamics of cells, organisms, and fluids are often modeled as probability distributions evolving over time. Reconstructing and extrapolating this evolution from unpaired snapshots requires assumptions about the under…
- 论文公开站arXiv
4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes
We introduce 4DCodeBench, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs. To accomplish this, agents must translate visu…
- 论文公开站arXiv
Less Decoder is More Encoder: Geometric Representation Learning from Novel View Synthesis
This paper examines the role of Novel View Synthesis (NVS) in geometric representation learning. In principle, NVS should reason about 3D scene structure, thereby enabling transferable multi-view geometric representation…
- 开源项目公开站GitHub
rohitg00/ai-engineering-from-scratch — Learn it. Build it. Ship it for others.
- 开源项目公开站GitHub
m3y54m/Embedded-Engineering-Roadmap — Comprehensive roadmap for aspiring Embedded Systems Engineers, featuring a curated
- 开源项目公开站GitHub
rust-embedded/awesome-embedded-rust — Curated list of resources for Embedded and Low-level development in the Rust progr
- 开源项目公开站GitHub
TianxingChen/Embodied-AI-Guide — [Lumina具身智能社区] 具身智能技术指南 Embodied-AI-Guide
- 资讯公开站rss_arxiv_cs_ai
K-Dense BYOK: An Open-Source AI Research Assistant That Runs Locally and Keeps a Hash-Chained Lab Notebook
arXiv:2610.00074v1 Announce Type: new Abstract: K-Dense BYOK (bring your own keys) is a free, open-source AI research assistant for scientists in any field that runs on the researcher's own computer. The researcher supp…
- 论文公开站arXiv
Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning
Step-size selection remains a central challenge in large-scale neural network optimization; conservative steps slow convergence, while aggressive steps can destabilize it. We combine \textbf{Z}ero-and-\textbf{F}irst-\tex…
- 论文公开站arXiv
Hierarchical Continuous Diffusion Language Models
Discrete diffusion language models offer a compelling alternative to autoregressive generation for tasks demanding bidirectional reasoning and global constraint satisfaction. Yet they share a structural bottleneck: when …
- 论文公开站arXiv
Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents
Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real…
- 论文公开站arXiv
One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars
3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bottleneck and show that the animation of pretrained avatar models can be…
- 论文公开站arXiv
图像分类器是高效的自监督视频表征学习器
Image Classifiers are Efficient Self-Supervised Video Representation Learners
提出 VideoMSN,一种用于视频高效自监督时空表征学习的掩码孪生网络框架。它不依赖重型3D架构或重建式自编码器,而是复用标准图像分类器处理无标注数据。
- 开源项目公开站GitHub
openbq-org/OpenBB——面向分析师、量化与AI智能体的开放数据平台
openbq-org/OpenBB — Open Data Platform for analysts, quants and AI agents.
- 资讯公开站rss_arxiv_cs_ai
PowerZooJax:面向强化学习的基于JAX的电力系统基准
PowerZooJax: A JAX-based Power System Benchmark for Reinforcement Learning
arXiv:2609.36052v1 公告类型:新 摘要:电力系统运行是一个安全关键的序贯决策问题,因此是强化学习(RL)的天然试验平台。然而,现有的电力系统RL环境往往范围狭窄,且受限于基于CPU的仿真工作流,难以进行大规模评估。我们提出PowerZooJax,一个基于JAX的电力系统运行RL基准套件。它提供五个约束马尔可夫决策过程任务,涵盖发电、输电、配电、分布式能源和数据中心微电网。通过将潮流计算、经济调度、市场出清和设备动态重写…
- 开源项目公开站GitHub
python279/stock-predict — 完全自动化的国际新闻抓取、分析和通知系统。系统每天自动抓取世界各地的权威新闻源,使用大模型进行深度分析,预测世界局势并给出投资建议。
- 论文公开站arXiv
STEPQuant:Delta规则循环状态量化中误差何时何地重要
STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization
线性注意力用固定大小循环状态替代增长的KV缓存,但这些持久状态在并发服务下可能成为内存瓶颈。直接低精度量化循环状态常导致严重精度下降。本文提出STEPQuant。
- 论文公开站arXiv
简化上下文机器人学习:面向操作任务的民主化方案
In-context Robot Learning Made Simple: A Democratized Recipe for Manipulation Tasks
本文研究机器人上下文学习(ICL),这一新兴范式使机器人能从视觉演示中推断并执行任务。尽管前景广阔,但问题本身仍定义不清:视觉演示同时传达动作轨迹和物体信息。
- 论文公开站arXiv
技能空间射击用于自主机器人策略改进
Skill-Space Shooting for Autonomous Robot Policy Improvement
部署在物理世界的机器人须能在遇到新情况和失败时超越初始训练进行改进。为使改进跨任务扩展,须有效利用经验而无需人类演示每次纠正。