- 论文公开站arXiv
VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning
Self-improving policies continually expose new failure patterns, changing what their judges must be able to verify. However, current fixed judges constrain both optimization feedback and the discovery of useful training …
- 论文公开站arXiv
AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model
Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, be…
- 论文公开站arXiv
Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?
Large language models (LLMs) can solve many narrow tasks, but querying them separately for millions of related instances can be prohibitively expensive. Can LLM agents autonomously create cheaper solutions for such workl…
- 资讯公开站rss_arxiv_cs_ai
SkillScriptBench: Benchmarking Self-Evolution of Executable Agent Skill Packages Beyond Markdown
arXiv:2610.04008v1 Announce Type: new Abstract: Executable Agent Skills combine natural-language instructions and scripts into reusable packages for LLM agents, and revising them requires fixing errors without breaking …
- 资讯公开站Digitimes
DIGITIMES Insight: AI servers will carry nearly twice as many CPUs per accelerator by 2027
Demand for agentic AI in commercial and consumer markets was rising rapidly, and the market appeal of consumer services such as Meta Muse was expanding. In agentic AI tasks, large language model (LLM) inference was mainl…
- 资讯公开站Ars Technica
OpenAI agents tried to hack Wikipedia tools and flooded it with traffic
- 论文公开站arXiv
TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts
We introduce TasteVal, a benchmark to evaluate the experimental research taste of frontier models. We define research taste as the ability to pick interesting problems to solve, design experiments, and interpret experime…
- 论文公开站arXiv
CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling
Open-source web agents are now strong enough to execute realistic browser tasks, but training them with reinforcement learning still depends on weak supervision: binary task success is too sparse for credit assignment, w…
- 论文公开站arXiv
MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents
Memory has become integral to the LLM agent ecosystem, supporting information retention and reuse across interactions. However, most existing agent memory systems construct memory in a query-agnostic manner, which can in…
- 论文公开站arXiv
Recursive Video In-Context Learning for Agentic Robot
LLM agents that orchestrate frozen vision-language-action (VLA) policies improve across episodes through text memory, which records what the agent did but not how the task is done. A demonstration video shows it, but fit…
- 论文公开站arXiv
One Figure, Every Canvas: Editable Flowchart Relayout via Agentic Pipeline
Pipeline figures in ML papers must be repurposed across many canvases, including paper columns, 16:9 slides, portrait posters, 1:1 social teasers, 9:16 phone previews. Each format imposes a different aspect ratio on the …
- 资讯公开站rss_arxiv_cs_ai
When Terminal-Agent Training Stalls: Demystifying Data Generation and Verification Challenge
arXiv:2610.02405v1 Announce Type: new Abstract: Using a frontier model like Claude Opus as a meta-agent to generate terminal tasks and verifiers for RL training is increasingly common. Yet a runnable Docker image and ex…
- 资讯公开站rss_arxiv_cs_ai
DeReAct: Decomposed Reasoning and Acting for Reliable AI Agents
arXiv:2610.02351v1 Announce Type: new Abstract: ReAct-based agents typically rely on a single LLM policy to propose actions, interact with the environment, and decide when a task is complete. This coupling makes action …
- 资讯公开站rss_arxiv_cs_ai
Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents
arXiv:2610.02330v1 Announce Type: new Abstract: Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condition subsequent deci…
- 资讯公开站Digitimes
Anthropic IPO filing flags legal risks over AI agents
Anthropic said in its IPO filing that as artificial intelligence (AI) systems gain greater autonomous capabilities, the risks rise as well, while current law still does not clearly define who is liable when AI agents go …
- 资讯公开站Entrepreneur
How AI Agents Are Helping Companies Scale Their Marketing
Here's how AI agents can help companies scale marketing by combining skills, tools and context into an agentic marketing operating model.
- 资讯公开站Ars Technica
MCP for agent-to-agent comms may be the riskiest protocol you've never heard of
- 资讯公开站CNX Software
Save $70 on GEEKOM A5 Pro 2026 Edition mini PC during Prime Day Sale (Sponsored)
The GEEKOM A5 Pro 2026 Edition is a pocket-sized 0.47L mini PC available at a $70 discount for the company’s Prime Day Sale until October 7. Powered by the 6-core Ryzen 5 7530U (up to 4.5GHz), it handles browsing, …
- 资讯公开站MIT Technology Review
Connecting AI agents to enterprise knowledge
For all the data that AI systems continually amass and analyze, enterprise AI agents often suffer from a curious shortcoming: a lack of knowledge. More than data, knowledge is the understanding of what the data means in …
- 资讯公开站MIT Technology Review
Bringing predictive analytics to the agentic AI era
In 2026, the question for enterprise AI is no longer whether predictive models can outperform statistical forecasts—that argument is settled. The big question now is how to enable predictive systems to act on their own c…
- 资讯公开站EE Times
GPT-Synopsys Combines IC Design EDA with Agentic AI
- 资讯公开站Semiconductor Engineering
The Agentic AI Super Cycle
- 资讯公开站Digitimes
DeepSeek Harness challenges Agent lock-in with Claude Code Mods bridge and open plugin architecture
DeepSeek has released a new version of DeepSeek Harness, adding an experimental compatibility layer for Anthropic's Claude Code Mods and sharpening its broader effort to build an open, highly extensible software layer fo…
- 论文公开站arXiv
From Mixing to Tearing: Graph Decomposition in Decentralized Optimization via Message Passing
We study the minimization of sums of smooth strongly convex functions over undirected graphs, with each function held by one agent and communication restricted to neighbors in the graph. Existing decentralized methods, w…
- 论文公开站arXiv
4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes
We introduce 4DCodeBench, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs. To accomplish this, agents must translate visu…