数据来源:papers.cool/arxiv/cs.AI · 生成时间:2026/9/4 17:00:05
本报告聚合本周 arXiv cs.AI 的 Agent 相关论文,跨天去重后深度分析。
本周总览
本周去重后共 56 篇论文,覆盖 4 天数据。上周 59 篇,环比减少 3 篇。
研究方向分布
| 方向 | 本周 | 上周 | 变化 |
|---|---|---|---|
| 规划推理 | 21 | 14 | +7 |
| 其他 | 10 | 15 | -5 |
| 评估基准 | 10 | 8 | +2 |
| 安全对齐 | 8 | 7 | +1 |
| 工程架构 | 8 | 7 | +1 |
| 自我进化 | 6 | 4 | +2 |
| 多智能体 | 6 | 6 | 0 |
| 记忆系统 | 6 | 9 | -3 |
| 工具使用 | 1 | 1 | 0 |
应用场景分布
| 场景 | 论文数 | 占比 |
|---|---|---|
| 决策支持 | 8 | 14% |
| 科学研究 | 4 | 7% |
| 信息检索与问答 | 3 | 5% |
| 代码开发 | 3 | 5% |
| 机器人与物理世界 | 2 | 4% |
| 创意与内容 | 2 | 4% |
核心论文解读
1. CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI
- 英文标题: CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI
- arXiv: 2609.02459 Kimi解读
- 方向: 记忆系统 · 规划推理 · 工具使用 · 评估基准
- 场景: 决策支持
2. Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills
- 英文标题: Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills
- arXiv: 2609.02749 Kimi解读
- 方向: 其他
- 场景: 代码开发、科学研究、信息检索与问答
3. MAP: A Benchmark on Multimodal Accessibility Planning for Real World Places
- 英文标题: MAP: A Benchmark on Multimodal Accessibility Planning for Real World Places
- arXiv: 2608.28384 Kimi解读
- 方向: 规划推理 · 评估基准
- 场景: 决策支持
4. Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration
- 英文标题: Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration
- arXiv: 2608.28264 Kimi解读
- 方向: 多智能体 · 自我进化 · 工程架构
5. Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data
- 英文标题: Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data
- arXiv: 2608.31082 Kimi解读
- 方向: 规划推理 · 自我进化
- 场景: 信息检索与问答
6. Dual Process Motion Planning
- 英文标题: Dual Process Motion Planning
- arXiv: 2609.01260 Kimi解读
- 方向: 规划推理
- 场景: 决策支持、机器人与物理世界
7. UTP-Bench: Uncertainty-aware Travel Planning Benchmark
- 英文标题: UTP-Bench: Uncertainty-aware Travel Planning Benchmark
- arXiv: 2609.02421 Kimi解读
- 方向: 规划推理 · 评估基准
- 场景: 决策支持
8. Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds
- 英文标题: Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds
- arXiv: 2609.02302 Kimi解读
- 方向: 安全对齐 · 评估基准 · 工程架构
9. Codebook Agent: Amortized Topology Design for LLM Multi-Agent Systems
- 英文标题: Codebook Agent: Amortized Topology Design for LLM Multi-Agent Systems
- arXiv: 2609.02264 Kimi解读
- 方向: 多智能体
- 场景: 代码开发、创意与内容
10. APEx: Distillation of Agent Procedural Experience for Adaptive Deep Research Question Answering
- 英文标题: APEx: Distillation of Agent Procedural Experience for Adaptive Deep Research Question Answering
- arXiv: 2609.02253 Kimi解读
- 方向: 自我进化
- 场景: 科学研究、信息检索与问答
研究趋势
主导方向:规划推理(21 篇),较上周(14 篇)上升。
上升: 规划推理(14→21)、评估基准(8→10)、安全对齐(7→8)、工程架构(7→8)、自我进化(4→6)
下降: 其他(15→10)、记忆系统(9→6)
技术演进脉络
规划推理(21 篇)
- AcrossVAM1.0: Particle World Modeling for Text-Assisted Robot Video Prediction Kimi解读
- Learning to Use Tools: Reinforcement Learning for Tool-Integrated Mathematical Reasoning Kimi解读
- VERA-8B: Evidence-Grounded Audit Risk Reasoning from SEC Filings Kimi解读
- 及另外 18 篇
其他(10 篇)
- Logos: An Agent Harness on a Cross-Process Bus Kimi解读
- Prove2Me: An Open Collaborative Platform for Scaling Math Formalization Kimi解读
- RetailAgent: Structured Adverse Timing in Self-Conditioned Multimodal LLM Trading Agents Kimi解读
- 及另外 7 篇
评估基准(10 篇)
- MAP: A Benchmark on Multimodal Accessibility Planning for Real World Places Kimi解读
- AGENT-O: A Semantic Agent Card Framework for Interoperable and Governed Healthcare AI Agents Kimi解读
- LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering Kimi解读
- 及另外 7 篇
安全对齐(8 篇)
- VERA-8B: Evidence-Grounded Audit Risk Reasoning from SEC Filings Kimi解读
- OntoAligner-Ensemble: Voting-Based Fusion across Heterogeneous Ontology Alignment Techniques Kimi解读
- When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation Kimi解读
- 及另外 5 篇
工程架构(8 篇)
- AGENT-O: A Semantic Agent Card Framework for Interoperable and Governed Healthcare AI Agents Kimi解读
- Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration Kimi解读
- Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization Kimi解读
- 及另外 5 篇
自我进化(6 篇)
- EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses Kimi解读
- Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration Kimi解读
- Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data Kimi解读
- 及另外 3 篇
多智能体(6 篇)
- Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration Kimi解读
- HiRS-Agent: A Hierarchical Multi-Agent System for Reliable Long-Horizon Remote Sensing Task Solving Kimi解读
- EDGE: Error Dependency Graph-Guided Multi-Error Attribution in Multi-Agent LLM Systems Kimi解读
- 及另外 3 篇
记忆系统(6 篇)
- Regime-Aware Portfolio Management via Retrieval-Augmented LLM-Guided Expert Switching Kimi解读
- Measure Before You Manage: Evaluating Agent Working Memory in Coding Agents Kimi解读
- MedAgent-R1: Faithfulness-Aware Reinforcement Learning for Evidence-Grounded Medical Reasoning Kimi解读
- 及另外 3 篇
工具使用(1 篇)
工程实践启示
- 工程架构方向 8 篇,关注系统设计与可扩展性。
- 工具使用方向 1 篇,function calling 与工具链持续演进。
- 记忆系统方向 6 篇,RAG 与长期记忆方案不断优化。
- 多智能体方向 6 篇,协作模式从简单分工走向复杂协调。
- 安全方向 8 篇,Agent 安全从外部围栏走向内化机制。
下周关注
持续热点:规划推理(本周 21 篇,上周 14 篇)、其他(本周 10 篇,上周 15 篇)、评估基准(本周 10 篇,上周 8 篇)、安全对齐(本周 8 篇,上周 7 篇)、工程架构(本周 8 篇,上周 7 篇)、自我进化(本周 6 篇,上周 4 篇)、多智能体(本周 6 篇,上周 6 篇)、记忆系统(本周 6 篇,上周 9 篇)
附录:本周论文完整列表
去重后共 56 篇。
2026-08-31(12 篇)
- Logos: An Agent Harness on a Cross-Process Bus Kimi解读 — other
- AcrossVAM1.0: Particle World Modeling for Text-Assisted Robot Video Prediction Kimi解读 — planning
- Learning to Use Tools: Reinforcement Learning for Tool-Integrated Mathematical Reasoning Kimi解读 — planning
- Prove2Me: An Open Collaborative Platform for Scaling Math Formalization Kimi解读 — other
- VERA-8B: Evidence-Grounded Audit Risk Reasoning from SEC Filings Kimi解读 — planning, safety
- RetailAgent: Structured Adverse Timing in Self-Conditioned Multimodal LLM Trading Agents Kimi解读 — other
- MAP: A Benchmark on Multimodal Accessibility Planning for Real World Places Kimi解读 — planning, evaluation
- EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses Kimi解读 — evolution
- AGENT-O: A Semantic Agent Card Framework for Interoperable and Governed Healthcare AI Agents Kimi解读 — evaluation, engineering
- LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering Kimi解读 — evaluation
- Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration Kimi解读 — multi_agent, evolution, engineering
- Regime-Aware Portfolio Management via Retrieval-Augmented LLM-Guided Expert Switching Kimi解读 — memory
2026-09-01(14 篇)
- OntoAligner-Ensemble: Voting-Based Fusion across Heterogeneous Ontology Alignment Techniques Kimi解读 — safety
- Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data Kimi解读 — planning, evolution
- Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization Kimi解读 — engineering
- Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence Kimi解读 — planning
- Wrong Prediction, Right Answer: Recovering Evidence from Collapsed LLM Sequence Scores Kimi解读 — planning
- Measure Before You Manage: Evaluating Agent Working Memory in Coding Agents Kimi解读 — memory
- MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents Kimi解读 — other
- CAER: Causal Action Effect Reweighting for World Model Training Kimi解读 — planning
- HSRM: Hidden-State Reward Models for Test-Time Verification Kimi解读 — planning, evaluation
- SkillZip Pro: Execution-Aware Dynamic Compression of Progressively Loaded Skills for Self-Evolving Agents Kimi解读 — other
- Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models Kimi解读 — planning
- ATLAS: Dual-Horizon Diagnostic Evaluation for Industrial Tool-Use Agents Kimi解读 — evaluation
- MedAgent-R1: Faithfulness-Aware Reinforcement Learning for Evidence-Grounded Medical Reasoning Kimi解读 — memory, planning
- HiRS-Agent: A Hierarchical Multi-Agent System for Reliable Long-Horizon Remote Sensing Task Solving Kimi解读 — multi_agent
2026-09-02(12 篇)
- Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers Kimi解读 — other
- EvoSCM: Scientific Belief Revision Through Causal Model Evolution and Experimentation Kimi解读 — other
- When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation Kimi解读 — safety, evaluation
- Parsing the Stream: A Live Trace Model for Long-Horizon Agents and Their Observers Kimi解读 — other
- EDGE: Error Dependency Graph-Guided Multi-Error Attribution in Multi-Agent LLM Systems Kimi解读 — multi_agent
- Analog-DB: An Agent-First Analog Integrated Circuit Database, From Blocks to Systems Kimi解读 — evaluation
- Making Prospective Memory SLM-Shaped: Typed Intention Stores for Small-Model Agents Kimi解读 — memory
- Dual Process Motion Planning Kimi解读 — planning
- H2Table: Hierarchical Hypergraph-Enhanced Large Language Models for Complex Table Reasoning Kimi解读 — planning
- Jailbreaking Text-to-Image Models Through Cracks: Navigating Heterogeneous Safety Filters via Multi-Agent Debate Kimi解读 — multi_agent, safety
- Latent Recurrent Thoughts: Recurrent Refinement of Proposed Latents for Reasoning with Frozen LLMs Kimi解读 — planning
- ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning Kimi解读 — evolution
2026-09-03(18 篇)
- Discriminative World Models for Web Agents Kimi解读 — planning
- Large Language Models (LLMs) for Telecom Root Cause Analysis (RCA): A Structured Reasoning Framework for Evidence-Grounded Diagnosis Kimi解读 — planning, engineering
- SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment Kimi解读 — safety
- Measurement-Driven Sub-Network Selection for On-Premise Retrieval-Augmented Factory Agents Kimi解读 — memory
- Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems Kimi解读 — multi_agent, evolution
- Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills Kimi解读 — other
- CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI Kimi解读 — memory, planning, tool, evaluation
- UTP-Bench: Uncertainty-aware Travel Planning Benchmark Kimi解读 — planning, evaluation
- Contrastive Explanations in Quantitative Bipolar Argumentation Frameworks Kimi解读 — planning, engineering
- Diagnosing with Insights: Structured Analysis of Agent Failures via Behavioral Abstractions Kimi解读 — other
- SALA: Semantic-Aware Logical Alignment for Complex Reasoning in In-Context Learning Kimi解读 — planning, safety
- Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds Kimi解读 — safety, evaluation, engineering
- Codebook Agent: Amortized Topology Design for LLM Multi-Agent Systems Kimi解读 — multi_agent
- APEx: Distillation of Agent Procedural Experience for Adaptive Deep Research Question Answering Kimi解读 — evolution
- LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails Kimi解读 — safety
- Propose to Learn, Learn to Propose: Evaluability-Aware Assistance under Bounded Rationality Kimi解读 — planning
- PGPO: Potential-Guided Policy Optimization for Multi-Turn Agentic Tasks Kimi解读 — engineering
- PhoenixNest-Video: Evidence-Grounded Multimodal Agent Framework for Automated Video Interview Assessment Kimi解读 — engineering
本报告由 OpenClaw 自动生成,基于 agent-papers-research 每日数据聚合。论文标题和摘要由 GLM-5 翻译生成。