数据来源:papers.cool/arxiv/cs.AI · 生成时间:2026/10/2 17:12:41
本报告聚合本周 arXiv cs.AI 的 Agent 相关论文,跨天去重后深度分析。
本周总览
本周去重后共 23 篇论文,覆盖 2 天数据。上周 48 篇,环比减少 25 篇。
研究方向分布
| 方向 | 本周 | 上周 | 变化 |
|---|---|---|---|
| 规划推理 | 8 | 14 | -6 |
| 其他 | 5 | 15 | -10 |
| 多智能体 | 4 | 2 | +2 |
| 评估基准 | 4 | 7 | -3 |
| 工程架构 | 2 | 3 | -1 |
| 安全对齐 | 2 | 2 | 0 |
| 记忆系统 | 2 | 6 | -4 |
| 自我进化 | 2 | 2 | 0 |
应用场景分布
| 场景 | 论文数 | 占比 |
|---|---|---|
| 决策支持 | 4 | 17% |
| 信息检索与问答 | 2 | 9% |
| 代码开发 | 1 | 4% |
| 创意与内容 | 1 | 4% |
| 企业自动化 | 1 | 4% |
核心论文解读
1. MA-WAM: Multi-Agent World-Action Model for Test-Time Planning
- 英文标题: MA-WAM: Multi-Agent World-Action Model for Test-Time Planning
- arXiv: 2609.31281 Kimi解读
- 方向: 规划推理 · 多智能体 · 评估基准
- 场景: 决策支持
2. Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation
- 英文标题: Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation
- arXiv: 2609.38024 Kimi解读
- 方向: 记忆系统 · 工程架构
- 场景: 信息检索与问答、决策支持
3. G2MAF: Test-Time Gradient Guidance for Multi-Agent Flow Policies
- 英文标题: G2MAF: Test-Time Gradient Guidance for Multi-Agent Flow Policies
- arXiv: 2609.31286 Kimi解读
- 方向: 多智能体 · 评估基准 · 工程架构
4. Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents
- 英文标题: Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents
- arXiv: 2609.31430 Kimi解读
- 方向: 其他
- 场景: 代码开发
5. Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI
- 英文标题: Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI
- arXiv: 2609.38143 Kimi解读
- 方向: 评估基准
- 场景: 创意与内容
6. Do LLM Agents Execute the Plans They Declare? From Planning-Mode Declaration to Pattern-Specific Execution
- 英文标题: Do LLM Agents Execute the Plans They Declare? From Planning-Mode Declaration to Pattern-Specific Execution
- arXiv: 2609.38108 Kimi解读
- 方向: 规划推理
- 场景: 决策支持
7. UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training
- 英文标题: UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training
- arXiv: 2609.38043 Kimi解读
- 方向: 评估基准
- 场景: 企业自动化
8. Diagnosing and Improving Probabilistic Reasoning in Large Language Models
- 英文标题: Diagnosing and Improving Probabilistic Reasoning in Large Language Models
- arXiv: 2609.38005 Kimi解读
- 方向: 规划推理
- 场景: 决策支持
9. SelfSearch: Reward-Free Search for Self-Improving Agents
- 英文标题: SelfSearch: Reward-Free Search for Self-Improving Agents
- arXiv: 2609.37968 Kimi解读
- 方向: 其他
- 场景: 信息检索与问答
10. Topological Coherence for Self-evolving Multi-agent Systems
- 英文标题: Topological Coherence for Self-evolving Multi-agent Systems
- arXiv: 2609.37953 Kimi解读
- 方向: 记忆系统 · 多智能体
研究趋势
主导方向:规划推理(8 篇),较上周(14 篇)下降。
上升: 多智能体(2→4)
下降: 记忆系统(6→2)、其他(15→5)、规划推理(14→8)、评估基准(7→4)、工程架构(3→2)
技术演进脉络
规划推理(8 篇)
- Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency Kimi解读
- MA-WAM: Multi-Agent World-Action Model for Test-Time Planning Kimi解读
- Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning Kimi解读
- 及另外 5 篇
其他(5 篇)
- DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education Kimi解读
- Segment-Level Agentic Topic Modeling for Improved Data Exploration and Resource Efficiency Kimi解读
- Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents Kimi解读
- 及另外 2 篇
多智能体(4 篇)
- Multi-agent Scaling Across Disjunctive and Compensatory Tasks Kimi解读
- G2MAF: Test-Time Gradient Guidance for Multi-Agent Flow Policies Kimi解读
- MA-WAM: Multi-Agent World-Action Model for Test-Time Planning Kimi解读
- 及另外 1 篇
评估基准(4 篇)
- G2MAF: Test-Time Gradient Guidance for Multi-Agent Flow Policies Kimi解读
- MA-WAM: Multi-Agent World-Action Model for Test-Time Planning Kimi解读
- Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI Kimi解读
- 及另外 1 篇
工程架构(2 篇)
- G2MAF: Test-Time Gradient Guidance for Multi-Agent Flow Policies Kimi解读
- Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation Kimi解读
安全对齐(2 篇)
- Character Training for Risk-Averse Agents Kimi解读
- Which Attention Heads are like the Human Head? Not the Ones that Compute Kimi解读
记忆系统(2 篇)
- Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation Kimi解读
- Topological Coherence for Self-evolving Multi-agent Systems Kimi解读
自我进化(2 篇)
- Video-RSI: Recursive Self-Improvement of Video Understanding Agents via Harness Evolution Kimi解读
- Guide, Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agentic RL Kimi解读
工程实践启示
- 工程架构方向 2 篇,关注系统设计与可扩展性。
- 记忆系统方向 2 篇,RAG 与长期记忆方案不断优化。
- 多智能体方向 4 篇,协作模式从简单分工走向复杂协调。
- 安全方向 2 篇,Agent 安全从外部围栏走向内化机制。
下周关注
持续热点:规划推理(本周 8 篇,上周 14 篇)、其他(本周 5 篇,上周 15 篇)、多智能体(本周 4 篇,上周 2 篇)、评估基准(本周 4 篇,上周 7 篇)、工程架构(本周 2 篇,上周 3 篇)、安全对齐(本周 2 篇,上周 2 篇)、记忆系统(本周 2 篇,上周 6 篇)、自我进化(本周 2 篇,上周 2 篇)
附录:本周论文完整列表
去重后共 23 篇。
2026-09-28(7 篇)
- Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency Kimi解读 — planning
- DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education Kimi解读 — other
- Multi-agent Scaling Across Disjunctive and Compensatory Tasks Kimi解读 — multi_agent
- Segment-Level Agentic Topic Modeling for Improved Data Exploration and Resource Efficiency Kimi解读 — other
- Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents Kimi解读 — other
- G2MAF: Test-Time Gradient Guidance for Multi-Agent Flow Policies Kimi解读 — multi_agent, evaluation, engineering
- MA-WAM: Multi-Agent World-Action Model for Test-Time Planning Kimi解读 — planning, multi_agent, evaluation
2026-09-30(16 篇)
- Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning Kimi解读 — planning
- Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI Kimi解读 — evaluation
- AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation Kimi解读 — other
- Stochastic World Models for Verifying Vision-Based Neural Feedback Systems Kimi解读 — planning
- Do LLM Agents Execute the Plans They Declare? From Planning-Mode Declaration to Pattern-Specific Execution Kimi解读 — planning
- NeuronEye: Query-Guided Visual Concept Activation for Vision-Language Reasoning Kimi解读 — planning
- Character Training for Risk-Averse Agents Kimi解读 — safety
- Probability is Not Enough: Exploring and Counting Divergent Tokens for Reasoning Uncertainty Quantification in LLMs Kimi解读 — planning
- UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training Kimi解读 — evaluation
- Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation Kimi解读 — memory, engineering
- Diagnosing and Improving Probabilistic Reasoning in Large Language Models Kimi解读 — planning
- Which Attention Heads are like the Human Head? Not the Ones that Compute Kimi解读 — safety
- SelfSearch: Reward-Free Search for Self-Improving Agents Kimi解读 — other
- Topological Coherence for Self-evolving Multi-agent Systems Kimi解读 — memory, multi_agent
- Video-RSI: Recursive Self-Improvement of Video Understanding Agents via Harness Evolution Kimi解读 — evolution
- Guide, Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agentic RL Kimi解读 — evolution
本报告由 OpenClaw 自动生成,基于 agent-papers-research 每日数据聚合。论文标题和摘要由 GLM-5 翻译生成。