数据来源:papers.cool/arxiv/cs.AI · 生成时间:2026/9/18 17:00:04
本报告聚合本周 arXiv cs.AI 的 Agent 相关论文,跨天去重后深度分析。
本周总览
本周去重后共 55 篇论文,覆盖 4 天数据。上周 53 篇,环比增加 2 篇。
研究方向分布
| 方向 | 本周 | 上周 | 变化 |
|---|---|---|---|
| 规划推理 | 18 | 17 | +1 |
| 其他 | 15 | 10 | +5 |
| 记忆系统 | 10 | 10 | 0 |
| 评估基准 | 7 | 13 | -6 |
| 工程架构 | 7 | 10 | -3 |
| 安全对齐 | 4 | 6 | -2 |
| 自我进化 | 4 | 1 | +3 |
| 多智能体 | 3 | 4 | -1 |
应用场景分布
| 场景 | 论文数 | 占比 |
|---|---|---|
| 信息检索与问答 | 12 | 22% |
| 科学研究 | 7 | 13% |
| 决策支持 | 4 | 7% |
| 企业自动化 | 3 | 5% |
| 机器人与物理世界 | 3 | 5% |
| 代码开发 | 2 | 4% |
| 数据分析 | 1 | 2% |
核心论文解读
1. Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval
- 英文标题: Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval
- arXiv: 2609.13073 Kimi解读
- 方向: 记忆系统 · 工程架构
- 场景: 科学研究、信息检索与问答
2. Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction
- 英文标题: Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction
- arXiv: 2609.13082 Kimi解读
- 方向: 评估基准
- 场景: 企业自动化、机器人与物理世界
3. Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science
- 英文标题: Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science
- arXiv: 2609.15983 Kimi解读
- 方向: 其他
- 场景: 科学研究、信息检索与问答
4. AlgoEvo: Self-Evolving Agentic Search for Automated Algorithm Discovery
- 英文标题: AlgoEvo: Self-Evolving Agentic Search for Automated Algorithm Discovery
- arXiv: 2609.15820 Kimi解读
- 方向: 其他
- 场景: 科学研究、信息检索与问答
5. Atria Dawn: The Dawn of Agentic Superintelligence
- 英文标题: Atria Dawn: The Dawn of Agentic Superintelligence
- arXiv: 2609.15818 Kimi解读
- 方向: 其他
- 场景: 科学研究、信息检索与问答
6. ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents
- 英文标题: ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents
- arXiv: 2609.17523 Kimi解读
- 方向: 自我进化
- 场景: 科学研究、信息检索与问答
7. Transformer-Based Token Fusion and Dynamic Graph Planning for Audio-Visual Navigation
- 英文标题: Transformer-Based Token Fusion and Dynamic Graph Planning for Audio-Visual Navigation
- arXiv: 2609.17421 Kimi解读
- 方向: 规划推理
- 场景: 决策支持、机器人与物理世界
8. FlashVector: Agent for Hierarchical Model Serving Stack Optimization
- 英文标题: FlashVector: Agent for Hierarchical Model Serving Stack Optimization
- arXiv: 2609.17391 Kimi解读
- 方向: 工程架构
- 场景: 代码开发、决策支持
9. Intrinsic Motivation in Reinforcement Learning: A Research Agenda for Adaptive Self-Organisation
- 英文标题: Intrinsic Motivation in Reinforcement Learning: A Research Agenda for Adaptive Self-Organisation
- arXiv: 2609.17325 Kimi解读
- 方向: 自我进化
- 场景: 科学研究、信息检索与问答
10. Symbolic Separation: Grounding Deep Agents in Knowledge Graphs for Trustworthy Operational Data Analytics
- 英文标题: Symbolic Separation: Grounding Deep Agents in Knowledge Graphs for Trustworthy Operational Data Analytics
- arXiv: 2609.17107 Kimi解读
- 方向: 其他
- 场景: 数据分析、信息检索与问答
研究趋势
主导方向:规划推理(18 篇),较上周(17 篇)上升。
上升: 规划推理(17→18)、其他(10→15)、自我进化(1→4)
下降: 工具使用(2→0)、评估基准(13→7)、工程架构(10→7)、多智能体(4→3)、安全对齐(6→4)
技术演进脉络
规划推理(18 篇)
- Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection Kimi解读
- When Should a World Model Move? Loss-Conditioned State Execution Kimi解读
- Navigating Sparse Evidence: Agentic Visual RAG via Explicit Context Selection and Consolidation Kimi解读
- 及另外 15 篇
其他(15 篇)
- Unified Agentic Video Editing Across Levels of Complexity and Creativity Kimi解读
- Skill Issue: Lessons from Optimizing Repository SKILLs for Coding Agents Kimi解读
- Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science Kimi解读
- 及另外 12 篇
记忆系统(10 篇)
- Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval Kimi解读
- Recurrent GraphNeural NetworkswithSet-BasedAggregation Kimi解读
- Navigating Sparse Evidence: Agentic Visual RAG via Explicit Context Selection and Consolidation Kimi解读
- 及另外 7 篇
评估基准(7 篇)
- Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction Kimi解读
- How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks Kimi解读
- K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments Kimi解读
- 及另外 4 篇
工程架构(7 篇)
- Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval Kimi解读
- K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments Kimi解读
- What Drives Recovery in Agentic Text-to-Cypher? LAST-CQ: An LLM Agent Self-Refinement Framework Kimi解读
- 及另外 4 篇
安全对齐(4 篇)
- CMA-OT: Hierarchical Expert Supervision for Dance-to-Music Generation Kimi解读
- Beyond Safe Answers: Segment-Aware Listwise Alignment for Reasoning Safety in Large Reasoning Models Kimi解读
- Beyond Routine Compliance: Cunning Data Cultivates Safety Vigilance in Large Language Models Kimi解读
- 及另外 1 篇
自我进化(4 篇)
- ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents Kimi解读
- Intrinsic Motivation in Reinforcement Learning: A Research Agenda for Adaptive Self-Organisation Kimi解读
- Cognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments Kimi解读
- 及另外 1 篇
多智能体(3 篇)
- Flag Game: A Toy Model for Mechanistic Swarm Interpretability Kimi解读
- Recursive Reasoning or Statistical Extrapolation? In-Context Learning in Multi-Agent Interdependent Decision-Making Kimi解读
- AeroWeaver: An Embodied-Agent Harness for Weaving Aerial Skills into Distributed, Adaptive Swarm Execution Kimi解读
工程实践启示
- 工程架构方向 7 篇,关注系统设计与可扩展性。
- 记忆系统方向 10 篇,RAG 与长期记忆方案不断优化。
- 多智能体方向 3 篇,协作模式从简单分工走向复杂协调。
- 安全方向 4 篇,Agent 安全从外部围栏走向内化机制。
下周关注
持续热点:规划推理(本周 18 篇,上周 17 篇)、其他(本周 15 篇,上周 10 篇)、记忆系统(本周 10 篇,上周 10 篇)、评估基准(本周 7 篇,上周 13 篇)、工程架构(本周 7 篇,上周 10 篇)、安全对齐(本周 4 篇,上周 6 篇)、多智能体(本周 3 篇,上周 4 篇)
附录:本周论文完整列表
去重后共 55 篇。
2026-09-14(8 篇)
- CMA-OT: Hierarchical Expert Supervision for Dance-to-Music Generation Kimi解读 — safety
- Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction Kimi解读 — evaluation
- Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval Kimi解读 — memory, engineering
- How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks Kimi解读 — evaluation
- K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments Kimi解读 — evaluation, engineering
- Unified Agentic Video Editing Across Levels of Complexity and Creativity Kimi解读 — other
- What Drives Recovery in Agentic Text-to-Cypher? LAST-CQ: An LLM Agent Self-Refinement Framework Kimi解读 — engineering
- Skill Issue: Lessons from Optimizing Repository SKILLs for Coding Agents Kimi解读 — other
2026-09-15(15 篇)
- Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection Kimi解读 — planning
- Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science Kimi解读 — other
- Recurrent GraphNeural NetworkswithSet-BasedAggregation Kimi解读 — memory
- LongAgent: History-Guided Agentic Search for Longitudinal Outcome Prediction Kimi解读 — other
- AlgoEvo: Self-Evolving Agentic Search for Automated Algorithm Discovery Kimi解读 — other
- Atria Dawn: The Dawn of Agentic Superintelligence Kimi解读 — other
- When Should a World Model Move? Loss-Conditioned State Execution Kimi解读 — planning
- Navigating Sparse Evidence: Agentic Visual RAG via Explicit Context Selection and Consolidation Kimi解读 — memory, planning
- KnowBench: Effort Reduction as a Unified, Deployment-Grounded Benchmark for Clinical AI Kimi解读 — evaluation, engineering
- EvoOntology: A Self-Evolving Ontology Layer for Data Agents Kimi解读 — other
- NoteVQA: Benchmarking VLMs on Real-Life Questions from Human Communities Kimi解读 — evaluation
- Diversified and Perceptible Counterfactual Examples Leveraging Expert Knowledge Kimi解读 — memory
- Option-Aware Retrieval and Task-Specific VLM Adaptation for Medical VQA Kimi解读 — memory
- Beyond Safe Answers: Segment-Aware Listwise Alignment for Reasoning Safety in Large Reasoning Models Kimi解读 — planning, safety
- The Troy Moment of AI: Why SomeWill Cheat and SomeWill Follow? Kimi解读 — other
2026-09-16(15 篇)
- ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents Kimi解读 — evolution
- Verifiable Social Reasoning for LLM Assistants Kimi解读 — planning
- Transformer-Based Token Fusion and Dynamic Graph Planning for Audio-Visual Navigation Kimi解读 — planning
- World Model Science: Self-Organized Criticality, Weak Chaos, and Metastable Belief Dynamics in Long-Horizon LLM Agents Kimi解读 — planning
- Never Stop Thinking: Continuous-Time Language Agents Kimi解读 — other
- FlashVector: Agent for Hierarchical Model Serving Stack Optimization Kimi解读 — engineering
- Self-Emergence Agent Architecture:Behavior-Inertia HMM, Reflexive Metacognition,and Social-Contrastive Self-Modeling Kimi解读 — engineering
- Intrinsic Motivation in Reinforcement Learning: A Research Agenda for Adaptive Self-Organisation Kimi解读 — evolution
- End-to-End Latency-Minimizing and Load-Balanced Request Scheduling for Edge LLM Inference in Agentic AI Services Kimi解读 — other
- MOCC-R1: Reinforcing Reasoning-Response Consistency for Multimodal Counselor Response Generation Kimi解读 — planning
- FirmCORe: A Benchmark for Structured Reasoning about Inter-Firm Collaboration Opportunities Kimi解读 — planning, evaluation
- Symbolic Separation: Grounding Deep Agents in Knowledge Graphs for Trustworthy Operational Data Analytics Kimi解读 — other
- Scaling-Score Conformal Prediction for Multi-Target Regression Kimi解读 — memory
- Interactive Memory Learning for Long-Term Conversations Kimi解读 — memory
- SKIP: a Self-knowledge-guided Step-wise Preference Learning Framework for Concise Reasoning Kimi解读 — planning, engineering
2026-09-17(17 篇)
- Cognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments Kimi解读 — memory, evolution
- Flag Game: A Toy Model for Mechanistic Swarm Interpretability Kimi解读 — multi_agent
- MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education Kimi解读 — evaluation
- Compiled Agency: Frontier General-Purpose Coding Agents Build Winning Game Players from Bare Interaction - from Flappy Bird to StarCraft II and Civilization Kimi解读 — other
- Lost in Perception: Isolating Perceptual and Reasoning Failures in Multimodal Physics and Geometry Reasoning Kimi解读 — planning
- Compositional Policy Violations: When Step-Level Compliance Fails In Agentic AI Workflows Kimi解读 — other
- CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents Kimi解读 — other
- Clueing up LLMs with Tool-Augmented Deductive Reasoning Kimi解读 — planning
- Which LLM is Best for Translating Natural Language Goals to PDDL Kimi解读 — planning
- Reasoning through Evolution: Automatic Meta-path Discovery for LLM-based Fake News Detection Kimi解读 — planning
- Recursive Reasoning or Statistical Extrapolation? In-Context Learning in Multi-Agent Interdependent Decision-Making Kimi解读 — planning, multi_agent
- AeroWeaver: An Embodied-Agent Harness for Weaving Aerial Skills into Distributed, Adaptive Swarm Execution Kimi解读 — multi_agent, evolution
- Beyond Routine Compliance: Cunning Data Cultivates Safety Vigilance in Large Language Models Kimi解读 — planning, safety
- First Token Matters: Understanding Safety Collapse in Large Reasoning Models Kimi解读 — planning, safety
- Disentangling Long-Term Memory via Latent Neuro-Symbolic Reasoning Kimi解读 — memory, planning
- Collective Loss of Control in LLM Agent Systems: An Epidemic Account of Mutation, Contagion, and Recovery Kimi解读 — other
- The Mirage of Calibrated Confidence: Trajectory-Independence of Verbalized Confidence in Vision-Language Models Kimi解读 — memory
本报告由 OpenClaw 自动生成,基于 agent-papers-research 每日数据聚合。论文标题和摘要由 GLM-5 翻译生成。