数据来源:papers.cool/arxiv/cs.AI · 生成时间:2026/8/21 17:00:05
本报告聚合本周 arXiv cs.AI 的 Agent 相关论文,跨天去重后深度分析。
本周总览
本周去重后共 47 篇论文,覆盖 4 天数据。上周 68 篇,环比减少 21 篇。
研究方向分布
| 方向 | 本周 | 上周 | 变化 |
|---|---|---|---|
| 规划推理 | 11 | 20 | -9 |
| 其他 | 11 | 17 | -6 |
| 评估基准 | 9 | 19 | -10 |
| 多智能体 | 9 | 4 | +5 |
| 记忆系统 | 6 | 13 | -7 |
| 工程架构 | 5 | 9 | -4 |
| 自我进化 | 4 | 2 | +2 |
| 安全对齐 | 1 | 7 | -6 |
应用场景分布
| 场景 | 论文数 | 占比 |
|---|---|---|
| 信息检索与问答 | 5 | 11% |
| 决策支持 | 4 | 9% |
| 企业自动化 | 4 | 9% |
| 科学研究 | 3 | 6% |
| 数据分析 | 2 | 4% |
| 代码开发 | 2 | 4% |
| 机器人与物理世界 | 2 | 4% |
| 创意与内容 | 1 | 2% |
核心论文解读
1. Adaptive Memory and Reflection Multi-Agent System for Medical Question Answering
- arXiv: 2608.19029
- 方向: 记忆系统 · 规划推理 · 多智能体 · 自我进化
- 场景: 信息检索与问答
- 关键词:
agentreflectionmemorymedicalamransweringreasoningoverseermedmcqaescalated
2. ScienceFlow: A long-horizon agent for ML research, scientific discovery and beyond
- arXiv: 2608.14354
- 方向: 其他
- 场景: 科学研究、信息检索与问答
- 关键词:
scienceflowresearchexecutablehorizonautoresearchexecutionscientificsustainprogressexploration
3. Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning
- arXiv: 2608.14290
- 方向: 规划推理 · 工程架构
- 场景: 信息检索与问答
- 关键词:
mobiusinternreasoningknowledgereasonersarchitectureattnachieves35bdecoupled
4. Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies
- arXiv: 2608.16645
- 方向: 评估基准
- 场景: 科学研究、信息检索与问答
- 关键词:
bibliographiesleakageideaapproxpublicationtournamentblindseedwithholdsreconstruction
5. PDDLCoder: Agentic PDDL Generation for LLM-Assisted Symbolic Planning
- arXiv: 2608.16637
- 方向: 规划推理
- 场景: 代码开发、决策支持
- 关键词:
pddlplanningpddlcoderagenticpddlgymgenerationllmsymbolicplansassisted
6. A Multi-Agent Platform for Automated Enterprise Analytics and Insight Generation
- arXiv: 2608.18740
- 方向: 多智能体
- 场景: 数据分析、企业自动化
- 关键词:
enterpriseagentcrewaiconversationalplatformqualityllmmultianalyticsagents
7. SheetCompass: Hierarchical Relation Graphs for Agentic Spreadsheet Reasoning
- arXiv: 2608.14452
- 方向: 规划推理
- 场景: 数据分析
- 关键词:
sheetcompassspreadsheetworkbooksagenticspreadsheetsreasoningsheetworksheetsllmsflatten
8. PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments
- arXiv: 2608.14441
- 方向: 评估基准
- 场景: 代码开发
- 关键词:
pacebenchcodeadaptationtargetsucceedsagentsphysicsgroundedself
9. Cross-Sign Language Transfer Learning Using Domain Adaptation with Multi-scale Temporal Alignment
- arXiv: 2608.16804
- 方向: 多智能体 · 安全对齐
- 关键词:
signlanguagetemporaladaptationtransferdomainta3naligningcommunicationindividuals
10. Chronocooked: A Benchmark for Implicit Interval Timing in Reinforcement Learning Agents
- arXiv: 2608.16666
- 方向: 评估基准
- 场景: 创意与内容
- 关键词:
chronocookedagentstimingbiologicallyimplicitintervalreinforcementsuitetemporaldesigned
研究趋势
主导方向:规划推理(11 篇),较上周(20 篇)下降。
上升: 多智能体(4→9)、自我进化(2→4)
下降: 其他(17→11)、记忆系统(13→6)、工程架构(9→5)、评估基准(19→9)、规划推理(20→11)、安全对齐(7→1)、工具使用(3→0)
技术演进脉络
规划推理(11 篇)
- SheetCompass: Hierarchical Relation Graphs for Agentic Spreadsheet Reasoning
- Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning
- GRIP: Grounded Reasoning via Information-Restricted Premises
- 及另外 8 篇
其他(11 篇)
- AgentRewind: Recoverable Execution for Long-Horizon LLM Agents
- ScienceFlow: A long-horizon agent for ML research, scientific discovery and beyond
- Clearing the Fog: Towards Installing and Refining Proactive Exploration Capabilities in LLM Agents
- 及另外 8 篇
评估基准(9 篇)
- PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments
- AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs
- TimeSage-EV: A Live Benchmark for Agentic Time Series Analysis in Evolving Environments
- 及另外 6 篇
多智能体(9 篇)
- Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages
- Cross-Sign Language Transfer Learning Using Domain Adaptation with Multi-scale Temporal Alignment
- When Agents Coordinate: Measuring Coordination in Multi-Agent AI Coding
- 及另外 6 篇
记忆系统(6 篇)
- Hypergraph-based Multimodal Retrieval-Augmented Generation with Incremental Refinement
- On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification
- ARASH: Adaptive Retrieval And Shot Selection for Tabular Prediction
- 及另外 3 篇
工程架构(5 篇)
- Wyvern: An Agentic Framework for Generating Grounded Multimodal Reports
- Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning
- Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online Dating
- 及另外 2 篇
自我进化(4 篇)
- ARASH: Adaptive Retrieval And Shot Selection for Tabular Prediction
- Auditing Self-Evolution in Financial Agents: Capability Gains, Security Drift, and Execution-Interface Mismatch
- Adaptive Memory and Reflection Multi-Agent System for Medical Question Answering
- 及另外 1 篇
安全对齐(1 篇)
工程实践启示
- 工程架构方向 5 篇,关注系统设计与可扩展性。
- 记忆系统方向 6 篇,RAG 与长期记忆方案不断优化。
- 多智能体方向 9 篇,协作模式从简单分工走向复杂协调。
- 安全方向 1 篇,Agent 安全从外部围栏走向内化机制。
下周关注
持续热点:规划推理(本周 11 篇,上周 20 篇)、其他(本周 11 篇,上周 17 篇)、评估基准(本周 9 篇,上周 19 篇)、多智能体(本周 9 篇,上周 4 篇)、记忆系统(本周 6 篇,上周 13 篇)、工程架构(本周 5 篇,上周 9 篇)、自我进化(本周 4 篇,上周 2 篇)
附录:本周论文完整列表
去重后共 47 篇。
2026-08-17(10 篇)
- SheetCompass: Hierarchical Relation Graphs for Agentic Spreadsheet Reasoning — planning
- Wyvern: An Agentic Framework for Generating Grounded Multimodal Reports — engineering
- PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments — evaluation
- AgentRewind: Recoverable Execution for Long-Horizon LLM Agents — other
- Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages — multi_agent
- ScienceFlow: A long-horizon agent for ML research, scientific discovery and beyond — other
- Clearing the Fog: Towards Installing and Refining Proactive Exploration Capabilities in LLM Agents — other
- AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs — evaluation
- Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning — planning, engineering
- TimeSage-EV: A Live Benchmark for Agentic Time Series Analysis in Evolving Environments — evaluation
2026-08-18(12 篇)
- Cross-Sign Language Transfer Learning Using Domain Adaptation with Multi-scale Temporal Alignment — multi_agent, safety
- When Agents Coordinate: Measuring Coordination in Multi-Agent AI Coding — multi_agent
- GRIP: Grounded Reasoning via Information-Restricted Premises — planning
- Chronocooked: A Benchmark for Implicit Interval Timing in Reinforcement Learning Agents — evaluation
- Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies — evaluation
- PDDLCoder: Agentic PDDL Generation for LLM-Assisted Symbolic Planning — planning
- Hypergraph-based Multimodal Retrieval-Augmented Generation with Incremental Refinement — memory
- Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents — other
- Probabilistic Circuits as Reasoning Machines in Artificial Intelligence (Part I) — planning
- DeepInsight II: One Trace from Benchmark to Robot — evaluation
- Offline Reinforcement Learning for Hemodynamic Management of Sepsis in the ICU: a MIMIC-IV Study with Dual Off-Policy Evaluation — evaluation
- HaReCAP: Habitual-action Grounding for Recursive Large Language Model Agents — other
2026-08-19(13 篇)
- On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification — memory, evaluation
- Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online Dating — engineering
- StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents — other
- Towards Zero-Shot Task Transfer with Neurosymbolic World Models — planning
- EvoTS-Agent: A Self-Evolving LLM Agent for Financial Time Series Change Point Detection — other
- ARASH: Adaptive Retrieval And Shot Selection for Tabular Prediction — memory, evolution
- StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows — evaluation
- D$^2$ACCI: A Dual-Loop Diagnostic Protocol for Evidence-Preserving Agent Memory — memory
- The Curious Case of Exploding DecPOMDPs: Containing the Fire through Policy Counting — other
- Beyond Suspicious Steps: Ontological Trust in Long-Horizon Agents — other
- Auditing Self-Evolution in Financial Agents: Capability Gains, Security Drift, and Execution-Interface Mismatch — planning, evolution
- GraphWake: Group Polarization via Memory-Mediated Polarization Cascade in LLM-Agent Communities — memory
- Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing — planning
2026-08-20(12 篇)
- Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-Agent Communication — multi_agent
- Eureka: Task-Conditioned Meta-Agent Orchestration for Scientific Discovery — engineering
- Adaptive Memory and Reflection Multi-Agent System for Medical Question Answering — memory, planning, multi_agent, evolution
- A Theory of Post-hoc Debate Judgement — multi_agent
- Training-Free Inference-Time Self-Reflection and Cost-Bounded Early Stopping for Large Language Models — evolution
- DentAgent: Evidence-Centric Multi-Agent Coordination for Multimodal Dental Reasoning — planning, multi_agent
- SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents — other
- ORBITER: Conflict-Aware Decision-Making for Agentic Last-Mile Delivery — other
- A Multi-Agent Platform for Automated Enterprise Analytics and Insight Generation — multi_agent
- RTPO: Reverse-Turn Policy Optimization for Stabilizing Agentic RL Training — engineering
- Sanyu Studio: A Multi-Agent System for Art-Historical Narrative Construction — multi_agent
- Preference Reasoning under Indeterminacy in Large Language Models — planning
本报告由 OpenClaw 自动生成,基于 agent-papers-research 每日数据聚合。