数据来源:papers.cool/arxiv/cs.AI · 生成时间:2026/8/28 17:00:05
本报告聚合本周 arXiv cs.AI 的 Agent 相关论文,跨天去重后深度分析。
本周总览
本周去重后共 48 篇论文,覆盖 4 天数据。上周 63 篇,环比减少 15 篇。
研究方向分布
| 方向 | 本周 | 上周 | 变化 |
|---|---|---|---|
| 规划推理 | 13 | 17 | -4 |
| 其他 | 10 | 16 | -6 |
| 评估基准 | 7 | 14 | -7 |
| 安全对齐 | 7 | 1 | +6 |
| 记忆系统 | 7 | 8 | -1 |
| 多智能体 | 5 | 9 | -4 |
| 自我进化 | 4 | 8 | -4 |
| 工程架构 | 4 | 7 | -3 |
| 工具使用 | 1 | 1 | 0 |
应用场景分布
| 场景 | 论文数 | 占比 |
|---|---|---|
| 企业自动化 | 3 | 6% |
| 决策支持 | 3 | 6% |
| 信息检索与问答 | 2 | 4% |
| 代码开发 | 2 | 4% |
| 科学研究 | 1 | 2% |
| 机器人与物理世界 | 1 | 2% |
| 数据分析 | 1 | 2% |
核心论文解读
1. SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning
- arXiv: 2608.23493
- 方向: 规划推理 · 自我进化 · 工程架构
- 场景: 决策支持
- 关键词:
srporeflectivereflectionpolicyselfteacheraimehorizonreasoninginternalizes
2. VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences
- arXiv: 2608.21357
- 方向: 评估基准
- 场景: 企业自动化
- 关键词:
vialsartifactssciencesvisualworkflowsscientistsinterpretinterpretationlifeimages
3. ReFrame: Evidence-Guided Test-Time Safety Alignment in Multimodal Large Language Models
- arXiv: 2608.21100
- 方向: 安全对齐 · 评估基准
- 关键词:
reframemultimodalsafetymllmsalignmentmllmutilityevidenceawarenessoversensitivity
4. ReWorld: An Interactive World Model with Long-Horizon Memory
- arXiv: 2608.23565
- 方向: 记忆系统 · 规划推理
- 关键词:
interactivereworldheadsworldmemorychunkhorizonwantsplacescache
5. EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards
- arXiv: 2608.23525
- 方向: 评估基准
- 场景: 科学研究
- 关键词:
earthversescientificagentshazardsearthevidenceansweracrossreproduciblesystems
6. Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty
- arXiv: 2608.23497
- 方向: 规划推理 · 安全对齐
- 关键词:
safetyreasoningrimsdpdirectionshiftsmisalignmentpenaltytuningharmful
7. SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL
- arXiv: 2608.24870
- 方向: 工程架构
- 场景: 决策支持
- 关键词:
spotokenpromptpolicystreamagenticactorwhitensalfworldasynchronous
8. CAFE: Self-Improving Search Agents Need Co-Evolving Feedback
- arXiv: 2608.24794
- 方向: 其他
- 场景: 信息检索与问答
- 关键词:
feedbackcafeagentsearchcriticimprovingagentscouplesrequestoutcome
9. Evidence Blindness in Direct Corpus Interaction: Persistent Navigation with AtlasNav
- arXiv: 2608.24764
- 方向: 其他
- 场景: 机器人与物理世界
- 关键词:
corpusatlasnavevidenceinteractionnavigationblindnessdciagenticdirectbudgets
10. Joint Optimization of Tool Creation and Use for Large Language Model Agents
- arXiv: 2608.24571
- 方向: 工程架构
- 场景: 决策支持
- 关键词:
toola3b30bcreationqwen3schemasinvokesmithwriteschema
研究趋势
主导方向:规划推理(13 篇),较上周(17 篇)下降。
上升: 安全对齐(1→7)
下降: 规划推理(17→13)、工程架构(7→4)、评估基准(14→7)、其他(16→10)、多智能体(9→5)、记忆系统(8→7)、自我进化(8→4)
技术演进脉络
规划推理(13 篇)
- ReWorld: An Interactive World Model with Long-Horizon Memory
- Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty
- SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning
- 及另外 10 篇
其他(10 篇)
- Don't Solve, Just Compare: Tiny Advisors for Runtime Intervention in LLM Agents
- Prime Agent: A Self-Improving RLM Harness
- SkillAlchemy: Open-World Agent Skill Creation
- 及另外 7 篇
评估基准(7 篇)
- VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences
- ReFrame: Evidence-Guided Test-Time Safety Alignment in Multimodal Large Language Models
- CellPath-Bench: A Multidimensional Benchmark for Whole-Slide Cellular Representations in Pathology Foundation Models
- 及另外 4 篇
安全对齐(7 篇)
- CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment
- ReFrame: Evidence-Guided Test-Time Safety Alignment in Multimodal Large Language Models
- Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty
- 及另外 4 篇
记忆系统(7 篇)
- ReWorld: An Interactive World Model with Long-Horizon Memory
- Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses
- Multi-Granularity Context-Enhanced RAG over Multimodal Knowledge Graphs
- 及另外 4 篇
多智能体(5 篇)
- Pivot-and-Station Multi-Agent Path Finding: Solvability, Complexity, and Algorithms
- SwarmWorld: Stigmergic technological evolution in societies of language-model agents
- ProgRouter: Online Progress-Guided Orchestration for Multi-Agent LLM Workflows under Quality-Cost Tradeoffs
- 及另外 2 篇
自我进化(4 篇)
- SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning
- MediSkill-Evo: Process-Constrained Self-Evolution for Evidence-Grounded Clinical Interaction
- Meta$^n$: Recursive Self-Improvement through Emergent Depth
- 及另外 1 篇
工程架构(4 篇)
- SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning
- SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL
- Joint Optimization of Tool Creation and Use for Large Language Model Agents
- 及另外 1 篇
工具使用(1 篇)
工程实践启示
- 工程架构方向 4 篇,关注系统设计与可扩展性。
- 工具使用方向 1 篇,function calling 与工具链持续演进。
- 记忆系统方向 7 篇,RAG 与长期记忆方案不断优化。
- 多智能体方向 5 篇,协作模式从简单分工走向复杂协调。
- 安全方向 7 篇,Agent 安全从外部围栏走向内化机制。
下周关注
持续热点:规划推理(本周 13 篇,上周 17 篇)、其他(本周 10 篇,上周 16 篇)、评估基准(本周 7 篇,上周 14 篇)、记忆系统(本周 7 篇,上周 8 篇)、多智能体(本周 5 篇,上周 9 篇)、自我进化(本周 4 篇,上周 8 篇)、工程架构(本周 4 篇,上周 7 篇)
附录:本周论文完整列表
去重后共 48 篇。
2026-08-24(5 篇)
- VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences — evaluation
- CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment — safety
- ReFrame: Evidence-Guided Test-Time Safety Alignment in Multimodal Large Language Models — safety, evaluation
- CellPath-Bench: A Multidimensional Benchmark for Whole-Slide Cellular Representations in Pathology Foundation Models — evaluation
- Don't Solve, Just Compare: Tiny Advisors for Runtime Intervention in LLM Agents — other
2026-08-25(11 篇)
- ReWorld: An Interactive World Model with Long-Horizon Memory — memory, planning
- Prime Agent: A Self-Improving RLM Harness — other
- EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards — evaluation
- Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty — planning, safety
- SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning — planning, evolution, engineering
- SkillAlchemy: Open-World Agent Skill Creation — other
- MediSkill-Evo: Process-Constrained Self-Evolution for Evidence-Grounded Clinical Interaction — evolution
- Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Learning — other
- Apodex 1.1: Scaling Agentic Intelligence for Complex Work — other
- Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data — planning
- Cognitive Profiling of LRMs' Reasoning Traces Using Bloom's Taxonomy — planning
2026-08-26(15 篇)
- Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses — memory
- SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL — engineering
- CAFE: Self-Improving Search Agents Need Co-Evolving Feedback — other
- Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought — planning
- StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing — safety
- Evidence Blindness in Direct Corpus Interaction: Persistent Navigation with AtlasNav — other
- Meta$^n$: Recursive Self-Improvement through Emergent Depth — evolution
- Parason: Revealing Subtask and Trial Parallelism in LLM Reasoning — planning
- Pivot-and-Station Multi-Agent Path Finding: Solvability, Complexity, and Algorithms — multi_agent
- PhysMLLMs: Spatial Priors for Unified Referring Segmentation and Grounded Reasoning of Images and Videos — planning
- Joint Optimization of Tool Creation and Use for Large Language Model Agents — engineering
- EviDx: Evidence-Aware Active Diagnosis with Scaffolded LLM Agents — other
- When "Must" Becomes "Maybe": Constraint Weakening in LLM Agent Workflows — other
- Discovering Adaptive Transmission Programs for Collective Innovation — evolution
- Neurosymbolic Alignment for Physiologically-Safe Clinical Language Models — safety
2026-08-27(17 篇)
- SwarmWorld: Stigmergic technological evolution in societies of language-model agents — multi_agent
- Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems — planning
- AsymSpec: Context-Asymmetric Speculative Decoding for Agentic LLMs — evaluation
- ProgRouter: Online Progress-Guided Orchestration for Multi-Agent LLM Workflows under Quality-Cost Tradeoffs — multi_agent
- Multi-Granularity Context-Enhanced RAG over Multimodal Knowledge Graphs — memory
- Quantitative Analysis of $ω$-Regular Robust MDPs — other
- LivingRAG: Augmenting Graph RAG with Experience — memory, planning
- Candidate supply and answer selection shape the value of LLM judging in multi-agent systems — multi_agent
- How Robust Are Automated Fact-Checking Systems? A Cross-Benchmark Evaluation — evaluation
- Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems — multi_agent
- LocalLSTC: A Long Short-Term Control Architecture for Locally Deployed GUI Agents — planning, engineering
- Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models — planning
- CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval — memory
- PonsRAG: A Pons-Inspired RAG Bridging Cognitive Islands for Coordinated Long Narrative Reasoning — memory, planning
- Training Alignment Auditors via Reinforcement Learning — safety
- Distance Is Not Enough: Forget-Retain Alignment Gap Predicts LLM Relearning Robustness — memory, safety
- Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents — tool, evaluation
本报告由 OpenClaw 自动生成,基于 agent-papers-research 每日数据聚合。