数据来源:papers.cool/arxiv/cs.AI · 生成时间:2026/8/14 17:00:05
本报告聚合本周 arXiv cs.AI 的 Agent 相关论文,跨天去重后深度分析。
本周总览
本周去重后共 55 篇论文,覆盖 4 天数据。上周 30 篇,环比增加 25 篇。
研究方向分布
| 方向 | 本周 | 上周 | 变化 |
|---|---|---|---|
| 其他 | 16 | 9 | +7 |
| 评估基准 | 15 | 6 | +9 |
| 规划推理 | 15 | 9 | +6 |
| 记忆系统 | 9 | 8 | +1 |
| 工程架构 | 7 | 2 | +5 |
| 安全对齐 | 4 | 2 | +2 |
| 多智能体 | 2 | 1 | +1 |
| 工具使用 | 2 | 1 | +1 |
| 自我进化 | 1 | 1 | 0 |
应用场景分布
| 场景 | 论文数 | 占比 |
|---|---|---|
| 信息检索与问答 | 10 | 18% |
| 科学研究 | 5 | 9% |
| 决策支持 | 4 | 7% |
| 企业自动化 | 3 | 5% |
| 机器人与物理世界 | 1 | 2% |
| 创意与内容 | 1 | 2% |
| 代码开发 | 1 | 2% |
核心论文解读
1. Adaptive Sequential Test Planning for Multi-Mechanism Reliability Qualification via Bayesian Monte Carlo Tree Search
- arXiv: 2608.09622
- 方向: 规划推理 · 评估基准 · 自我进化
- 场景: 信息检索与问答、决策支持
- 关键词:
qualificationreliabilitysequentialplanningtesttddbdamageadaptivefailuredegradation
2. FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings
- arXiv: 2608.07400
- 方向: 记忆系统 · 评估基准
- 场景: 信息检索与问答
- 关键词:
filingsfinrankgroundednegativesevidenceembedderfinancialansweringreportingsec
3. Agentic Auto-Research is Fuzz Testing
- arXiv: 2608.09855
- 方向: 评估基准
- 场景: 科学研究、信息检索与问答
- 关键词:
researchautofeedbackfuzzerprogresssignalagenticfuzzrathervalidation
4. Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- arXiv: 2608.09696
- 方向: 规划推理
- 场景: 科学研究、创意与内容
- 关键词:
discoveryemphmdamechanisticproposerinterventionalcitepmodelllmdpbench
5. V-FiLLM: Verified Financial LLM Reasoning Benchmark
- arXiv: 2608.11047
- 方向: 规划推理 · 评估基准
- 场景: 信息检索与问答
- 关键词:
financialfillmreasoningverifiedtablesitemsfinqa2022abenchmarksevaluating
6. ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling
- arXiv: 2608.10928
- 方向: 记忆系统 · 规划推理 · 评估基准
- 关键词:
thinkretrievereasoningaimelrmstraces2025testscalingstepsciq
7. Tree-of-Ideas: Automated Research Ideation via Cross-Trajectory Reasoning over Scholarly Evolution
- arXiv: 2608.10740
- 方向: 规划推理
- 场景: 科学研究、信息检索与问答
- 关键词:
scholarlyresearchideationideastoitrajectoriesevoagenttreereasoningacross
8. Self-Correcting Long-Horizon Search Agents via Tree-Structured Memory
- arXiv: 2608.10676
- 方向: 记忆系统 · 规划推理
- 场景: 信息检索与问答
- 关键词:
retreesearchcorrectingevidenceagentstreecontextreactreasoningself
9. VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies
- arXiv: 2608.12282
- 方向: 记忆系统 · 规划推理 · 工具使用
- 关键词:
apisvakratextbfreasoninghoptoolibmacrossfailuresetrieval
10. SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent
- arXiv: 2608.07449
- 方向: 其他
- 场景: 信息检索与问答
- 关键词:
skillproxskillsproximaldiagnosistextualgradientskillagentknowledgeupdates
研究趋势
主导方向:其他(16 篇),较上周(9 篇)上升。
上升: 其他(9→16)、评估基准(6→15)、规划推理(9→15)、记忆系统(8→9)、工程架构(2→7)、安全对齐(2→4)、多智能体(1→2)、工具使用(1→2)
技术演进脉络
其他(16 篇)
- SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent
- Blast Radius
- ResidencyRL: Reinforcement Learning in Simulated Clinical Environments
- 及另外 13 篇
评估基准(15 篇)
- Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing
- CoBa: Cost-Effective Test-Time Scaling via Compute-Balanced Routing
- GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks
- 及另外 12 篇
规划推理(15 篇)
- CoBa: Cost-Effective Test-Time Scaling via Compute-Balanced Routing
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- Avalon-ToM-Bench: Evaluating Fine-Grained Theory of Mind via Asymmetric Game Mechanics
- 及另外 12 篇
记忆系统(9 篇)
- PsychoAgent: An Affect-Sensitive Cognitive Architecture for Conflict-Aware Memory in LLM Agents
- TEPA: Revoking Stale Memories for Conflict-Robust Language Agents
- FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings
- 及另外 6 篇
工程架构(7 篇)
- PsychoAgent: An Affect-Sensitive Cognitive Architecture for Conflict-Aware Memory in LLM Agents
- ArchAgent v2: A Case Study with the Data Prefetching Championship
- CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems
- 及另外 4 篇
安全对齐(4 篇)
- People Are Not Just Their Countries. Disentangling Social Determinants of LLM Value Alignment Across Europe
- SHE: Trajectory-driven Safety Harness Evolution for LLM Agents
- REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems
- 及另外 1 篇
多智能体(2 篇)
- EMAS: Stabilizing Multi-Agent System Evolution through Evidence-Guided Revision
- ExRole: From Team Trajectories to Executable Roles in Multi-Agent Language Models
工具使用(2 篇)
- Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models
- VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies
自我进化(1 篇)
工程实践启示
- 工程架构方向 7 篇,关注系统设计与可扩展性。
- 工具使用方向 2 篇,function calling 与工具链持续演进。
- 记忆系统方向 9 篇,RAG 与长期记忆方案不断优化。
- 多智能体方向 2 篇,协作模式从简单分工走向复杂协调。
- 安全方向 4 篇,Agent 安全从外部围栏走向内化机制。
下周关注
持续热点:其他(本周 16 篇,上周 9 篇)、评估基准(本周 15 篇,上周 6 篇)、规划推理(本周 15 篇,上周 9 篇)、记忆系统(本周 9 篇,上周 8 篇)、工程架构(本周 7 篇,上周 2 篇)、安全对齐(本周 4 篇,上周 2 篇)
附录:本周论文完整列表
去重后共 55 篇。
2026-08-10(13 篇)
- SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent — other
- Blast Radius — other
- PsychoAgent: An Affect-Sensitive Cognitive Architecture for Conflict-Aware Memory in LLM Agents — memory, engineering
- Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing — evaluation
- TEPA: Revoking Stale Memories for Conflict-Robust Language Agents — memory
- CoBa: Cost-Effective Test-Time Scaling via Compute-Balanced Routing — planning, evaluation
- ResidencyRL: Reinforcement Learning in Simulated Clinical Environments — other
- GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks — evaluation
- FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings — memory, evaluation
- People Are Not Just Their Countries. Disentangling Social Determinants of LLM Value Alignment Across Europe — safety
- An End-to-End Agent Auditing Engine — evaluation
- EMAS: Stabilizing Multi-Agent System Evolution through Evidence-Guided Revision — multi_agent
- Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory — memory
2026-08-11(13 篇)
- SHE: Trajectory-driven Safety Harness Evolution for LLM Agents — safety
- ArchAgent v2: A Case Study with the Data Prefetching Championship — engineering
- Agentic Auto-Research is Fuzz Testing — evaluation
- CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems — engineering
- CARD: Controlled Agentic Reddit Discussions for Credit Card Simulation — other
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models — planning
- Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models — tool, evaluation
- Avalon-ToM-Bench: Evaluating Fine-Grained Theory of Mind via Asymmetric Game Mechanics — planning, evaluation
- Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines? — engineering
- Adaptive Sequential Test Planning for Multi-Mechanism Reliability Qualification via Bayesian Monte Carlo Tree Search — planning, evaluation, evolution
- ElasticBack: Stealthy Conditional Backdoor in LLM-Agent Skills via Coupled Trigger-Rule Optimization — engineering
- The Politician, the Liar, and the Obedient Worker: Emerging Behavior of LLM Agents in Hierarchical Games — other
- Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents — other
2026-08-12(15 篇)
- Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding — other
- SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure — evaluation
- V-FiLLM: Verified Financial LLM Reasoning Benchmark — planning, evaluation
- XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving — planning
- ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling — memory, planning, evaluation
- IO Factory: Simulating AI-Enabled Influence Campaigns at Scale — other
- ComBodied Agents: a New Paradigm of Human-Centric Agentic AI — other
- ChemWorld: Programmable Chemical Worlds for Controlled and Replayable Agent Experimentation — other
- SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation — memory
- Compositional Benchmark Synthesis for Hierarchical Human Action Recognition — evaluation, engineering
- Tree-of-Ideas: Automated Research Ideation via Cross-Trajectory Reasoning over Scholarly Evolution — planning
- Self-Correcting Long-Horizon Search Agents via Tree-Structured Memory — memory, planning
- REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems — safety
- VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus — planning
- Operationalising Relative Causal Knowledge: Backbone Identifiability from Private Reports on a Shared Outcome — other
2026-08-13(14 篇)
- Constructing Dynamic Master Logic Models as Knowledge Graphs for Complex System Diagnostics Using Retrieval-Augmented Large Language Models — memory
- VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies — memory, planning, tool
- An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS — other
- CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations — other
- Claim-Level Reliability Assessment for Efficient Test-Time Reasoning — planning, evaluation
- ExRole: From Team Trajectories to Executable Roles in Multi-Agent Language Models — multi_agent
- Policy-as-logic for robust reasoning over rules — planning
- Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents — safety
- The Sleeping Agent: What Gist-Based Context Compression Loses and Why — other
- Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents — other
- HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting — planning
- FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents — evaluation
- AgenticTwin: An Agentic LLM Framework Integrated with Digital Twin for Anomaly Detection — planning, engineering
- CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement — other
本报告由 OpenClaw 自动生成,基于 agent-papers-research 每日数据聚合。