数据来源:papers.cool/arxiv/cs.AI · 生成时间:2026/9/11 17:00:05
本报告聚合本周 arXiv cs.AI 的 Agent 相关论文,跨天去重后深度分析。
本周总览
本周去重后共 45 篇论文,覆盖 3 天数据。上周 66 篇,环比减少 21 篇。
研究方向分布
| 方向 | 本周 | 上周 | 变化 |
|---|---|---|---|
| 规划推理 | 14 | 23 | -9 |
| 评估基准 | 13 | 11 | +2 |
| 记忆系统 | 9 | 7 | +2 |
| 其他 | 9 | 15 | -6 |
| 工程架构 | 6 | 8 | -2 |
| 安全对齐 | 4 | 8 | -4 |
| 多智能体 | 3 | 7 | -4 |
| 工具使用 | 2 | 1 | +1 |
应用场景分布
| 场景 | 论文数 | 占比 |
|---|---|---|
| 信息检索与问答 | 6 | 13% |
| 决策支持 | 5 | 11% |
| 企业自动化 | 4 | 9% |
| 科学研究 | 3 | 7% |
| 机器人与物理世界 | 1 | 2% |
核心论文解读
1. SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
- 英文标题: SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
- arXiv: 2609.09113 Kimi解读
- 方向: 其他
- 场景: 科学研究、企业自动化、信息检索与问答
2. Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models
- 英文标题: Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models
- arXiv: 2609.05381 Kimi解读
- 方向: 记忆系统 · 规划推理 · 评估基准
3. MeClear: Cooperative Game-Theoretic Attribution and Risk-Aware Memory Clearance for Long-Horizon LLM Agents
- 英文标题: MeClear: Cooperative Game-Theoretic Attribution and Risk-Aware Memory Clearance for Long-Horizon LLM Agents
- arXiv: 2609.09115 Kimi解读
- 方向: 记忆系统 · 多智能体 · 安全对齐
4. GoAnt: Quality-Diversity Multi-Agent Search for Alpha Factor Discovery in Market Microstructure Data
- 英文标题: GoAnt: Quality-Diversity Multi-Agent Search for Alpha Factor Discovery in Market Microstructure Data
- arXiv: 2609.08719 Kimi解读
- 方向: 多智能体
- 场景: 科学研究、信息检索与问答
5. CLAMP: Constrained Decoding for Vision-Language Embodied Planning
- 英文标题: CLAMP: Constrained Decoding for Vision-Language Embodied Planning
- arXiv: 2609.08602 Kimi解读
- 方向: 规划推理
- 场景: 决策支持、机器人与物理世界
6. ConvMem: Convolutional Memory for Long-Context Reasoning
- 英文标题: ConvMem: Convolutional Memory for Long-Context Reasoning
- arXiv: 2609.10441 Kimi解读
- 方向: 记忆系统 · 规划推理
- 场景: 信息检索与问答
7. From Symbolic Perception to Logical Deduction: A Framework for Guiding Language Models in Geometric Reasoning
- 英文标题: From Symbolic Perception to Logical Deduction: A Framework for Guiding Language Models in Geometric Reasoning
- arXiv: 2609.10335 Kimi解读
- 方向: 规划推理 · 评估基准 · 工程架构
8. Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe
- 英文标题: Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe
- arXiv: 2609.05395 Kimi解读
- 方向: 工具使用 · 评估基准
9. Large Language Models for HVAC Operations in Building Energy Systems: A Critical Review of Methods, Applications, and Deployment Readiness
- 英文标题: Large Language Models for HVAC Operations in Building Energy Systems: A Critical Review of Methods, Applications, and Deployment Readiness
- arXiv: 2609.05314 Kimi解读
- 方向: 工程架构
- 场景: 企业自动化
10. Commonsense Reasoning in Computer Vision: Foundations, Recent Advancements, and Future Directions
- 英文标题: Commonsense Reasoning in Computer Vision: Foundations, Recent Advancements, and Future Directions
- arXiv: 2609.05257 Kimi解读
- 方向: 规划推理
- 场景: 信息检索与问答
研究趋势
主导方向:规划推理(14 篇),较上周(23 篇)下降。
上升: 评估基准(11→13)、记忆系统(7→9)、工具使用(1→2)
下降: 其他(15→9)、规划推理(23→14)、安全对齐(8→4)、自我进化(6→0)、工程架构(8→6)、多智能体(7→3)
技术演进脉络
规划推理(14 篇)
- Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models Kimi解读
- GUT: Quantifying and Optimizing the Reasoning Uncertainty of LLMs via Graph Complexity Kimi解读
- Commonsense Reasoning in Computer Vision: Foundations, Recent Advancements, and Future Directions Kimi解读
- 及另外 11 篇
评估基准(13 篇)
- Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe Kimi解读
- Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models Kimi解读
- Testing Interchangeability in LLM Agent Teams Kimi解读
- 及另外 10 篇
记忆系统(9 篇)
- Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models Kimi解读
- Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability Kimi解读
- MeClear: Cooperative Game-Theoretic Attribution and Risk-Aware Memory Clearance for Long-Horizon LLM Agents Kimi解读
- 及另外 6 篇
其他(9 篇)
- CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents Kimi解读
- Trace2Tower: Transition-Aware EigenTrace Induction of Multi-Level Skills for LLM Agents Kimi解读
- Substrate-Aware AI Agents: Execution Context as a First-Class Input Kimi解读
- 及另外 6 篇
工程架构(6 篇)
- Large Language Models for HVAC Operations in Building Energy Systems: A Critical Review of Methods, Applications, and Deployment Readiness Kimi解读
- SkillAdam: Stable and Efficient Skill Evolution for Agents Kimi解读
- AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems Kimi解读
- 及另外 3 篇
安全对齐(4 篇)
- MeClear: Cooperative Game-Theoretic Attribution and Risk-Aware Memory Clearance for Long-Horizon LLM Agents Kimi解读
- Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training Kimi解读
- Beyond Surface Imitation: Contrastive Modeling for Reasoning Path Alignment in Multimodal In-Context Learning Kimi解读
- 及另外 1 篇
多智能体(3 篇)
- CABAL: Multi-Agent Simulacra for Tracing the Effects of Collusive Bidding in Peer Review Kimi解读
- MeClear: Cooperative Game-Theoretic Attribution and Risk-Aware Memory Clearance for Long-Horizon LLM Agents Kimi解读
- GoAnt: Quality-Diversity Multi-Agent Search for Alpha Factor Discovery in Market Microstructure Data Kimi解读
工具使用(2 篇)
- Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe Kimi解读
- API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces Kimi解读
工程实践启示
- 工程架构方向 6 篇,关注系统设计与可扩展性。
- 工具使用方向 2 篇,function calling 与工具链持续演进。
- 记忆系统方向 9 篇,RAG 与长期记忆方案不断优化。
- 多智能体方向 3 篇,协作模式从简单分工走向复杂协调。
- 安全方向 4 篇,Agent 安全从外部围栏走向内化机制。
下周关注
持续热点:规划推理(本周 14 篇,上周 23 篇)、评估基准(本周 13 篇,上周 11 篇)、记忆系统(本周 9 篇,上周 7 篇)、其他(本周 9 篇,上周 15 篇)、工程架构(本周 6 篇,上周 8 篇)、安全对齐(本周 4 篇,上周 8 篇)、多智能体(本周 3 篇,上周 7 篇)
附录:本周论文完整列表
去重后共 45 篇。
2026-09-07(12 篇)
- Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe Kimi解读 — tool, evaluation
- Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models Kimi解读 — memory, planning, evaluation
- CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents Kimi解读 — other
- Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability Kimi解读 — memory
- Large Language Models for HVAC Operations in Building Energy Systems: A Critical Review of Methods, Applications, and Deployment Readiness Kimi解读 — engineering
- GUT: Quantifying and Optimizing the Reasoning Uncertainty of LLMs via Graph Complexity Kimi解读 — planning
- Testing Interchangeability in LLM Agent Teams Kimi解读 — evaluation
- Trace2Tower: Transition-Aware EigenTrace Induction of Multi-Level Skills for LLM Agents Kimi解读 — other
- Commonsense Reasoning in Computer Vision: Foundations, Recent Advancements, and Future Directions Kimi解读 — planning
- Do LLMs Exhibit Coherent Knowledge Structures in Mathematical Reasoning? A Perspective from Knowledge Space Theory Kimi解读 — planning
- Substrate-Aware AI Agents: Execution Context as a First-Class Input Kimi解读 — other
- CABAL: Multi-Agent Simulacra for Tracing the Effects of Collusive Bidding in Peer Review Kimi解读 — multi_agent
2026-09-09(16 篇)
- Procedural Graphs: Self-Evolving Execution Structures for LLM Agents Kimi解读 — other
- ExecCritic: Learn to Test, Test to Improve for Coding Agents Kimi解读 — evaluation
- MeClear: Cooperative Game-Theoretic Attribution and Risk-Aware Memory Clearance for Long-Horizon LLM Agents Kimi解读 — memory, multi_agent, safety
- SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research? Kimi解读 — other
- Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training Kimi解读 — memory, safety
- Answer-Distribution Trajectories: A Stochastic-Dynamics View of LLM Reasoning Kimi解读 — planning
- Deposon: An Auditable, Conservation-Guaranteed, Game-Theoretically Tested Scattering Layer over LLM Reasoning Paths Kimi解读 — planning, evaluation
- PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving Kimi解读 — evaluation
- SkillAdam: Stable and Efficient Skill Evolution for Agents Kimi解读 — engineering
- API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces Kimi解读 — tool, evaluation
- Closing the Consistency Gap: Self-Evolving Agents That Learn to Stay on Course Kimi解读 — other
- GoAnt: Quality-Diversity Multi-Agent Search for Alpha Factor Discovery in Market Microstructure Data Kimi解读 — multi_agent
- CLAMP: Constrained Decoding for Vision-Language Embodied Planning Kimi解读 — planning
- Graph-Based Personalized Memory for LLM Agents: Representation, Evolution, Retrieval, and Evaluation Kimi解读 — memory, evaluation
- A Three-Tier Persona Vector for Controllable User Simulation in Agentic Evaluation Kimi解读 — evaluation
- AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems Kimi解读 — engineering
2026-09-10(17 篇)
- JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task Composition Kimi解读 — other
- ConvMem: Convolutional Memory for Long-Context Reasoning Kimi解读 — memory, planning
- Fortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMs Kimi解读 — memory
- From Symbolic Perception to Logical Deduction: A Framework for Guiding Language Models in Geometric Reasoning Kimi解读 — planning, evaluation, engineering
- TRACE: Training Reasoning Agents for Causal Exploration with Synthesized Rewards Kimi解读 — planning
- What Should an Agent Forget? Separating What Is Stored from What Is Used Kimi解读 — other
- Why Sample What You Can Enumerate? Exact Policy Optimization for Genomic Tool Selection Kimi解读 — engineering
- Beyond Surface Imitation: Contrastive Modeling for Reasoning Path Alignment in Multimodal In-Context Learning Kimi解读 — planning, safety
- Kernel-Managed Shared Memory for System-Wide Personalization Kimi解读 — memory
- Agent-Based ML-LLM Fusion with Self-Optimizing Prompts for Plateau Weather Alerts Kimi解读 — other
- Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States Kimi解读 — evaluation
- OntologyAligner: Ontology-Aligned Retrieval and Hierarchy-Guided Large Language Model Reranking for Biomedical Ontology Normalization Kimi解读 — memory
- Belief-State Engine: Augmenting LLMs for Principled Planning Under Partial Observability Kimi解读 — planning
- Structural Process Supervision for Latent Chain-of-Thought Reasoning Kimi解读 — planning, safety
- AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI Agents Kimi解读 — evaluation, engineering
- The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM Agents Kimi解读 — evaluation
- Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward Kimi解读 — planning
本报告由 OpenClaw 自动生成,基于 agent-papers-research 每日数据聚合。论文标题和摘要由 GLM-5 翻译生成。