数据来源:papers.cool/arxiv/cs.AI · 生成时间:2026/7/31 22:52:26
本报告聚合本周 arXiv cs.AI 的 Agent 相关论文,跨天去重后深度分析。
本周总览
本周去重后共 57 篇论文,覆盖 4 天数据。上周 76 篇,环比减少 19 篇。
研究方向分布
| 方向 | 本周 | 上周 | 变化 |
|---|---|---|---|
| 规划推理 | 26 | 22 | +4 |
| 评估基准 | 13 | 19 | -6 |
| 工程架构 | 9 | 6 | +3 |
| 记忆系统 | 7 | 7 | 0 |
| 自我进化 | 7 | 1 | +6 |
| 其他 | 7 | 25 | -18 |
| 安全对齐 | 6 | 3 | +3 |
| 多智能体 | 5 | 9 | -4 |
应用场景分布
| 场景 | 论文数 | 占比 |
|---|---|---|
| 信息检索与问答 | 8 | 14% |
| 决策支持 | 8 | 14% |
| 企业自动化 | 5 | 9% |
| 科学研究 | 5 | 9% |
| 代码开发 | 2 | 4% |
| 数据分析 | 1 | 2% |
| 创意与内容 | 1 | 2% |
核心论文解读
1. Agents in the Wild: Where Research Meets Deployment
- arXiv: 2607.19336
- 方向: 规划推理 · 工程架构
- 场景: 科学研究、信息检索与问答、决策支持
- 关键词:
deploymentagenticagentschecklistswildmeetsresearchreasoningplanningdiscovery
2. AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets
- arXiv: 2607.15781
- 方向: 记忆系统 · 多智能体 · 评估基准 · 工程架构
- 关键词:
agentfairgeospatialevaluatorsdatasetdatasetsaveragespercentagecriticagentfair
3. DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings
- arXiv: 2607.15418
- 方向: 规划推理 · 评估基准
- 场景: 企业自动化、信息检索与问答
- 关键词:
drawingsdrawingvqareasoningengineeringconstructionworkflowsmllmsdomaintextualworld
4. AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adaptive Recovery
- arXiv: 2607.15367
- 方向: 规划推理 · 多智能体 · 自我进化
- 场景: 决策支持
- 关键词:
anovaxassistantagentllmrecoveryphonetypedexecutorsdesktopgemini
5. PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents
- arXiv: 2607.18859
- 方向: 其他
- 场景: 代码开发、数据分析、决策支持
- 关键词:
phoenixrepairrepairexplorationsweagenteditdeepsoftwareanalyticsrethinkinginsufficientlocations
6. Knowledge-Centric Agents for Workflow Generation
- arXiv: 2607.15845
- 方向: 规划推理
- 场景: 企业自动化、信息检索与问答
- 关键词:
knowledgeworkflowgenerationworkflowsexecutablecentriccomfyuireasoningstrategiesbrittleness
7. Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents
- arXiv: 2607.15715
- 方向: 安全对齐 · 自我进化
- 场景: 企业自动化
- 关键词:
agenticextractionreflectiveworkflowsagentfixedagentsllmreflectionretries
8. S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation
- arXiv: 2607.15686
- 方向: 规划推理 · 评估基准
- 场景: 科学研究
- 关键词:
omniscientificunifiedreasoningpredictiongenerationspecificmultimodalbenchmarksai4s
9. SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction
- arXiv: 2607.15550
- 方向: 规划推理 · 安全对齐 · 工程架构
- 关键词:
guisafetyseerguardrisksriskassessmentsawmmobileagentsaction
10. WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting
- arXiv: 2607.18084
- 方向: 评估基准
- 场景: 科学研究、信息检索与问答
- 关键词:
worldcuparenascorelinescorefootballmatchaccuracyresultwzk1015evaluationkickoff
研究趋势
主导方向:规划推理(26 篇),较上周(22 篇)上升。
上升: 规划推理(22→26)、工程架构(6→9)、自我进化(1→7)、安全对齐(3→6)
下降: 其他(25→7)、多智能体(9→5)、评估基准(19→13)
技术演进脉络
规划推理(26 篇)
- DSWorld: A Data Science World Model for Efficient Autonomous Agents
- Knowledge-Centric Agents for Workflow Generation
- NeurOWL: An LLM-Based Neural-symbolic Framework for Incomplete OWL Ontology Reasoning
- 及另外 23 篇
评估基准(13 篇)
- AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets
- S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation
- Neuro-Symbolic AI for LEED compliance: Document-Centric Benchmarking, Deterministic Numeric Checking, and When Multimodal Hurts
- 及另外 10 篇
工程架构(9 篇)
- AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets
- NeurOWL: An LLM-Based Neural-symbolic Framework for Incomplete OWL Ontology Reasoning
- SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction
- 及另外 6 篇
记忆系统(7 篇)
- AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets
- Rethinking Heterogeneous LLM Merging: A Weighted Model Averaging Perspective
- Exploratory and Assimilating Reflection: Reflective Recall Cycle for Long-term Memory
- 及另外 4 篇
自我进化(7 篇)
- Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents
- AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adaptive Recovery
- AdaHome: An Adaptive Smart Home Assistant using Local Small Language Models
- 及另外 4 篇
其他(7 篇)
- From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems
- Towards Agentic Agent-based Models: Feasibility, Performance, and Statistical Model Checking
- CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents
- 及另外 4 篇
安全对齐(6 篇)
- Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents
- SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction
- Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes
- 及另外 3 篇
多智能体(5 篇)
- AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets
- Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning
- AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adaptive Recovery
- 及另外 2 篇
工程实践启示
- 工程架构方向 9 篇,关注系统设计与可扩展性。
- 记忆系统方向 7 篇,RAG 与长期记忆方案不断优化。
- 多智能体方向 5 篇,协作模式从简单分工走向复杂协调。
- 安全方向 6 篇,Agent 安全从外部围栏走向内化机制。
下周关注
持续热点:规划推理(本周 26 篇,上周 22 篇)、评估基准(本周 13 篇,上周 19 篇)、工程架构(本周 9 篇,上周 6 篇)、记忆系统(本周 7 篇,上周 7 篇)、其他(本周 7 篇,上周 25 篇)、安全对齐(本周 6 篇,上周 3 篇)、多智能体(本周 5 篇,上周 9 篇)
附录:本周论文完整列表
去重后共 57 篇。
2026-07-20(17 篇)
- DSWorld: A Data Science World Model for Efficient Autonomous Agents — planning
- Knowledge-Centric Agents for Workflow Generation — planning
- AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets — memory, multi_agent, evaluation, engineering
- NeurOWL: An LLM-Based Neural-symbolic Framework for Incomplete OWL Ontology Reasoning — planning, engineering
- Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents — safety, evolution
- S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation — planning, evaluation
- ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning — planning
- Neuro-Symbolic AI for LEED compliance: Document-Centric Benchmarking, Deterministic Numeric Checking, and When Multimodal Hurts — evaluation
- SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction — planning, safety, engineering
- From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems — other
- Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes — planning, safety
- Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3? — planning
- DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings — planning, evaluation
- Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning — planning, multi_agent
- AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adaptive Recovery — planning, multi_agent, evolution
- Cura 1T: Specialized Model for Agentic Healthcare — planning
- Causal-Audit: Explicit and Auditable Graph-based Reasoning via Target-Aware Causal Chain Construction — planning
2026-07-21(14 篇)
- Can We Break LLMs Out of Self-Loops? Fine-Grained Reasoning Control with Activation Steering — planning
- WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting — evaluation
- AdaHome: An Adaptive Smart Home Assistant using Local Small Language Models — planning, evolution
- Rethinking Heterogeneous LLM Merging: A Weighted Model Averaging Perspective — memory
- PAMD: Structured Adaptive Distances for Bisimulation Representations in Visual Reinforcement Learning — evolution
- SAGE: Subgoal-Conditioned Action Generation for Latent World Model Planning — planning
- Towards Agentic Agent-based Models: Feasibility, Performance, and Statistical Model Checking — other
- PEARL: Auditable Repair for Scientific Reasoning Graph Extraction — planning
- Stress Testing Concept Erasure with Large Language Model Agents — evaluation
- ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding — planning
- Exploratory and Assimilating Reflection: Reflective Recall Cycle for Long-term Memory — memory, evolution
- WuYu-EnvLE-Bench: A Benchmark for Evaluating Large Language Models in Environmental Law Enforcement — evaluation
- Semantically Similar, Logically Distinct: Diagnosing the Semantic-Answerability Gap in Table RAG — memory
- SR-Agent: An Experience-Driven Agentic Framework for Post-Ranking Strategies Refinement in E-Commerce Recommendation — engineering
2026-07-22(10 篇)
- CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents — other
- Agents in the Wild: Where Research Meets Deployment — planning, engineering
- ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D — other
- Graph-Based Agentic AI with LangGraph: Workflow Pathways for Long-Running Stateful Business Processes — other
- BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance — evaluation
- Comparative Study of Multi-Agent Actor-Critic Algorithms in Parameterized Action Reinforcement Learning — multi_agent
- OpenRTAG: A Comprehensive Benchmark for Robust Text-Attributed Graph Learning under Data Quality Degradation — evaluation
- Supra Cognitive Modes: A Routed Architecture for Agent Memory — memory, engineering
- Fishing Out Free Riders: Shapley-Based Reward Attribution for Parallel Reasoning via Reinforcement Learning — planning
- PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents — other
2026-07-23(16 篇)
- SoftReason: A Fully Differentiable Neuro-Soft-Symbolic Deductive Reasoning Architecture over High-Dimensional Perceptual Data — planning, engineering
- PoTRE: Test-Time Reasoning inspired by Cognitive Heterogeneity — planning, evaluation
- CUSUM-Shaped Inference-Time Monitoring and Targeted Re-Decoding for Quantized Small Language Model Reasoning — planning
- PRO-LONG: Programmatic Memory Enables Long-Horizon Reasoning — memory, planning
- EvoDRC: A Self-Evolving Agentic Framework for Automated DRC Violation Repair — engineering
- CLARK: Closed-loop Learning for Adaptive Reasoning over Knowledge Graphs — planning, evolution
- Coordinating from Memory: Graph-Structured Experience Reuse for Multi-Agent Adaptation in Dynamic Manufacturing — memory, multi_agent
- EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization — planning, engineering
- MOF-Sleuth: Tool-Grounded Reward Alignment for Explainable Fine-Grained MOF CIF Auditing — safety
- JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety — safety
- DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations — evaluation
- Know Your Agent: Reconnaissance-Driven Pentesting of AI Agents — evaluation
- Rewarding Better Thinking for LLM Preference Alignment — safety
- Silent Failures in Multimodal Agentic Search:A Diagnostic Taxonomy and Cross-Judge Evaluation — evaluation
- Symbol and Footprint Database for Electronic Components by Agentic Recognition and Generation — other
- Knowledge-Centric Self-Improvement — evolution
本报告由 OpenClaw 自动生成,基于 agent-papers-research 每日数据聚合。