Chain of News Digest

Chain of News 09/07/2026

09/07/2026
**Top Story** The recent discovery of illicit AI use at a major conference has led to the rejection of hundreds of papers, as reported by Nature. This incident highlights the growing concern of AI-generated content in academic research and the need for stricter guidelines and detection methods. The conference's decision to reject these papers demonstrates the importance of maintaining academic integrity and ensuring that research is conducted ethically. This development has significant implications for AI developers, as it emphasizes the need for transparency and accountability in AI-generated content. Furthermore, it raises questions about the role of AI in academic research and the potential consequences of relying on AI-generated content. As AI continues to advance, it is essential to establish clear guidelines and regulations to prevent such incidents in the future. **AI Models & Research** The paper "Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks" presents a systematic benchmark for evaluating the performance of agentic AI on computational imaging tasks. This research is significant for developers, as it provides a comprehensive understanding of the capabilities and limitations of agentic AI in handling complex imaging tasks. The study's findings have implications for the development of more advanced AI models that can effectively handle inverse problems and physics-based tasks. Another notable paper, "Measuring Intelligence Beyond Human Scale," explores the challenges of measuring intelligence beyond human capability and proposes new approaches for evaluating AI performance. This research is crucial for developers, as it highlights the need for more advanced evaluation methods that can assess AI capabilities beyond human scale. The paper "Learning social norms enhances compatibility in dynamic human-AI coordination" demonstrates the importance of social norms in human-AI coordination and highlights the need for AI systems that can learn and adapt to social norms. **Developer Tools & Frameworks** The introduction of AgentLens, a production-assessed benchmark for interactive code agents, provides developers with a valuable tool for evaluating the performance of code agents. With AgentLens, developers can assess the entire trajectory of a code agent's execution, rather than just its final output. This allows for more comprehensive evaluation and improvement of code agents. The development of Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1 is another significant update, as it enables developers to create more efficient and effective agent harnesses for abstract reasoning and generalization tasks. The release of Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics provides developers with a new framework for combining LLM agents with Computer Algebra Systems (CAS) for mathematical computations. **Industry & Business** A major conference's decision to reject hundreds of papers due to illicit AI use has significant implications for the academic and research communities. This incident highlights the need for stricter guidelines and regulations regarding AI-generated content in academic research. The development of AgentLens, a production-assessed benchmark for interactive code agents, has the potential to impact the way code agents are evaluated and developed in the industry. The introduction of Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1 may also influence the development of more efficient and effective agent harnesses for various applications. **Worth Watching** The paper "The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI" provides valuable insights into the economics of agentic AI development and the impact of orchestration design on token economics. The study's findings have significant implications for developers and businesses involved in agentic AI development. The development of Large Behavior Model, a promptable digital twin of the retail customer, is another interesting item that deserves attention. This model has the potential to revolutionize customer behavior modeling and provide more accurate predictions and recommendations. The research on "From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents" is also worth watching, as it explores the optimization of tool utilization for LLM agents and its potential applications in various domains.

Today's Stories

Today's articles

ArXiv cs.AI

From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents

Tool utilization enables Large Language Model (LLM) agents to interact with the real world and resolve complex tasks. However, existing agent frameworks predominantly rely on static toolsets composed of granular atomic actions (e.g., basic file I/O or single-turn search), which forces agents to reinvent low-level logic for every recurring workflow, leading to increased reasoning overhead and failure rates.

09/07/2026
ArXiv cs.AI

Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1

Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionary search, exhaustive sampling, extended chain-of-thought), or benchmark-specific training in which small models are fine-tuned on ARC data, often with task-specialized architectures. We study a third regime: an open-weight model in non-thinking mode (DeepSeek V3.2) under a strict budget, with no ARC-specific fine-tuning.

09/07/2026
ArXiv cs.AI

Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics

Recent advances in AI for Mathematics have focused largely on autoformalization and theorem proving, leaving the role of Computer Algebra Systems (CAS) in agentic LLM workflows underexplored. We propose a ReAct-style agentic setup that combines LLM reasoning with verifiable feedback from SageMath, together with Context7 for the up-to-date documentation.

09/07/2026
ArXiv cs.AI

Learning social norms enhances compatibility in dynamic human-AI coordination

Humans continuously coordinate with others in dynamic interactions, often through implicit, hard-to-quantify social norms that act as shared tacit expectations among interacting agents. As AI agents, including large language models (LLMs), become embedded in daily life, they increasingly participate in such interactions and reshape social interaction structures. Yet they often fail to coordinate with humans in an effective, considerate, and natural manner.

09/07/2026
ArXiv cs.AI

Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks

Vision-language models (VLMs) and agentic AI have shown strong performance on semantic visual tasks, but it remains unclear whether they can handle the physics and inverse problems that underlie computational imaging. We present ImagingBench, a benchmark of 20 computational imaging tasks spanning five categories: ray and wave optics, image signal processing, inverse reconstruction, computational sensing, and calibration.

09/07/2026
ArXiv cs.AI

AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit -- did the task pass? -- but the people who actually use these agents experience the entire trajectory: how the agent follows instructions, uses its tools, verifies its own work, recovers from mistakes, and talks to them along the way. AgentLens evaluates that whole trajectory.

09/07/2026
ArXiv cs.AI

The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI

Agentic AI development today runs on token maxing: buying capability with tokens -- longer reasoning traces, more turns, wider tool payloads, bigger replayed contexts -- so tokens per task grow faster than task value. Falling per-token prices mask the pattern; total spend rises anyway. We argue the decisive lever against token maxing is the harness: the orchestration layer that assembles context, exposes tools, sequences turns, delegates work, and carries enterprise observability and governance.

09/07/2026
ArXiv cs.AI

Large Behavior Model: A Promptable Digital Twin of the Retail Customer

Customer behavior modeling underpins recommendation, marketing, and decision support, yet existing approaches either optimize predictive accuracy without explaining decisions or simulate users without grounding them in real behavioral data. We present the Large Behavioral Model (LBM) that learns customer decision making directly from large-scale retail transactions through a unified Person-Environment formulation.

09/07/2026
ArXiv cs.AI

Measuring Intelligence Beyond Human Scale

How can we measure intelligence beyond human capability? Human-authored benchmarks saturate, and above human capability, examiners may not know which tasks are both hard and verifiable. We argue that this difficulty is inherent to absolute-scale evaluation and propose a new paradigm based on relative measurement in which models generate public challenges that separate other systems.

09/07/2026
GNews AI EN

Major conference catches illicit AI use — and rejects hundreds of papers - Nature

Major conference catches illicit AI use — and rejects hundreds of papers Nature

25/03/2026