DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
The authors introduce DocOps, a deterministically verifiable benchmark for evaluating autonomous agents on complex document operations, and analyze the operational failure modes of frontier models...
1. Introduction: From Chatbots to Active Workspace Participants
Large language model (LLM) agents are transitioning from “passive conversationalists” to active participants in digital workspaces, where they are increasingly tasked with manipulating complex artifacts such as spreadsheets, presentations, and PDFs. As Jiang et al. observe, while progress in agentic orchestration is rapid, there remains a critical lack of quantitative understanding regarding how these agents interact with structured documents beyond simple read-only extraction.
The researchers identify a fundamental limitation in prior evaluation paradigms: static frameworks treat documents as fixed knowledge repositories, while workflow-oriented evaluations treat them as transient data payloads. Neither approach treats the document as a “first-class computational object”—a stateful artifact where operations are discrete state transitions that must maintain native-format validity. To address this, the authors introduced “DocOps,” a benchmark released in self-contained Harbor bundles (containerized environments) designed to provide a diagnostic environment for measuring document integrity and agent reliability.
2. The DocOps Taxonomy: Mapping the Operational Space
The paper deconstructs document operations using a two-axis taxonomy that separates atomic capabilities from workflow depth. This framework allows researchers to localize whether a system failure originates from localized structural disruption or a collapse in long-horizon planning.
The Operation Axis This axis categorizes document interaction into three families of atomic primitives:
- Content: Evaluates textual and numeric semantics, including extraction, grounded generation, and formula computation.
- Format: Focuses on the presentation state, encompassing style consistency, layout control, and theme transfer.
- Structure: Targets the organization of native document objects, such as hierarchy editing (e.g., Word outline levels), slide reordering, and the manipulation of native tables or worksheets.
The Difficulty Axis The taxonomy employs a four-tier complexity gradient to measure an agent’s state-tracking and planning capacity:
L1 Atomic to L4 Cross-doc State Tracking
- L1 (Atomic): Local execution of a single operation.
- L2 (Composite): Multiple atomic operations within a single artifact, testing local dependency modeling.
- L3 (Workflow): Maintaining global consistency across multi-step edits in a single complex document.
- L4 (Cross-doc): Producing a final artifact through state tracking and information alignment across multiple source documents.
3. Moving Beyond “LLM-as-a-Judge”: The Deterministic Verifier
Traditional benchmarks often rely on “LLM-as-a-judge” mechanisms that frequently overlook subtle structural corruption. The DocOps methodology replaces these with a deterministic verifier that directly inspects output files using native document libraries. To ensure the benchmark accommodates various valid execution paths, the verifier uses targeted predicates rather than whole-file exact matching.
| Predicate Type | Function | Example from Source |
|---|---|---|
| Structural Predicates | Validates native states and internal object hierarchies. | Verifying executable formulas in Excel or outline levels in Word. |
| Linguistic Anchors | Verifies required content without requiring exact character matches. | Using task-specific diagnostic keywords to confirm content accuracy. |
| Preservation Predicates | Ensures out-of-scope document elements remain unmodified. | Detecting the unauthorized deletion of hidden worksheets or protected styles. |
The authors report that this verifier maintains a 95.31% agreement rate with human audits. In a rigorous stress test involving 180 controlled mutations of 36 verified outputs, the system detected 96.67% of violations across content, structure, and preservation requirements.
4. The Reality Check: Empirical Performance Limits
Jiang et al. systematically evaluated frontier models (GPT-5.5, GPT-5.4, Claude Sonnet 4.6) and open-source models (DeepSeek-V4-Pro, Qwen3.6) across diverse agentic harnesses. The data indicates that frontier systems remain far from fully reliable in document manipulation.
The highest recorded success rate was 0.671, achieved by GPT-5.5 using the Codex harness with skills. Despite its status as a frontier configuration, the model failed nearly one-third of the tasks. Performance degrades sharply as task difficulty increases, demonstrating a breakdown in long-horizon state management. For GPT-5.5, average pass rates across difficulty levels were:
- L1 (Atomic): 0.725
- L2 (Composite): 0.600
- L3 (Workflow): 0.492
- L4 (Cross-doc): 0.237
5. A Taxonomy of Failure: Three Modes of Document Corruption
Through trajectory analysis, the researchers identified three pervasive failure modes that lead to “silent corruption”—outputs that appear visually correct but are functionally or structurally compromised. These failures present a significant safety risk in enterprise environments by creating a false sense of task completion.
- Long-Term State Tracking Failure: Agents lose track of the global document state during multi-step sequences.
- Technical Scenario: In PDF manipulation tasks, an agent may successfully execute localized edits on individual pages but produce a final file where the global sequence order is incorrect or components are misplaced.
- Shallow Semantic Verification: Agents accept outputs that are visually plausible but lack underlying functional accuracy. This is particularly dangerous as it evades surface-level human review.
- Technical Scenario: In spreadsheet workflows, an agent may calculate a correct numerical value but replace a required dynamic formula with a static value, breaking the document’s future computational utility.
- Destructive Editing: Agents “downgrade” complex object trees into flat text or break structural metadata.
- Technical Scenario: When tasked with inserting a column in a PowerPoint table, the agent may “visually patch” the layout by overlaying a text box rather than updating the native XML object tree. This renders the document difficult to edit or automate in future steps.
According to Figure 5 in the paper, semantic verification gaps dominate, accounting for 42.54% of overall failures.
6. The Impact of the Harness: Open-Source vs. Frontier Models
The researchers found that the “harness”—the interface through which a model interacts with the document—is a primary determinant of success. The study compared DocTools (constrained API), Terminus-2 (Bash-based), Codex (CLI), and Claude Code (interactive coding).
The effectiveness of “Skill Injection” (providing explicit procedural guidance) was found to be non-uniform. While it acts as a catalyst for mid-tier open-source models—with Qwen3.5-27B seeing a 7.1 percentage point gain specifically under the Claude Code harness—it offers marginal utility for frontier models. The authors suggest that rigid adherence to prescribed “skills” can actually limit the ability of highly capable models to find alternative problem-solving strategies.
Furthermore, the paper highlights the impact of “coupling.” Success rates in Excel “plummet to near zero” in complex tasks because the format is tightly coupled; a single error in a formula reference propagates throughout the entire workbook state. In contrast, low-coupling formats like PDF, which rely on modular page-level state changes, proved more robust.
7. Conclusion: Shifting the Focus for Document-Centric AI
The DocOps benchmark demonstrates that agent failure is frequently caused by a collapse in state maintenance rather than a simple inability to follow instructions. For AI safety researchers, the primary takeaway is the urgent need for state-aware, non-destructive agents capable of maintaining consistency across highly coupled digital environments.
The paper identifies several future directions for the development of robust agents:
- Expanding benchmarks to include workflows requiring live external services.
- Evaluating agent reliability in collaborative editing environments where the document state is dynamic.
- Integrating interactive user clarification loops to resolve ambiguous document states before corruption occurs.
- Prioritizing the preservation of native structural invariants over surface-level visual plausibility to prevent silent data degradation.
Read the full paper on arXiv · PDF