Daily Paper

FORGE-plus: Force-Budgeted Recovery for Contact-Rich Assembly with a Frozen LLM Supervisor

The authors evaluate a two-layer assembly framework in which a frozen text-only LLM specifies force limits and selects recovery maneuvers from textual force signatures while a low-level controller...

arXiv:2607.21227 Empirical Study

Kyupaeck Jeff Rah, Midum Oh

contact-rich-assemblyforce-constrained-manipulationfailure-recoveryllm-supervisionreinforcement-learning

1. Introduction: The Fragility Problem in Robotic Assembly

In robotic assembly, a fundamental tension exists between operational speed and mechanical gentleness. High-speed insertion often requires substantial force to overcome friction and misalignment, yet excessive force frequently results in the destruction of the components being handled. Rah and Oh identify a critical “Press Harder” paradox: standard recovery strategies that effectively seat robust parts, such as steel bolts, tend to destroy fragile counterparts like nylon clips or glass bottles.

Current failure-recovery systems often rely on Large Language Models (LLMs) or Vision-Language Models (VLMs) to re-plan at a task level using visual data. However, the authors argue that the subtle physical differences distinguishing a successful seat from a destructive jam—such as a cross-thread or a burr—are often invisible to cameras. To resolve these failures safely, a system must “feel” the interaction.

The authors introduced FORGE-plus, a two-layer framework designed to solve two persistent questions: who sets the force ceiling for a specific object, and how should a robot recover from failure without exceeding damage limits? By decoupling semantic reasoning from low-level control, the system ensures that physical safety is maintained even during complex recovery attempts.

2. The Two-Layer Architecture: Separating Authority from Logic

The FORGE-plus architecture is strictly divided into a “Slow Layer” for reasoning and a “Fast Layer” for high-frequency execution. This hierarchy ensures that the high-level reasoner provides strategic guidance without the authority to override physical safety constraints.

  1. The Slow Layer (Frozen LLM): Operating at approximately 0.1–1 Hz, a frozen, text-only LLM performs two distinct roles. First, as a Budget-Setter, it assigns a per-object force ceiling (FmaxF_{max}) based on textual identity (material, mass, and geometry) before execution. Second, as a Recovery-Selector, it analyzes a textual force signature upon failure to choose a maneuver from a fixed action menu, specifying bounded parameters for the chosen action.
  2. The Fast Layer (Low-level Controller): Operating at 60–120 Hz, this layer contains the reinforcement learning (RL) skill and an operational-space controller. It incorporates a ForceClamp that serves as the ultimate safety authority, saturating all commanded forces to the FmaxF_{max} provided by the slow layer.

The authors emphasize that the LLM is text-only and does not utilize images; it relies entirely on “compact textual force signatures” to interpret the physical environment.

3. Design Invariants: Ensuring “Non-Circular” Safety

To ensure the evaluation of the system was not “circular”—where success might be guaranteed by providing the agent with the answer in advance—the authors established three design invariants:

  • Hidden Breaking Force (FbreakF_{break}): The actual force threshold at which an object breaks is sampled randomly per episode and remains hidden from the RL policy and the LLM. It is visible only to the evaluator.
  • Immutable Budgets: Once the initial FmaxF_{max} is set based on object identity, the recovery layer is strictly forbidden from increasing it. Recovery reallocates motion rather than relaxing force limits.
  • Fast-Loop Authority: Safety is structural rather than “model-flavored.” The ForceClamp in the fast loop enforces limits regardless of LLM output, backed by a global hard cap of 120 N as a defense against hallucinations.

4. “Feeling” the Failure: The Power of Force Signatures

Rah and Oh contend that force traces are superior to vision for diagnosing assembly failures. Jams caused by wedges or cross-threads may appear identical to a camera but produce distinct physical profiles. The system utilizes a compact text force signature to summarize these events for the LLM, including:

  • Peak axial force and net insertion depth.
  • Lateral bias and axial rising trends.
  • Contact persistence and slip events.

This approach enables a sophisticated safety breakthrough: the detection of a “contactless hover.” This occurs when a policy correctly refuses a geometrically impossible insertion (such as a tilted gear). While invisible to standard force-threshold detectors, the signature makes this refusal legible, allowing the LLM to route the failure to a regrasp—the only maneuver that fixes a tilted grip.

5. Empirical Results: Success at Sub-Millimeter Clearance

The authors evaluated FORGE-plus on gear insertion tasks with a 0.4 mm diametral clearance. A unified checkpoint successfully navigated clean gates for both fragile and robust objects across two different grippers.

FORGE-plus Performance Highlights (Clean Gates)

Gripper / MilestoneObject ClassSuccess RateBreakage RateMean Peak Force
Robotiq 2F-140 (Clean)Fragile ABS256/2560%15.9 N
Robotiq 2F-140 (Clean)Steel Gear256/2560%35.6 N
Franka Panda (Clean)Fragile ABS200/2000%13.8 N
Table-Pick Flow (2F-140)Fragile ABS64/640%5.4 N

A vital finding emerged from the Table-Pick Flow. The authors discovered that removing non-physical assists (such as “teleporting” the gear into the gripper) and using a real friction-based pick actually lowered peak insertion forces to a 5.4 N mean. This occurred because a physical pick centered the part more accurately in the pads than a kinematic pin, proving that realistic physics can enhance force economy.

While Clean Gate success was 100%, the authors noted that under injected in-grip slip, the recovery success rates were 40% for the 2F-140 and 64% for the Franka Panda.

6. The Failure of “Press-Harder” and Oracle Baselines

The authors contrasted FORGE-plus with a “Press-Harder” baseline, which mimics a common recovery strategy of escalating force by 25% upon failure. This baseline demonstrated the “two faces” of Press-Harder failure:

  • Futile: On the Robotiq 2F-140, the strategy was 100% ineffective (zero success, 100% timeouts). Because the force was increased without fixing the underlying misalignment, the tilted bore never took the load, and the extra force was useless.
  • Destructive: On the Franka Panda, the strategy was catastrophic, resulting in a 96% breakage rate for fragile parts.

The study also highlighted a failure of the “Oracle” baseline, which set the budget exactly at the breaking threshold (Fbreak−ϵF_{break} - \epsilon). This baseline still caused breakage in 49.8% of cases. The authors quantified that low-level controller overshoot peaks at 45–54 N (approximately 1.5x the budget). Thus, identity-derived budgets must be conservative enough to account for physical overshoot rather than just matching breaking thresholds.

7. Negative Results: Why PPO is Not Enough

In the spirit of Failure-First research, the authors detailed several unsuccessful designs:

  • PPO Failure: Standard Proximal Policy Optimization (PPO) failed to solve the 0.4 mm clearance task on the 2F-140. This was due to an “exploration-damage trade-off”: the random noise required for the robot to search for the hole was high enough to destroy the fragile part before any learning occurred.
  • The Learned Release Head: The authors identified that “post-open samples are label poison.” Because the environment latches the first release, any data collected after the gripper opens confuses the learner. The authors detailed three specific failed designs:
    1. Frozen-trunk/all-timesteps: Learned “open if already open” and never fired in deployment.
    2. Pre-open-only labels/frozen trunk: Failed because the arm-trained trunk had already discarded the seat-state signal.
    3. Trunk fine-tuning with distillation loss: Resulted in a conflict between the two losses, pinning the false-negative rate at 50%.
  • Success: The authors eventually succeeded using an “Input-Skip Head” that read raw observations directly, bypassing the main network trunk.

8. Conclusion: Lessons for AI Safety and Robotics

The authors conclude that for contact-rich assembly, safety should be structural and enforced by low-level hardware clamps rather than being a “model-flavored” suggestion. While LLMs are excellent at categorizing objects and selecting recovery maneuvers, they must operate within immutable physical boundaries.

Key Takeaways for Practitioners:

  • Identity-Derived Budgets: Force limits must account for physical controller overshoot (approximately 1.5x the commanded budget) rather than merely matching a material’s breaking threshold.
  • Signature Discrimination: Force signatures are vital for resolving failures like “contactless hovers,” which are invisible to standard vision systems but legible through contact data.
  • Recovery Invariants: Structural safety is maintained by strictly forbidding the relaxation of force limits during recovery maneuvers; the system must reallocate motion, not force.

Limitations: The authors emphasize that this is a simulation-only study where breakage is modeled as a scalar threshold. Future work is required to validate these findings on physical hardware and against naturalistic failure distributions.

Read the full paper on arXiv · PDF