TransGraspNet: Physically and Geometrically Consistent Manipulation of Transparent Labware
Proposes TransGraspNet, a framework enforcing boundary, surface, and physics consistency across perception, depth reconstruction, and grasp planning for robotic manipulation of transparent...
1. Introduction: The Invisible Risk in Laboratory Automation
Deploying autonomous “robot scientists” to execute complex experimental procedures in chemistry and biology laboratories requires robotic manipulators to handle transparent laboratory glassware, such as beakers, flasks, and test tubes. However, transparent objects introduce severe optical ambiguity: they lack internal visual texture, and their visual appearance is dominated by light refractions, specular highlights, and background distortions in RGB images. Furthermore, commodity RGB-D depth sensors frequently fail to capture glass surfaces, resulting in missing depth returns and severe boundary corruption.
When these transparent vessels contain liquid, robotic handling becomes a safety-critical manipulation task. Minor perceptual or geometric errors easily lead to tilted grasps, causing unstable transport, vessel slippage, or hazardous liquid spillage. To address these vulnerabilities, Hu et al. propose TransGraspNet, a framework designed to eliminate cascading cross-stage errors by enforcing geometry–physics consistency across perception, depth reconstruction, and grasp planning.
2. The Cascading Failure Chain in Standard Modular Pipelines
Standard autonomous manipulation systems typically rely on modular, cascaded pipelines that separate perception, depth reconstruction, and grasp planning into independent optimization tasks. Hu et al. demonstrate that in transparent object manipulation, optimizing each module in isolation leads to severe cross-stage failure propagation. Errors accumulate across module interfaces: an initial boundary error in 2D perception degrades 3D structural depth reconstruction, which distorts surface normal calculations and ultimately forces the grasp planner to select physically unstable grasps.
The authors document this cascading failure chain and illustrate how TransGraspNet intervenes at each stage:
Cross-Stage Error Propagation in Modular Manipulation Pipelines
| Modular Pipeline Stage & Failure Mode | Downstream Cascading Impact & TransGraspNet Fix |
|---|---|
| Perception Stage: Imperfect or blurred object boundaries caused by background texture leakage and reflection noise. | Cascading Impact: Induces depth bleeding across object edges during 3D reconstruction. TransGraspNet Fix: Employs an Edge-Guided Boundary head with Morphological Erosion supervision to deliver sharp contour priors. |
| Depth Reconstruction Stage: Incomplete depth returns and distorted depth surfaces across transparent interiors. | Cascading Impact: Corrupts surface normal estimations, ruining contact geometry and force-closure analysis. TransGraspNet Fix: Integrates an Edge-Guided Attention Gate (EGAG) and Masked Geometric Retention () loss to preserve surface curvature and surface normal fidelity. |
| Grasp Planning Stage: Task-agnostic grasp scoring focused strictly on local visual confidence, friction, or contact depth. | Cascading Impact: Generates tilted or off-center grasps that fail or cause liquid spillage under dynamic motion. TransGraspNet Fix: Applies geometry–physics aware refinement incorporating principal axis/centroid alignment and wrench-space stability constraints. |
3. The Three Consistency Principles of TransGraspNet
To break this error propagation chain, TransGraspNet enforces consistency across perception, reconstruction, and execution through three coupled mechanisms.
3.1 Edge-Guided Boundary Consistency (Perception)
To overcome visual feature sparsity and reflection noise, Hu et al. introduce boundary-consistent reasoning using explicit contour priors. The perception backbone consists of a ResNet-101 with a Feature Pyramid Network (FPN), augmented by an Enhanced Convolutional Block Attention Module (E-CBAM). E-CBAM suppresses background clutter and reflection noise by refining feature representations across spatial and channel dimensions via .
The network incorporates a dual-stream head featuring a dedicated Edge Branch alongside the standard mask branch. The Edge Branch predicts contour probability maps , which are fused with mask features (). Edge supervision is generated via morphological erosion:
The total loss incorporates a Dice-style loss for edge prediction. By forwarding predicted object masks and edge probability maps as explicit geometric priors to downstream modules, the system isolates structural object boundaries. On the RobotSci-Glass dataset, this edge-guided mechanism achieves an instance mask average precision () of 78.5% and a Boundary F-score of 65.3% (a 17.1% increase over the baseline Mask R-CNN).
3.2 Surface Consistency via Geometry-Aware Depth Completion (Reconstruction)
Reconstructing the 3D surface of transparent labware requires recovering interior depth without causing cross-boundary depth bleeding into the background. TransGraspNet adopts a TDCNet depth completion backbone integrated with an Edge-Guided Attention Gate (EGAG). EGAG uses the boundary priors forwarded from the perception stage to compute a gating scalar . This scalar adaptively regulates multi-modal feature fusion between RGB images and sparse depth maps, suppressing depth propagation across object edges while preserving smooth interior surfaces.
To prevent geometric collapse and maintain true physical shape (such as cylindrical curvature), the authors introduce a Masked Geometric Retention (MGR) loss:
In this formulation, enforces strong spatial gradient constraints over the object mask to retain true physical shape and curvature, while dampens background gradients () to prevent edge blur and background depth bleeding. Evaluated on the RobotSci-Glass Depth Golden Subset, this architecture reduces root mean square error (RMSE) to 18.1 mm (compared to 25.4 mm in baseline TDCNet) and cuts the mean surface normal error () nearly in half to 8.4° (down from 15.2°), providing accurate surface geometry for force-closure evaluation.
3.3 Physics-Aware Grasp Refinement (Manipulation)
Standard 6-DoF grasp generators (such as GraspNet-1Billion) output candidate grasps along with a raw candidate score , which measures baseline visual and geometric confidence. However, because optimizes only local friction or contact depth, Hu et al. show that relying solely on this raw score frequently selects tilted or off-center grasps. TransGraspNet re-scores candidate grasps by enforcing global physical and geometric constraints.
Using Principal Component Analysis (PCA) on the reconstructed 3D point cloud, the system extracts the object centroid () and principal vertical axis (). For a grasp candidate with center and approach vector , the refinement module evaluates five score factors:
- Radial Alignment (): Measures gripper proximity to the principal axis via .
- Angular Matching (): Penalizes deviations between the gripper approach vector and the ideal orthogonal vector via .
- Centroid Alignment (): Penalizes lateral and axial distance offsets relative to via .
- Antipodal Friction Closure (): Checks whether contact normals satisfy antipodal friction cone constraints via .
- Wrench-Space Robustness (): Evaluates the metric (), which measures the radius of the largest ball centered at the origin contained within the grasp wrench space. This metric explicitly quantifies the grasp’s capacity to resist arbitrary external disturbance forces and torques during dynamic transport.
The final score integrates the raw baseline visual confidence with the five geometry–physics refinement factors:
The weighting parameters were calibrated offline by Hu et al. via linear regression on 200 labeled candidate grasps to maximize correlation with actual physical success without requiring backpropagation training. Favoring upright, centroid-aligned poses, this refinement reduces grasp angular error from 22.5° to 3.8° and cuts the center offset along the vessel height from 35.2 mm to 8.5 mm compared to baseline visual confidence scoring.
4. Experimental Validation: Dataset and Real-World Platform Testing
Hu et al. validated TransGraspNet through benchmark evaluations, domain fine-tuning on a novel laboratory dataset, and physical robotic experiments.
- Dataset Innovation: To address the lack of ground-truth depth data for laboratory glassware, the authors constructed the RobotSci-Glass dataset covering 20 labware categories (including 15 transparent vessels). To capture true depth ground truth, the authors developed an “Opaque Coating” acquisition method: transparent vessels were blackened with spray coating, allowing an Intel RealSense camera to capture noise-free depth ground truth. The dataset includes a Perception Subset with 5,000+ annotated RGB-D images and a Depth Golden Subset of 200 representative scenes reserved for depth fine-tuning and quantitative evaluation.
- Physical Hardware Setup: Real-world validation was conducted on a physical robotic platform consisting of an AUBO i5 6-DoF robotic arm, a CTAG2F90C adaptive parallel gripper, and an eye-in-hand Intel RealSense D435i RGB-D camera.
- Four-Stage Evaluation Protocol: To ensure rigorous closed-loop safety evaluation, Hu et al. defined a manipulation trial as successful only if it satisfied a strict 4-stage execution sequence:
- Perception Success: Accurate instance segmentation, boundary detection, and 3D pose estimation.
- Grasp Success: Successful finger closure and lifting of the vessel to a height of 30 cm, followed by a 3-second static hold without slippage.
- Transport Success: Smooth horizontal movement across a 20 cm trajectory to the target placement zone.
- Placement Success: Upright, stable placement on the target surface without tipping, shaking, or liquid spillage.
- Grasp Success Rates: Across 100 full-loop physical manipulation trials, TransGraspNet achieved a 96.0% success rate in simple single-object scenes (48/50) and an 86.0% success rate in hard cluttered scenes (43/50), yielding an overall success rate of 91.0%.
- Dynamic Liquid Transport Test: To evaluate safety under dynamic inertial disturbances, the authors performed high-speed liquid transport trials using 50 mL Erlenmeyer flasks and 100 mL glass bottles filled to 50% capacity with liquid.
Dynamic Liquid Transport Performance: During horizontal transport across a 30 cm trajectory at a velocity of 0.5 m/s and an acceleration of 1.0 m/s², TransGraspNet achieved zero liquid spillage. The authors report that wrench-constrained scoring placed contact points near object centroids, minimizing inertial tilt moments and suppressing micro-slippage during rapid acceleration.
5. Key Takeaways for Embodied AI Safety
- Cross-Stage Co-Design is Essential: Independent optimization of modular vision and grasp models creates compounding failure chains across module interfaces. In safety-critical physical automation, downstream physical execution constraints must actively guide upstream perception and depth reconstruction.
- Boundary Priors Stop Depth Bleeding: Explicit 2D contour supervision is critical for resolving optical ambiguity in transparent or reflective settings. Guiding 3D depth completion with sharp boundary priors prevents depth bleeding across object edges and preserves surface normal fidelity.
- Task-Agnostic Grasps Fail Under Motion: Traditional grasp generators that score candidates solely on local visual confidence or contact depth yield unstable, tilted poses. Integrating wrench-space stability () and centroid alignment () into grasp evaluation is mandatory for safe, spill-free liquid transport.
Read the full paper on arXiv · PDF