Daily Paper

ReferTrack: Referring Then Tracking for Embodied Visual Tracking

ReferTrack introduces a referring-then-tracking framework that grounds embodied visual tracking by selecting an explicit bounding box before decoding navigation waypoints from a single forward-facing...

arXiv:2607.20061 Empirical Study

Hanjing Ye, Tianle Zeng, Jiazhao Zhang, Shaoan Wang et al.

embodied-visual-trackingvision-language-actiongrounded-referring-expressiontarget-followingsim-to-real-transfer

1. Introduction: The Challenge of Embodied Visual Tracking

Embodied Visual Tracking (EVT) requires a mobile agent to continuously pursue a specific target, defined by natural language instructions, using only egocentric onboard vision. Ye et al. characterize this task as a dual challenge of target identification—distinguishing the correct subject from potential distractors—and trajectory planning, which involves generating collision-free waypoints to maintain a consistent following distance.

The paper identifies a critical failure mode in current Vision-Language-Action (VLA) models: their reasoning often relies on abstract spatial latents during the chain-of-thought (CoT) process. These latents are difficult to supervise and frequently show weak alignment with explicit detections in the image space. For AI safety researchers and red-teamers, this opacity is problematic; when a model fails, it is often impossible to determine if the error was a perceptual failure in identification or a motor failure in navigation. By grounding the reasoning process in discrete image-space selections, the authors aim to improve the transparency and reliability of robotic following in dynamic, crowded environments.

2. The Paradigm Shift: Referring Then Tracking

To address the limitations of abstract reasoning, the researchers introduce ReferTrack, a framework that decouples tracking into an explicit “referring” phase followed by waypoint decoding. This paradigm transforms target identification into a constrained multiple-choice problem, which leverages the inherent grounding strengths of Vision-Language Models (VLMs).

The architecture is defined by three core components:

  • Candidate Catalog: At each timestep, an off-the-shelf detector (YOLO11 combined with the ByteTrack association tracker) identifies pedestrians in the forward-facing camera view. These are organized into an indexed catalog of bounding boxes (bboxes), including a virtual ⟨NO EXIST⟩ index for cases where the target is not visible.
  • Refer-CoT: The model predicts a single Refer-CoT token that selects the correct index from the candidate catalog. This provides a compact, supervised decision point in the image space, serving as an observable checkpoint for failure analysis.
  • TVBI tokens: To maintain target persistence, the model utilizes Temporal-Viewpoint-Bbox Indicator (TVBI) tokens. These tokens inject the geometric features of a sliding-window queue of previously selected bboxes into the visual history. When a target is unobserved in a historical frame, the researchers employ an “absence sentinel” coordinate of bt=[0,0,0,0]b_t = [0, 0, 0, 0] to distinguish missing-target states from active motion.

3. Architecture and Implementation

ReferTrack utilizes the Qwen3-4B Large Language Model (LLM) as its backbone. The visual perception branch employs a dual-encoder setup consisting of SigLIP and DINOv2. The features from these encoders are concatenated and processed via grid pooling to generate fine tokens (Vfine∈R64×CV_{fine} \in \mathbb{R}^{64 \times C}) for the current observation and coarse tokens (Vcoarse∈R4×CV_{coarse} \in \mathbb{R}^{4 \times C}) to retain broader historical context.

The visual history is managed through a sliding window of the latest HH frames. During training, the current frame’s fine tokens are deprived of explicit bbox injections, forcing the model to rely on historical TVBI cues and raw visual features for grounding. The policy is optimized using a weighted sum of three loss components:

  1. Trajectory Loss (LtrajL_{traj}): Minimizes the Mean Squared Error between predicted waypoints w^i=(x,y,θ)\hat{w}_i = (x, y, \theta) and expert waypoints.
  2. Referring Loss (LreferL_{refer}): A cross-entropy loss supervising the Refer-CoT classification over the discrete candidate vocabulary.
  3. Text Loss (LtextL_{text}): Standard cross-entropy for language-based grounding tasks.

4. Benchmarking Performance: EVT-Bench Results

Ye et al. evaluated ReferTrack on the EVT-Bench suite using a single forward-view camera protocol. The authors report that the model achieves state-of-the-art performance, matching or surpassing multi-camera baselines on identification-heavy splits.

TaskSuccess Rate (SR) ↑Tracking Rate (TR) ↑Collision Rate (CR) ↓
Single-Target Tracking (STT)89.4%92.5%1.6%
Distracted Tracking (DT)73.3%81.8%7.6%
Ambiguity Tracking (AT)74.1%85.7%7.7%

The paper’s ablation study (Table 2) provides empirical evidence regarding the “identification bottleneck.” The authors found that when using an “Oracle TVBI” (ground-truth bboxes), the Success Rate on Distracted Tracking reached 81.5%. The gap between this oracle and ReferTrack’s 73.3% suggests that identification remains the primary source of failure. Critically, when the Refer-CoT interface is removed entirely, the SR drops by 17.6% (falling to 55.7%), proving that the image-space referring interface is the fundamental source of the model’s robustness against distractors.

5. Bridging the Gap: Refer-QA and Sim-to-Real Transfer

To ensure robust real-world performance, the researchers co-trained the model on 1.3M navigation samples and 1.3M Refer-QA samples derived from the SYNTH-PEDES dataset. This strategy strengthens the model’s ability to ground complex appearance descriptions (e.g., “man in a white shirt and black pants”) into specific bbox indices.

Real-world deployments on the Unitree Go2 quadruped and Unitree G1 humanoid robots validated the model’s stability. The researchers observed:

  • Target Persistence: The model maintained tracking even when the camera field of view captured only partial subjects (e.g., the target’s lower body).
  • Partial Occlusion: The system remained stable when pedestrians detoured around obstacles or during multi-person interference.
  • Operational Efficiency: The perception-and-control loop operated at approximately 10.6 Hz, with the detection phase requiring only 12 ms per step.

6. Conclusion: Implications for AI Safety and System Reliability

The ReferTrack framework demonstrates that explicit image-space grounding significantly enhances the reliability of embodied agents. For the AI safety community, this research offers three critical takeaways:

  1. Observable Checkpoints: By shifting from abstract latents to explicit bbox selection, the model provides a “diagnostic checkpoint.” This allows for the clear attribution of failures to either perceptual misidentification (Refer-CoT error) or motor execution (trajectory error), preventing “covert” failures where a model appears to track while actually drifting in the latent space.
  2. System Efficiency: The authors achieve state-of-the-art results without the resource-heavy requirements of reinforcement learning refinement or expansive multi-camera arrays, proving that architectural grounding is a viable alternative to raw scaling.
  3. Grounding Stability: The integration of TVBI tokens and absence sentinels ensures that target motion is anchored in geometric reality, providing superior target persistence in crowded or occluded scenes.

The code for ReferTrack is available at: https://github.com/MedlarTea/referTrack.

Read the full paper on arXiv · PDF