Tharun V. PuthanveettilPortfolio ↗

Multimodal UAV autonomy · 2024

Track what
a person means.

Can a person identify a target by clicking, drawing a box, supplying an image, or writing text—and have one aerial system turn that intent into closed-loop motion?

Track Anything Rapter (TAR) combines pretrained vision models, segmentation, persistent tracking, re-identification, ROS 2 integration, visual servoing, and PX4 flight control. The project moves beyond showing a bounding box: it measures how the complete UAV trajectory agrees with a Vicon-derived target trajectory.

Authors
Tharun V. PuthanveettilFnu Obaid ur Rahman
PublicationarXiv:2405.11655 · May 2024
ValidationGazebo + physical VOXL2 M500 + Vicon

Abstract

A query is useful only if it survives the entire chain from human specification to perception, target persistence, control, and physical motion. TAR makes each transition explicit—and evaluates the final behavior as a trajectory.

PaperRead the preprint ↗RepositorySource, setup and demos ↗ReportMethods and result table ↗
LicenseMIT open-source implementation
Physical rolloutQuery → target → tracking → visual servoing → PX4

Method

Separate specification from persistence.

The query tells the system what to acquire. The tracker then carries that target through time. Re-detection is invoked when temporal tracking is no longer trustworthy, so the perception stack has an explicit recovery path.

01Specifytext · image · click · box
02DetectDINO · CLIP · descriptor match
03SegmentSAM · target mask
04Persist(Seg)AOT · re-detection
05Actvisual error · velocity command
Text query

Name the target

Language is embedded and compared against candidate visual regions, enabling open-vocabulary specification without training one detector per object.

Image query

Show the target

A reference crop supplies a visual descriptor that can be matched against the live scene and reused for later re-identification.

Click / box

Point to the target

Direct image-space interaction seeds segmentation and tracking when the user can see the desired object but cannot describe it reliably.

Three-level recovery ladder

LevelRecovery mechanismTrade-off
1 · Tracker recoveryThe temporal tracker attempts to reacquire the recent target.Fastest, but vulnerable after long disappearance or identity drift.
2 · Human re-seedingA new click or box explicitly identifies the target.Reliable supervision, but interrupts autonomy.
3 · Descriptor re-identificationStored visual descriptors search for the target after re-entry.Automatic recovery, bounded by appearance and range.

Implementation

From a workstation model to a flying robot.

Compute, streaming, perception, and flight control are distributed across the ground workstation and the VOXL2-enabled vehicle. ROS 2 messages define the handoff from image-space target error to vehicle commands.

Physical stack

  • VOXL2 M500 custom UAV
  • PX4 Autopilot and offboard control
  • RTSP camera stream from vehicle to ground station
  • ROS 2 Humble on workstation; ROS 2 Foxy on the drone-side stack
  • Vicon motion capture for ground-truth trajectory evaluation

Software stack

  • Python 3.10 · CUDA 12.2
  • DINO and CLIP for query-conditioned localization
  • SAM for target segmentation
  • AOT-family tracking for temporal persistence
  • ROS 2 vision and controller nodes
  • PX4 velocity/yaw command interface
  1. 01

    Validate the video path

    Confirm the live RTSP stream independently before involving perception or control.

  2. 02

    Run the vision node

    Resize and decode the stream, acquire the query, estimate target location, track it, and publish image-space error.

  3. 03

    Run the controller node

    Convert the target offset into proportional vehicle commands while maintaining the selected object near the desired image feature.

  4. 04

    Mirror the path in simulation

    Gazebo, PX4 SITL, RTSP proxying, QGroundControl, and scripted target motion reproduce the same perception–control boundary before physical flight.

Evaluation protocol

Measure behavior as a trajectory.

Per-frame overlap would evaluate only the vision output. TAR instead compares the time series of target and UAV positions so perception delay, controller response, target loss, recovery, and vehicle dynamics all remain visible in one closed-loop measure.

01 · Capture

Record synchronized paths

Vicon provides target and vehicle positions while the TAR stack flies the physical platform.

02 · Align

Use Dynamic Time Warping

DTW aligns trajectories that may progress at different rates rather than forcing frame-to-frame correspondence.

03 · Compare

Report mean aligned distance

Lower mean distance indicates closer path agreement for the complete query-to-control loop.

Physical tracking trace

The demonstration video combines the arena view with live diagnostic feeds. These frames show the target progressing across the workspace while the perception and tracking views remain observable in the inset.

Metric boundary: DTW is deliberately system-level. A higher value cannot by itself identify whether error came from target localization, tracker drift, streaming latency, controller gain, vehicle dynamics, or Vicon alignment.

Results

Occlusion is visible in the flight path.

The report compares target types, selection methods, DINO/VTM features, a VOXL AprilTag condition, and obstructed sequences. Lower mean DTW is better.

TargetConditionMean DTWReadout
TurtleDINO · unobstructed0.62 mLowest reported trajectory disagreement.
AprilTagDINO · unobstructed0.67 mFeature-based target specification remains close to reference.
AprilTagVTM · unobstructed0.80 mIntermediate result under the reported alternative.
TurtleBounding box0.98 mHigher disagreement than turtle/DINO.
AprilTagBounding box0.98 mSame reported mean as turtle bounding-box case.
VOXL AprilTagPlatform baseline0.80 mReported platform comparison.
AprilTagDINO · obstructed1.56 mObstruction increases trajectory disagreement.
TurtleDINO · obstructed1.65 mLargest reported mean DTW.
Bar chart comparing mean Dynamic Time Warping distance across TAR tracking cases
Reconstructed from the report table · mean DTW in meters · lower is better
Honest read: the recovery ladder prevents every target loss from becoming terminal, but the obstructed cases roughly double path disagreement. TAR demonstrates recovery-capable multimodal tracking—not obstruction-proof autonomy.

Contribution and resources

A collaborative system with attributable components.

The project is open source and includes physical and simulation run instructions, environment definitions, query assets, visualization media, and the full project report.

My work connected the perceptual result to closed-loop aerial behavior and made that behavior measurable.

I implemented the proportional high-level controller, ported and integrated the system across ROS 2 and MAVROS/PX4 interfaces, and designed the DTW-based trajectory evaluation against Vicon ground truth. The broader detection, tracking, perception, and vehicle integration were collaborative with Fnu Obaid ur Rahman.

Scope boundary. The reported tests cover a limited target set and controlled indoor conditions. Generalization to outdoor illumination, weather, long-range sensing, fast targets, crowds, and unseen object categories requires additional evaluation.

CitationTharun V. Puthanveettil and Fnu Obaid ur Rahman. “Track Anything Rapter (TAR).” arXiv:2405.11655, 2024.
Next project

How can explicit coordination states turn individual sensing into safe multi-robot behavior?

Open Auto-Platoon →