Tharun V. PuthanveettilPortfolio ↗

Connected autonomy · 2023–2024

Coordinate by
explicit state.

Can a group of small autonomous vehicles follow a leader, react coherently to a local obstacle, and resume without requiring a mechanically coupled platoon?

Auto-Platoon is a physical two-robot leader–follower system organized around “software latching”: explicit engage, follow, pause, disengage, and resume transitions shared through a networked coordination layer. The project integrates learned detection, image embeddings, monocular depth, feature tracking, Kalman estimation, obstacle sensing, and low-level motion.

Authors
Tharun V. PuthanveettilAbhijay SinghYashveer JainVinay BukkaSameer Arjun S.
PublicationarXiv:2405.11659 · May 2024
Physical scopeTwo Raspberry Pi mobile robots + host server

Research premise

A safe platoon is not merely a follower controller. It is a distributed agreement about when the group is engaged, when one agent’s observation should stop everyone, and when the formation may resume.

PaperRead the preprint ↗RepositoryImplementation and setup ↗DemonstrationPhysical result videos ↗
MaturityPhysical proof of concept · two robots
Two-robot physical validationFollow · obstacle-triggered pause · coordinated resume

Core idea

A latch sits between confidence and motion.

The perception stack proposes a relationship between leader and follower. The latch turns that continuous, uncertain estimate into an explicit coordination state. A robot may be visible without the platoon being engaged, and a cleared obstacle does not imply immediate motion until the re-engagement conditions are satisfied.

State 01

Disengaged

No valid leader relationship. The follower holds rather than extrapolating motion from stale perception.

State 02

Latched

Identity, distance, and communication conditions support a leader–follower relationship; following commands are permitted.

State 03

Paused

An obstacle or peer-stop message suppresses motion across the platoon while preserving the coordination context.

State 04

Reacquiring

Perception and estimation re-establish the leader and safe spacing before motion resumes.

State 05

Resumed

The latch returns to an engaged state only after the relevant obstacle and tracking conditions clear.

Design role

Group behavior

The state machine converts a local observation into a coordinated response instead of relying on independent reactive controllers.

Architecture

Perceive, estimate, decide, communicate, control.

The stack uses multiple perception routes because no single signal is sufficient across all situations. Appearance maintains identity, features support tracking, monocular depth estimates spacing, and obstacle sensing can override the following objective.

01PerceiveYOLOv7 · feature tracking
02RepresentMediaPipe image embedding
03EstimateMiDaS depth · Kalman filter
04Coordinatesocket messages · latch state
05Controldistance hold · stop · resume
Published Auto-Platoon process flow connecting perception, tracking, depth estimation, dynamic planning, cooperative sensing and close-range control
End-to-end process architectureCamera input passes through perception and tracking before planning, cooperative coordination, and physical control.Preprint · Fig. 2
ModulePurposeFailure it addresses
Custom YOLOv7Detect the leader vehicle and relevant scene objects.Initialization and reacquisition after tracking loss.
Feature trackingMaintain the target between heavier detections.Frame-to-frame identity continuity.
Image embedderCompare appearance for leader identity.Latching onto the wrong similar object.
MiDaS depthEstimate relative separation from monocular imagery.Following too closely or losing useful spacing.
Kalman filterRegularize noisy position and motion estimates.Control oscillation from frame-level noise.
Obstacle and peer messagesPropagate stop and resume state.One follower continuing while another must stop.

Architecture breakdown

The published diagrams expose the decisions between modules.

I designed the overall process architecture and the published subsystem architectures: the dynamic planner, cooperative sensing and communication topology, and close-range controller. The implementation and paper were collaborative; the contribution here is stated at the architecture level rather than implied by a generic system description.

Trace one module
01 · Perception and tracking

Establish identity before motion.

A custom YOLOv7 model initializes the leader. MobileNetV2 embeddings and cosine similarity associate it across frames, while a Kalman filter estimates centroid, scale, and motion. MiDaS supplies the relative depth used downstream.

Dataset
4,200 train · 400 validation · 200 test
Tracker
appearance embedding + state estimation
Range
relative depth + reference calibration
YOLOv7 training curves for the Auto-Platoon target detector
Detector trainingFig. 13
Leader robot with a persistent tracking identifier in the follower camera
Persistent leader trackFig. 15
02 · Dynamic planning

Choose follow, stop, or search.

Tracked target state and depth enter a planner that distinguishes a leader from an obstacle. A valid leader produces a follow plan; an unsafe depth produces stop-and-proceed behavior; target loss triggers leader search.

Dynamic planner flow chart selecting follow, stop-and-proceed, or leader-search behavior
Dynamic planner architectureFig. 7
03 · Cooperative sensing

Turn a local event into group state.

The main server processes images and returns tracking outputs. Followers report status to the leader’s local server and read the current system state, allowing one agent’s obstacle event to stop the coordinated system.

Network architecture connecting follower robots, leader local server and main perception server
Cooperative network topologyFig. 8
04 · Close-range control

Align first, then regulate distance.

The controller combines target state, IMU orientation, and monocular depth. It corrects angular misalignment, evaluates the leader-distance threshold, then commands trace or stop behavior.

Close-range controller flow chart using object tracking, orientation and depth data
Low-level controller architectureFig. 9

Intermediate evidence from the perception stack

The page now retains the artifacts that connect the architecture to the physical result: dataset construction, temporal tracking, depth estimation, and detector outputs.

Auto-Platoon software-latching state machine from leader acquisition through coordinated stop and resume
Software-latching state machine reconstructed from the published system description. The diagram communicates control logic; it is not a measured timing trace.

Physical implementation

Coordination lives across three computational layers.

A host computer performs the heavier perception and state logic. Raspberry Pi computers on the leader and follower bridge network decisions to local sensors and motors.

Host

Perception and coordination server

Python, OpenCV, learned models, socket communication, and state estimation run on the central machine.

Leader robot

Motion and group reference

The leader supplies the behavior that followers reproduce while also responding to coordinated stop messages.

Follower robot

Tracking and local safety

The follower estimates separation, commands its motors, detects dynamic obstacles, and publishes state changes back to the group.

  1. 01

    Acquire and identify the leader

    Detection, appearance, and temporal features establish the candidate leader before latching.

  2. 02

    Estimate relative motion

    Monocular depth and filtered image measurements approximate the state required by the following controller.

  3. 03

    Maintain a threshold distance

    The follower advances, holds, or slows as the estimated gap changes.

  4. 04

    Propagate a safety event

    A dynamic obstacle triggers a pause message so the leader and other followers do not continue independently.

  5. 05

    Revalidate and resume

    After the obstacle clears and the leader remains tracked, the latch permits coordinated motion again.

Evidence

A physical behavior result, not yet a benchmark.

The available artifact demonstrates the complete state sequence on two robots: leader acquisition, following, obstacle-triggered group stop, and automatic resume. The paper and repository establish integration; they do not report a sufficiently controlled repeated-trial matrix for statistical performance claims.

Physical two-robot demonstration · rendered at native aspect ratio
Auto-Platoon demonstration frame showing the physical leader and follower platforms
Physical platform frame · leader–follower proof of concept

Coordination sequence

Extracting the state transition from the video makes the implementation visible: nominal spacing, obstacle introduction, group pause, obstacle clearance, and resumed following.

DemonstratedLeader tracking and distance-based following

Physical video

DemonstratedObstacle-triggered coordinated stop

Physical video

DemonstratedAutomatic resume after clearance

Physical video

Not measuredLatency, repeatability, false latch and formation error

Needs logs

Evidence needed for a stronger claim: repeated trials with timestamped leader pose, follower pose, latch state, obstacle state, network delay, stop reaction time, resume delay, and false engage/disengage events.

Contribution and limitations

System integration is the contribution.

The project is best read as a connected-autonomy prototype: a complete perception-to-coordination-to-control stack and a concrete mechanism for turning confidence into group state.

Software latching made uncertainty and coordination explicit.

I designed the end-to-end process architecture and each published subsystem architecture: the perception-to-planning flow, dynamic planner, cooperative sensing and communication network, and close-range controller. I also integrated the perception, estimation, communication, and controller behavior into the physical leader–follower demonstration. The implementation and paper were collaborative, and all co-authors are credited above.

Limitations. Validation uses two small robots and a host-centric architecture. Scale to longer platoons, packet loss, network partitions, outdoor conditions, heterogeneous robots, high speed, and formal string stability was not established.

CitationTharun V. Puthanveettil, Abhijay Singh, Yashveer Jain, Vinay Bukka, and Sameer Arjun S. “Auto-Platoon: Freight by Example.” arXiv:2405.11659, 2024.
Next project

How can planning encode safety when the obstacle is living, deformable geometry?

Open precision weeding →