Tharun V. Puthanveettil

Robotics R&D Engineer

My research takes inspiration from how people learn, adapt, and develop skills. I work on robot learning, multimodal models, and the systems needed to study them on physical robots. My long-term aim is to build robots that improve through experience and understand human intent with the contextual sensitivity people use to understand one another—and act or learn from that understanding.

Tharun speaking at an event

Background

How the research
developed.

My earlier work in perception, multimodal systems, and physical robotics provides the technical basis for my current research in human-inspired robot learning.

2017–21

Data, perception, and early physical systems

Work in statistical learning, computer vision, soft robotics, human-facing AI, and ROS established experience in turning observations into decisions and actions.

Rainfall · Water quality · Soft robot · My Style · Try-on · AI Cricket · ROS
2021–23

Perception applied to manipulation and HRI

At IISc and UMD, I worked on depth estimation, precision weeding, and an exploratory imitation-learning system using demonstrations, speech, and gesture.

Depth · Precision weeding · ROS 2 · Speech + gesture + demonstration
2023–24

Multimodal autonomy and evaluation

I developed systems that combined partial evidence across sensors or viewpoints and connected perception to control, coordination, or anomaly detection.

PoseFusion · Fall detection · Auto-Platoon · TAR · Multimodal anomaly detection
2024–present

Robot learning on physical systems

My current work brings these threads together in bimanual manipulation: teleoperation, multimodal data, compliant control, learned policies, digital twins, and real-world evaluation.

Deformables · Bimanual manipulation · Compliance · Learned policies · Sim-to-real
See where these threads meet now

Questions guiding the research

01What representations support understanding of people, tasks, and environments?
02How can a robot adapt through experience without losing prior capability?
03How should capability be evaluated beyond task-level success?

Present · R&D engineer, AI & robotics

Research infrastructure

A shared system for studying teleoperation, data collection, learned policies, human intervention, simulation, deployment, and comparative policy evaluation on bimanual manipulation tasks.

I designed and implemented the workflow from teleoperation and synchronized collection through dataset conversion, policy training, evaluation, testing, and deployment.

The physical cell and digital twin use the same operator, recording, review, conversion, and policy interfaces, with separate hardware and physics backends. This supports controlled comparisons between simulation and real deployment.

Interventions and failed rollouts can be recorded, reviewed, converted, and returned to training. Policy evaluation is currently performed manually. I am also developing a common evaluation methodology that goes beyond aggregate success rate to compare model behavior more fairly.

The research platform

Implemented Six shared layers · physical and simulated backends · staged evaluation

The gated research platform Six shared platform layers crossed by seven promotion gates. One evidence path descends through the system and a failed authorization returns to capture as recovery data. OperatePerceiveTwinCaptureEvaluateOrchestrate multi-arm teleoperationscene statephysical-authoritativesynchronized evidencelearned policiesconfirmation-gated skills RECORDQCREVIEWFREEZECONVERTSHADOWAUTHORIZE RECOVER · RELABEL · RELEARN
Shared interfaces across physical and simulated backendsInspect a stage · policy evaluation is currently manual
Implemented

Modular ROS 2 infrastructure for bimanual manipulation, integrating teleoperation, perception, simulation, data capture, training, and deployment.

Continual learning

Human takeover, correction capture, review, dataset conversion, retraining, manual evaluation, and controlled policy rollout.

Experimental control

Shared interfaces across the real cell and digital twin, with staged testing and recorded evidence for each policy version.

Current research directions

Each line of work is marked by its current stage of development and validation.

  1. 01

    Continual learning through intervention

    Implemented workflow Human takeover, correction capture, review, conversion, retraining, testing, and deployment are connected. Evaluation is manual while automated policy evaluation is being developed.

    Refusal becomes the next data campaign
  2. 02

    Real → sim → real

    Active experiments I use the digital twin to study transfer across sensing, timing, control, contact, and visual variation. Current experiments examine representations and evaluation methods that remain useful when simulated and physical environments do not match closely.

    PHYSICALSIMULATEDONE INTERFACE
    Backend substitution · shared owners after the bar
  3. 03

    Policy explainability and representation

    Experiments completed I use attention maps, Grad-CAM, state-vector ablations, prompt comparisons, and comparisons across self-supervised and pretrained representations to study what influences policy outputs. These are diagnostic tests; they do not by themselves establish a causal explanation.

    ABLATEDPOLICYΔ
    Cut one input · measure the output response
  4. 04

    Egocentric demonstration learning

    In development We are developing a portable egocentric interface for recording human manipulation demonstrations. It reuses the teleoperation, synchronization, data review, and conversion infrastructure I built for the current manipulation system.

    Current stateShared data infrastructure implemented; demonstration-capture hardware and egocentric policy studies in development.
  5. 05

    Video-action models and inverse dynamics

    Offline evaluation completed I adapted a published video-action architecture to teleoperation data and completed offline evaluation. Physical evaluation has not yet been performed. I am continuing to study how predictive and inverse-dynamics models can support action learning and transfer.

    Next validationPhysical rollouts and controlled comparison with existing policy approaches.
  6. 06

    Policy evaluation beyond success rate

    Methodology in development I am evaluating leading open-source vision-language-action models across tasks and developing a common methodology for fairer comparison. The work treats aggregate success rate as incomplete and studies partial progress, failure modes, repeatability, recovery behavior, and sensitivity to control conditions.

    MODEL AMODEL B SAME TASK+ CONDITIONS OUTCOMEcompletion · progress BEHAVIORfailure · recovery CONTROLcompliance · repeatability
    Controlled comparisonTask completion · partial progress

Policy analysis instrument

Select an input source, then compare the nominal and controlled-perturbation views.

Input under study
Diagnostic view

Vision Which regions redirect the policy response when visual evidence changes?

Attention routing under input ablationFour cell fields compare the input, attention with all inputs, attention after one input is ablated, and the resulting difference. INPUT FIELD ATTENTION · ALL INPUTS ATTENTION · ONE ABLATED Δ · ROUTED ELSEWHERE
Controlled diagnostic schematic · attribution is not causationIllustrative field—not measured output

Selected systems

Projects and
technical contributions.

Selected work in robot learning, perception, control, multimodal interaction, and physical deployment.

PoseFusion architecture showing view-specific pose graphs, GCN encoders, semantic aggregation, and temporal transformer encodingArchitecture designed across two studies
Research poster · 2023Multi-view learning

Action recognition from partial evidence

I defined the core architecture across two studies: pose graphs encode spatial structure through GCNs, transformers model action over time, and attention fuses incomplete evidence across camera views. The work supported privacy-preserving fall detection and multi-view recognition under occlusion—then became the architectural starting point for my multimodal anomaly-detection work.

0.82 F1 under occlusion3 views fused semanticallyGCN → Transformer spatial to temporal
Chapter 03 · 2023Connected vehicles

Software-latched multi-robot autonomy

A safety-aware leader–follower stack combining learned detection, feature-based tracking, monocular depth, Kalman state estimation, cooperative sensing, and explicit engage–disengage transitions.

Three-actuator soft robot positioned along a pipe surface Prototype still
Chapter 01 · 2017Soft robotics · published

Pneumatic wall-climbing robot

A lightweight inspection prototype built around three single-bending pneumatic actuators—moving compliance from a control objective into the robot’s physical design.

View publication
Real UAV deployment
Chapter 03 · 2024Multimodal perception → control

Track Anything Rapter · TAR

One target.
Many ways to specify it.

I co-developed a ROS 2 aerial system that accepts a click, box, image, or text query; combines foundation models for detection and segmentation; then closes the loop through visual servoing on a PX4-enabled VOXL2 M500 drone.

  1. 01Specifytext · image · click · box
  2. 02PerceiveDINO · CLIP · SAM
  3. 03Tracktarget pose · re-detection
  4. 04Actvisual servoing · PX4
Chapter 02 · 2021–23Formative research exploration

Multimodal human instruction

Learning robot actions from demonstrations, speech, and gesture

Using the NatSGD dataset provided by my PhD guide, Snehesh Shrestha, I implemented a kitchen-object detector, a language-conditioned recurrent policy, and two experimental gesture pathways. The notebooks establish working model and training prototypes, while task success and the benefit of gesture remain unvalidated.

Contribution boundary: I did the notebook implementation; the NatSGD dataset and paper belong to Shrestha and his co-authors. The model builds on Stepputtis et al.'s language-conditioned imitation-learning method.

NatSGD overview showing synchronized speech, gesture, visual observations, and robot demonstrations

Related research contextNatSGD later formalized speech, gesture, and synchronized robot demonstrations at UMD. Image: Shrestha et al., CC BY 4.0.

Technical module · U-ASTROD

Multimodal anomaly detection

For the ASTROD research program, I architected the fusion pipeline that encoded LiDAR, odometry, and network traffic independently before combining their learned representations in a shared anomaly detector.

The published system paired modality-specific spatiotemporal encoders with feature-wise reconstruction scoring. On real robot attacks, that scoring increased multimodal detection accuracy from 72% to 98% and retained 97% accuracy with half the training data.

U-ASTROD process flow with actual network, odometry, and LiDAR samples feeding modality-specific encoders and a shared anomaly detector
U-ASTROD process flow · redrawn from the published architecture

Research map

One program, tested through different systems.

The projects below are not separate interests. Each isolates a capability needed for robots to perceive evidence, interpret people, learn skills, and act reliably.

Independent products · 2025–present

Independent software projects.

Tools I build outside my robotics work to explore local-first software, technical knowledge management, and human–agent workflows.

01 / 02Use arrows, tabs, or swipe

UN
Local-first · self-hosted · open source

UltraNote Lite

“The notebook I wanted
but couldn’t find.”

A single-user workspace for daily notes, tasks, projects, journaling, and research capture. It occupies the space between heavyweight cloud tools and bare Markdown: fast enough for every thought, structured enough for long-running work, and owned entirely by the person using it.

Local firstNo cloud, account, telemetry, or data lock-in.Capture firstOne shortcut routes a task, note, journal entry, or idea.Agent readyPeople and coding agents share one durable project record.
01Capture anywhereDesktop · PWA · share sheet · bookmarklet
02Organize deliberatelyDaily · projects · research · wiki-links
03Keep the source of truthOne Node process · one owned data file
04Collaborate with agentsContext · query · search · validated updates

Project archive

Additional work.

A filterable index of research systems, prototypes, course projects, and independent software.

04
Connected autonomy

Auto-Platoon ↗

Bio-inspired leader–follower coordination using software latching.

05
Multimodal UAV

Track Anything Rapter ↗

Click, image, text, and box-conditioned tracking with foundation models and visual servoing.

TAR multimodal detection and tracking examples
12
Egocentric interaction

AI Cricket Coach ↗

Action recognition and immersive feedback for batting practice.

13
Product intelligence

My Style ↗

An AI-assisted shopping and personal-style application.

16
Cyber-physical security · published

Multimodal Anomaly Detection ↗

Modality-specific LiDAR, odometry, and network encoders fused for abnormal-behavior detection.

U-ASTROD multimodal fusion architecture
17
Local-first product

UltraNote Lite ↗

A self-hosted workspace for daily capture, research, projects, and human–agent collaboration.

UltraNote research workspace
19
Local-first product

Intra-Chat ↗

A LAN-hosted team hub for conversation, project knowledge, files, and equipment records.

Intra-Chat team conversation interface

Research record

Selected publications.

Peer-reviewed publications, preprints, and posters, with links to the available manuscripts and project material.

2021 · Chapter 01Forecasting

A Univariate Data Analysis Approach for Rainfall Forecasting

V. P. Tharun, P. Ramya, S. Renuga Devi

How can lightweight time-series analysis provide useful early warning of unexpected rainfall?

Develops an efficient and implementable approach for short-term rainfall forecasting.

2019 · Chapter 01Regression

Prediction of Rainfall Using Data Mining Techniques

V. P. Tharun, R. Prakash, S. R. Devi

How do regression and statistical models compare when predicting regional rainfall intensity?

Evaluates multiple regression approaches using relative error on rainfall data from Coonoor.

2018 · Chapter 01Classification

A Comparative Study of Various Classification Techniques to Determine Water Quality

R. Prakash, V. P. Tharun, S. R. Devi

Which classical learning method best characterizes groundwater quality from measured samples?

Compares decision trees, K-nearest neighbours, and support-vector machines on groundwater data.

2018 · Chapter 01Assessment

Water Quality Assessment Using Data Mining Techniques

P. Ramya, V. P. Tharun, S. Renuga Devi

How can data-mining methods turn raw water measurements into an actionable quality assessment?

Applies classical learning techniques to environmental water-quality analysis.

2018 · Chapter 01Time series

Rainfall Forecasting Using Time Series Analysis

V. P. Tharun, P. Ramya, S. Renuga Devi

What temporal structure in historical rainfall can be converted into a practical forecast?

Uses time-series analysis to model and forecast rainfall patterns.

2017 · Chapter 01Soft robotics

Wall-Climbing Robot Using Soft Robotics

L. P. Pratap, P. M. Shailendrasingh, A. Anand, V. P. Tharun

Can compliant mechanisms produce a lightweight inspection robot for difficult pipe surfaces?

Develops a soft-robotic wall-climbing prototype for inspection inside and outside pipes.

View complete Google Scholar profile

Experience

Research and engineering roles.

01

Present

R&D Engineer · AI & Robotics

Bimanual deformable-object manipulation, multimodal perception, force-aware compliance, demonstration-driven policies, digital twins, sim-to-real evaluation, and ROS 2 deployment.

02

University of Maryland

Robotics & AI Researcher

Mobile-robot navigation security, abnormal-behavior detection, action recognition, and multimodal human–robot interaction.

03

Indian Institute of Science

Research Assistant

Human-collaborative agricultural robotics, manipulation, and perception-driven precision weeding.

04

Wipro Innovations Lab

AI Team Lead

Computer vision, mobile AI, egocentric interaction, mixed reality, and rapid prototyping across emerging technologies.

Technical storytelling

Writing on robotics, AI, and their human context.

I write about the assumptions behind models and demonstrations, the engineering between a benchmark and deployment, and the human consequences of technical choices.

Human in the Loom is an independent publication for longer essays on these themes.

Human in the LoomIndependent publication ↗

Research direction

Robots that can understand people and learn from them.

My long-term aim is to develop algorithms that allow robots to interpret human intent in ways that support action and learning.

My work approaches this problem through policy understanding, representation learning, world models, embodied experience, human interaction, continual learning, and rigorous evaluation.

I am particularly interested in methods that remain interpretable and testable when they move from offline benchmarks to physical systems.

I hold an M.Eng. in Robotics from the University of Maryland and bring experience across academic research, industrial R&D, and real-world robotics deployment.

Collaborate

Research and engineering
collaboration.

Contact me