Representation
learning
What has a policy learned, and which inputs influence its actions?
Robotics R&D Engineer
My research takes inspiration from how people learn, adapt, and develop skills. I work on robot learning, multimodal models, and the systems needed to study them on physical robots. My long-term aim is to build robots that improve through experience and understand human intent with the contextual sensitivity people use to understand one another—and act or learn from that understanding.
What has a policy learned, and which inputs influence its actions?
How can embodied demonstrations become useful robot experience?
How can predicted outcomes and inverse dynamics support action learning?
How can human instruction and correction help robots acquire skills?
Background
My earlier work in perception, multimodal systems, and physical robotics provides the technical basis for my current research in human-inspired robot learning.
Work in statistical learning, computer vision, soft robotics, human-facing AI, and ROS established experience in turning observations into decisions and actions.
Rainfall · Water quality · Soft robot · My Style · Try-on · AI Cricket · ROSAt IISc and UMD, I worked on depth estimation, precision weeding, and an exploratory imitation-learning system using demonstrations, speech, and gesture.
Depth · Precision weeding · ROS 2 · Speech + gesture + demonstrationI developed systems that combined partial evidence across sensors or viewpoints and connected perception to control, coordination, or anomaly detection.
PoseFusion · Fall detection · Auto-Platoon · TAR · Multimodal anomaly detectionMy current work brings these threads together in bimanual manipulation: teleoperation, multimodal data, compliant control, learned policies, digital twins, and real-world evaluation.
Deformables · Bimanual manipulation · Compliance · Learned policies · Sim-to-realQuestions guiding the research
Present · R&D engineer, AI & robotics
A shared system for studying teleoperation, data collection, learned policies, human intervention, simulation, deployment, and comparative policy evaluation on bimanual manipulation tasks.
I designed and implemented the workflow from teleoperation and synchronized collection through dataset conversion, policy training, evaluation, testing, and deployment.
The physical cell and digital twin use the same operator, recording, review, conversion, and policy interfaces, with separate hardware and physics backends. This supports controlled comparisons between simulation and real deployment.
Interventions and failed rollouts can be recorded, reviewed, converted, and returned to training. Policy evaluation is currently performed manually. I am also developing a common evaluation methodology that goes beyond aggregate success rate to compare model behavior more fairly.
The research platform
Implemented Six shared layers · physical and simulated backends · staged evaluation
Modular ROS 2 infrastructure for bimanual manipulation, integrating teleoperation, perception, simulation, data capture, training, and deployment.
Human takeover, correction capture, review, dataset conversion, retraining, manual evaluation, and controlled policy rollout.
Shared interfaces across the real cell and digital twin, with staged testing and recorded evidence for each policy version.
Each line of work is marked by its current stage of development and validation.
Implemented workflow Human takeover, correction capture, review, conversion, retraining, testing, and deployment are connected. Evaluation is manual while automated policy evaluation is being developed.
Active experiments I use the digital twin to study transfer across sensing, timing, control, contact, and visual variation. Current experiments examine representations and evaluation methods that remain useful when simulated and physical environments do not match closely.
Experiments completed I use attention maps, Grad-CAM, state-vector ablations, prompt comparisons, and comparisons across self-supervised and pretrained representations to study what influences policy outputs. These are diagnostic tests; they do not by themselves establish a causal explanation.
In development We are developing a portable egocentric interface for recording human manipulation demonstrations. It reuses the teleoperation, synchronization, data review, and conversion infrastructure I built for the current manipulation system.
Offline evaluation completed I adapted a published video-action architecture to teleoperation data and completed offline evaluation. Physical evaluation has not yet been performed. I am continuing to study how predictive and inverse-dynamics models can support action learning and transfer.
Methodology in development I am evaluating leading open-source vision-language-action models across tasks and developing a common methodology for fairer comparison. The work treats aggregate success rate as incomplete and studies partial progress, failure modes, repeatability, recovery behavior, and sensitivity to control conditions.
Policy analysis instrument
Select an input source, then compare the nominal and controlled-perturbation views.
Vision Which regions redirect the policy response when visual evidence changes?
Selected systems
Selected work in robot learning, perception, control, multimodal interaction, and physical deployment.
Research and engineering for bimanual manipulation of deformable objects—coupling multimodal perception, force-aware compliant control, human demonstrations, learned action policies, high-fidelity simulation, and real-world ROS 2 deployment.
Open case study

A plant is not a rigid obstacle with a clean bounding box. This system first estimates roots, leaf spread, and bed geometry, then couples learned plant geometry with a hybrid inverse-kinematics controller so the manipulator can reach weeds while reducing unintended contact.
I co-designed this gripper with a mechanical engineer as part of the same research program. Vision-adaptive control positions the tool to reduce contact with non-weed plants, while the end effector combines mechanical plucking with targeted spraying or watering.
Architecture designed across two studiesI defined the core architecture across two studies: pose graphs encode spatial structure through GCNs, transformers model action over time, and attention fuses incomplete evidence across camera views. The work supported privacy-preserving fall detection and multi-view recognition under occlusion—then became the architectural starting point for my multimodal anomaly-detection work.
A safety-aware leader–follower stack combining learned detection, feature-based tracking, monocular depth, Kalman state estimation, cooperative sensing, and explicit engage–disengage transitions.
Prototype still
A lightweight inspection prototype built around three single-bending pneumatic actuators—moving compliance from a control objective into the robot’s physical design.
View publicationTrack Anything Rapter · TAR
I co-developed a ROS 2 aerial system that accepts a click, box, image, or text query; combines foundation models for detection and segmentation; then closes the loop through visual servoing on a PX4-enabled VOXL2 M500 drone.
Multimodal human instruction
Using the NatSGD dataset provided by my PhD guide, Snehesh Shrestha, I implemented a kitchen-object detector, a language-conditioned recurrent policy, and two experimental gesture pathways. The notebooks establish working model and training prototypes, while task success and the benefit of gesture remain unvalidated.
Contribution boundary: I did the notebook implementation; the NatSGD dataset and paper belong to Shrestha and his co-authors. The model builds on Stepputtis et al.'s language-conditioned imitation-learning method.
Related research contextNatSGD later formalized speech, gesture, and synchronized robot demonstrations at UMD. Image: Shrestha et al., CC BY 4.0.
Technical module · U-ASTROD
For the ASTROD research program, I architected the fusion pipeline that encoded LiDAR, odometry, and network traffic independently before combining their learned representations in a shared anomaly detector.
The published system paired modality-specific spatiotemporal encoders with feature-wise reconstruction scoring. On real robot attacks, that scoring increased multimodal detection accuracy from 72% to 98% and retained 97% accuracy with half the training data.

Research map
The projects below are not separate interests. Each isolates a capability needed for robots to perceive evidence, interpret people, learn skills, and act reliably.
Preserve useful structure when views, body regions, or sensor channels disappear.
Translate different forms of instruction into a target that a physical system can act upon.
Connect demonstrations, synchronized datasets, policy training, evaluation, and correction.
Use geometry, depth, simulation, prediction, and inverse dynamics to make action consequences testable.
Make local observations, human correction, compliance, and shared state influence group behavior.
Diagnose what a policy uses, how it fails, and whether a metric reflects physical competence.
Independent products · 2025–present
Tools I build outside my robotics work to explore local-first software, technical knowledge management, and human–agent workflows.
01 / 02Use arrows, tabs, or swipe
“The notebook I wanted
but couldn’t find.”
A single-user workspace for daily notes, tasks, projects, journaling, and research capture. It occupies the space between heavyweight cloud tools and bare Markdown: fast enough for every thought, structured enough for long-running work, and owned entirely by the person using it.



“The operating layer
for a physical team.”
An actively developed, self-hosted team hub for labs, workshops, and robotics crews. It keeps conversation, technical knowledge, files, equipment, and agent-generated context on the local network—without external accounts, cloud dependency, or telemetry.



Project archive
A filterable index of research systems, prototypes, course projects, and independent software.
Bimanual manipulation through multimodal perception, compliance, learned policies, simulation, and real-world deployment.
My language-conditioned policy and gesture-grounding experiments on the NatSGD dataset.

Weed perception, adaptive control, plant-safe hybrid-IK planning, and a co-designed end effector.
Bio-inspired leader–follower coordination using software latching.
Click, image, text, and box-conditioned tracking with foundation models and visual servoing.

Vision-based localization for aerial asset inspection and management.
YOLO-based visual inspection for detecting structural surface defects.
Superpixel segmentation for compact image representation.
Image outpainting through coordinate-conditioned neural representations.
Transfer learning for estimating scene depth from a single camera.
Visual apparel synthesis and personalized style evaluation.
Action recognition and immersive feedback for batting practice.
An AI-assisted shopping and personal-style application.
A bowler agent that adapts to batting styles and field configurations.
A lightweight soft-robotic prototype for pipe inspection.

Modality-specific LiDAR, odometry, and network encoders fused for abnormal-behavior detection.

A self-hosted workspace for daily capture, research, projects, and human–agent collaboration.

An exploratory outdoor-safety concept using asynchronous event streams.
A LAN-hosted team hub for conversation, project knowledge, files, and equipment records.

Research record
Peer-reviewed publications, preprints, and posters, with links to the available manuscripts and project material.
How can domain structure and learned evidence work together to secure autonomous systems?
Connects first-principles system knowledge with data-driven detection for more efficient cyber-physical security.
How can vision and learned task planning make plant–robot interaction more collaborative?
Combines vision, hybrid inverse kinematics, and a feedforward network for precision-farming task execution.
How should autonomous systems be evaluated when abnormal behavior appears across multiple modalities?
Architected the multimodal fusion pipeline within a team framework for anomaly detection and evaluation beyond controlled laboratory conditions.
How can autonomous followers coordinate reliably with a leader in dynamic environments?
Introduces software latching and a layered perception–decision–control architecture for real-time multi-agent coordination.
Can a UAV track what a person means, regardless of how they specify it?
Combines DINO, CLIP, SAM, multimodal queries, visual servoing, and motion control on a PX4-enabled aerial platform.
How can lightweight time-series analysis provide useful early warning of unexpected rainfall?
Develops an efficient and implementable approach for short-term rainfall forecasting.
How do regression and statistical models compare when predicting regional rainfall intensity?
Evaluates multiple regression approaches using relative error on rainfall data from Coonoor.
Which classical learning method best characterizes groundwater quality from measured samples?
Compares decision trees, K-nearest neighbours, and support-vector machines on groundwater data.
How can data-mining methods turn raw water measurements into an actionable quality assessment?
Applies classical learning techniques to environmental water-quality analysis.
What temporal structure in historical rainfall can be converted into a practical forecast?
Uses time-series analysis to model and forecast rainfall patterns.
Can compliant mechanisms produce a lightweight inspection robot for difficult pipe surfaces?
Develops a soft-robotic wall-climbing prototype for inspection inside and outside pipes.
Experience
Present
Bimanual deformable-object manipulation, multimodal perception, force-aware compliance, demonstration-driven policies, digital twins, sim-to-real evaluation, and ROS 2 deployment.
University of Maryland
Mobile-robot navigation security, abnormal-behavior detection, action recognition, and multimodal human–robot interaction.
Indian Institute of Science
Human-collaborative agricultural robotics, manipulation, and perception-driven precision weeding.
Wipro Innovations Lab
Computer vision, mobile AI, egocentric interaction, mixed reality, and rapid prototyping across emerging technologies.
Technical storytelling
I write about the assumptions behind models and demonstrations, the engineering between a benchmark and deployment, and the human consequences of technical choices.
Human in the Loom is an independent publication for longer essays on these themes.
Research direction
Robots that can understand people and learn from them.
My long-term aim is to develop algorithms that allow robots to interpret human intent in ways that support action and learning.
My work approaches this problem through policy understanding, representation learning, world models, embodied experience, human interaction, continual learning, and rigorous evaluation.
I am particularly interested in methods that remain interpretable and testable when they move from offline benchmarks to physical systems.
I hold an M.Eng. in Robotics from the University of Maryland and bring experience across academic research, industrial R&D, and real-world robotics deployment.
Collaborate