Tharun V. PuthanveettilPortfolio ↗

Cyber-physical security · 2024–2026

When one sensor lies,
the others hold evidence.

Can an autonomous robot detect behavior that is individually plausible in one modality but inconsistent across its physical and network traces?

Across two publications, this research progressed from a three-modal U-ASTROD system—LiDAR, odometry, and network traffic—to a more compact domain-informed detector using LiDAR-derived features and wheel velocities. Both studies evaluate attacks against autonomous robot operation rather than only synthetic tabular data.

Authors
Mahshid NooraniTharun V. PuthanveettilAsim ZoulkarniJack MirenziCharles D. GrodyJohn S. Baras
VenueGameSec 2024 · LNCS 14908 · pp. 306–325
PublishedSpringer · 11 October 2024

Abstract

A stealthy false obstacle may look legitimate in the point cloud. It is harder for that same attack to remain consistent with robot motion and the network activity that delivered it. The detector learns those cross-modal relationships from normal operation.

GameSec 2024Three-modal U-ASTROD study ↗IEEE ROSE 2026Domain-informed two-modal detector ↗
PlatformClearpath Husky · physical and simulated environments
EvidenceAttack runs · modality ablations · cross-environment tests
Complete U-ASTROD process flow with actual LiDAR, odometry and network samples, modality-specific encoders, late fusion and anomaly detection
U-ASTROD process flow · redrawn from the published architectureSpatial encoder · temporal encoders · late fusion · anomaly score

Threat model

A plausible perception can still be inconsistent behavior.

A man-in-the-middle attack intercepts the robot’s ROS point-cloud stream and injects false obstacles. The autonomy stack reacts to a scene that does not exist. LiDAR alone may accept the modified observation; odometry and network traffic provide independent context.

Physical trace

LiDAR geometry

Captures the spatial structure perceived by navigation. It is also the attacked channel in the false-obstacle scenario.

Behavior trace

Odometry through time

Reveals how the robot actually moved in response to the perceived environment.

Cyber trace

Network traffic

Exposes communication patterns that can change when messages are intercepted, altered, or replayed.

Attack construction

The attack is visible at several levels: the cyber interception path, the falsified point cloud, and the robot’s resulting physical response.

Detection target: the model flags abnormality; it does not identify the attack family or prove malicious intent. Diagnosis remains downstream work.

Model

Keep unlike signals separate until they are encoded.

Each modality has different structure and sampling behavior. The architecture therefore avoids early raw-feature concatenation: LiDAR geometry and sequential signals are encoded through modality-appropriate branches before a shared normal-behavior model.

01Windowsynchronized modality sequences
02Encode LiDARgraph/spatial representation
03Encode timeodometry · network sequence
04Fuse lateF₁ ⊕ F₂ ⊕ F₃
05Reconstructautoencoder · anomaly score

Normal-behavior learning

The autoencoder is trained on nominal robot operation. At inference, reconstruction error measures how far the observed multimodal state lies from the learned normal manifold.

anomaly ⇐ ‖V − AE(V)‖² > τ

Why vector-based scoring

A single scalar average can dilute a strong deviation in one feature when many others reconstruct well. The vector-based loss preserves individual feature contributions before the final decision, improving sensitivity to subtle cross-modal inconsistency.

Design choiceAlternativeReason
Modality-specific encodersRaw early concatenationPreserve geometry and temporal structure before fusion.
Late multimodal fusionLiDAR-only detectionAn attacker must remain consistent across partially independent channels.
Unsupervised normal trainingSupervised attack classificationDoes not require examples of every future attack during training.
Vector reconstruction scoreScalar mean lossPrevents local feature deviations from being averaged away.

Experimental system

Attacks were executed against a moving robot.

The evaluation instrumented a custom autonomous UGV platform and collected normal and attacked runs across its physical and cyber traces.

3modalities · LiDAR, odometry, network traffic
50 kgphysical Clearpath Husky platform
1 m/sreported platform speed
92training samples in the half-data experiment
  1. 01

    Instrument nominal autonomy

    Synchronize LiDAR, odometry, and observable network traffic while the Husky executes its navigation stack.

  2. 02

    Learn normal multimodal structure

    Train the modality encoders and shared autoencoder without requiring attack labels.

  3. 03

    Inject physical consequences through a cyber path

    MITM manipulation introduces false obstacle evidence into the ROS point-cloud channel and changes the robot’s navigation behavior.

  4. 04

    Compare reconstruction scoring

    Evaluate standard scalar loss against vector-based feature scoring on the real-platform attack sequences.

  5. 05

    Reduce nominal training data

    Repeat the training condition with half of the original data to test whether multimodal redundancy remains useful when data are constrained.

Published results

Fusion and scoring changed the operating point.

The strongest result is not the headline accuracy alone. It is that cross-modal structure and feature-preserving reconstruction remain effective even after the nominal training set is reduced to 92 samples.

72%standard scalar-loss / baseline condition
98%vector-based multimodal detection accuracy
+26 ptsabsolute accuracy increase in the reported comparison
97%accuracy with 50% training data
Bar chart comparing baseline, vector-based multimodal and half-data anomaly detection accuracy
Published values visualized for comparison · not a new experimental rerun

What each modality learned

PCA projections make the modality argument inspectable. Attack 1 produces clearer separation; the stealthier Attack 2 compresses the LiDAR distinction while odometry and network structure remain useful contextual evidence.

ConditionTraining dataAccuracyInterpretation
Standard scalar reconstruction scoreFull study condition72%Averaging reconstruction error can suppress subtle feature deviations.
Vector-based multimodal scoreFull study conditionUp to 98%Feature-wise contributions improve attack sensitivity in the reported setup.
Vector-based multimodal score50% · 92 samples97%Cross-modal redundancy retained performance under constrained nominal data.

Per-modality feature extraction

These values expose why late fusion matters: the attacked modality can be highly discriminative in one attack and degrade substantially in a stealthier variation, while secondary modalities provide different evidence.

AttackModalityAccuracyF1ARRNPR
Attack 1LiDAR1.0001.0001.0001.000
Attack 1Odometry0.9520.9510.8301.000
Attack 1Network0.9400.9400.9200.957
Attack 2LiDAR0.8600.8700.6001.000
Attack 2Odometry0.9520.9510.8301.000
Attack 2Network1.0001.0001.0001.000

2026 continuation

Domain knowledge reduced the sensing burden.

The IEEE ROSE study retained the central cross-modal idea while using a compact architecture built around LiDAR-derived features and wheel velocities. It expanded evaluation to indoor, outdoor, unseen-attack, temporal-window, and cross-environment conditions.

Authors
Asim ZoulkarniMahshid NooraniJohn S. BarasJack MirenziTharun V. PuthanveettilCharles D. Grody
VenueIEEE ROSE 2026
DOI10.1109/ROSE69181.2026.11570974
>99%reported accuracy in selected in-distribution conditions
≥559 Hzreported inference frequency · Ryzen 9 7950X · W=10
2modalities · LiDAR features + wheel velocities
3attack variations · baseline, stealthy, flooding

Cross-environment evidence

The most useful test is not the strongest in-distribution number. It is whether a detector trained in one environment preserves attack sensitivity in another.

Training → testingNormal W=1Normal W=10Baseline W=1Baseline W=10Stealthy W=1Stealthy W=10
Indoor → outdoor97.7582.1480.5399.3282.4788.68
Outdoor → indoor70.7682.3887.9795.3094.7696.46
Interpretation: temporal aggregation helps several attack conditions but can reduce normal-class or true-positive performance elsewhere. The paper therefore supports condition-dependent operating-point selection—not a universal “longer windows are better” claim.
CitationZoulkarni, A., Noorani, M., Baras, J. S., Mirenzi, J., Puthanveettil, T. V., and Grody, C. D. “Efficient Integration of Domain Knowledge and Data-Driven Methods for Securing Autonomous Cyber-Physical Systems.” IEEE ROSE, 2026. DOI: 10.1109/ROSE69181.2026.11570974.

Contribution and boundaries

A shared result with a specific architectural contribution.

This research belongs to the larger U-ASTROD program and was supported by the Army Research Laboratory under Cooperative Agreement W911NF-23-2-0040.

I architected the multimodal sequential fusion pipeline used by the larger system.

My core contribution was the model architecture that independently represented LiDAR, odometry, and network sequences before fusing them for shared anomaly detection. The robot platform, attack infrastructure, dataset creation, broader research program, and publication were collaborative; all co-authors are credited above.

Limitations. The study uses one UGV and a bounded attack library. Network features assume traffic is observable. The detector reports abnormality rather than attack semantics, and performance under unseen platforms, sensor suites, benign distribution shifts, adaptive attackers, and long-duration operation remains open.

CitationNoorani, M., Puthanveettil, T. V., Zoulkarni, A., Mirenzi, J., Grody, C. D., and Baras, J. S. “Multimodal Anomaly Detection for Autonomous Cyber-Physical Systems Empowering Real-World Evaluation.” GameSec 2024, LNCS 14908, pp. 306–325. DOI: 10.1007/978-3-031-74835-6_15.
Reproducibility upgrade needed: publishing synchronized modality windows, split membership, threshold selection, per-attack confusion matrices, run-to-run variation, and inference timing would allow the figures to be regenerated rather than reconstructed from the paper.
Current continuation

Return to the platform that now applies evidence governance to learned manipulation policies.

Open robot-learning platform →