Tharun V. PuthanveettilPortfolio ↗

Multimodal perception · 2023

Reasoning from
partial evidence.

Can pose-based action models remain informative when the viewpoint changes or important parts of the body disappear?

Two connected studies examine the same representation problem at different scales. PoseFusion preserves each camera view long enough to aggregate complementary evidence. The fall-recognition study isolates what spatial graphs, camera viewpoint, body-region occlusion, and temporal sampling contribute to a binary safety-relevant decision.

PoseFusion authors
Abhijay SinghTharun V. PuthanveettilAniruddh BalramSathish V.Sai S.
Fall studyAbhijay Singh · Tharun V. Puthanveettil
AffiliationUniversity of Maryland, College Park

Research premise

Humans rarely discard every useful observation because one view is blocked. These studies test whether a learned representation can also preserve incomplete, source-specific evidence before combining it through time.

PosterPoseFusion publication ↗RepositoryMulti-view training code ↗RepositoryFall-recognition code ↗
ContributionCore multimodal sequential architecture and experimental framing
PoseFusion architecture showing view-specific pose graphs, graph convolution, aggregation and temporal transformer encoding
PoseFusion architectureView-specific graph encoding → semantic aggregation → temporal reasoning

PoseFusion

Fuse views after they become meaningful.

The system does not flatten all cameras into a single raw tensor. Each view is first represented as a sequence of skeletal graphs. Spatial encoders preserve joint topology; an aggregation module combines view embeddings; a transformer then models the action over time.

01Extract3D skeleton · 25 joints
02Graph encodejoint topology · per view
03Aggregatemean · linear · attention
04Encode timemulti-head transformer
05Classifyaction class

Dataset construction

  • NTU RGB+D skeleton sequences
  • 60 action classes: 40 daily living, 9 health-related, 11 mutual actions
  • 40 subjects and three camera views
  • 25 joints with 3D coordinates per frame
  • Offline preprocessing into fixed sequence tensors

Occlusion protocol

Structured joint groups are masked in selected views to simulate partial evidence rather than random independent dropout. The multi-view loader then pairs complementary views so the aggregation mechanism can be tested under controlled corruption.

The repository exposes both single-view and multi-view model paths, dataset preprocessing, occlusion augmentation, training, validation, and test loaders.

PoseFusion results

The benefit appears when evidence is incomplete.

The clean single-view case remains strong. Multi-view aggregation earns its complexity under occlusion—and the aggregation strategy matters.

0.85best reported aggregation accuracy · attention
0.82multi-view F1 under occlusion
0.76single-view F1 under occlusion
256effective attention embedding in reported ablation

Aggregation comparison from the poster

AggregatorClass F1Class precisionClass recallAccuracy
Mean0.72 · 0.78 · 0.840.61 · 1.00 · 0.840.86 · 0.64 · 0.840.79
Linear0.72 · 0.86 · 0.880.61 · 1.00 · 0.910.86 · 0.76 · 0.840.82
Attention0.66 · 0.90 · 0.840.61 · 1.00 · 0.840.73 · 0.82 · 0.840.85
Interpretation: the result is conditional, not universal. On clean observations, one good view can outperform a noisier multi-view combination. The reported advantage appears when complementary views preserve information that another view has lost.

Pose-based fall recognition

Decompose spatial structure from temporal motion.

The companion study compares a temporal transformer operating on pose coordinates with a GCN–Transformer pipeline that first encodes the skeletal graph. It then deliberately removes body regions, views, and frames to locate failure.

01 · Representation

Transformer only

Pose features enter temporal self-attention directly, establishing whether sequence modeling alone captures enough structure.

02 · Representation

GCN + Transformer

Graph convolution uses skeletal connectivity before temporal encoding, testing whether explicit spatial relationships improve recognition.

03 · Robustness

Structured ablations

View shift, torso/lower/upper-body occlusion, and increasing frame skip reveal which evidence the classifier depends on.

UR Fall dataset

  • 123 samples
  • Front-view fall versus non-fall classification
  • 60 / 30 / 10 train–validation–test split
  • Transformer and GCN–Transformer comparison

NTU RGB+D subset

  • Six selected actions mapped to fall / non-fall
  • Training on one view; evaluation on alternate views
  • Occlusion groups applied by body region
  • Temporal stride increased to stress motion loss

Ablation matrix

The failure pattern is the result.

Top-line classification alone would hide the most useful findings: torso visibility, temporal density, and camera angle do not contribute equally.

ExperimentConditionF1What it suggests
Architecture · URTransformer only0.8933Temporal encoding already captures a strong signal.
Architecture · URGCN + Transformer0.9033Explicit joint topology yields a modest gain on the reported split.
Architecture · NTUTransformer only0.9911The selected binary task is comparatively separable.
Architecture · NTUGCN + Transformer1.0000Perfect score is split-specific, not a deployment claim.
Cross-viewView 2 / View 30.9733 / 0.9198Viewpoint shift is not uniform.
OcclusionTorso masked0.6559Torso geometry is particularly important for this fall representation.
OcclusionLower body masked0.937The classifier retains more discriminative evidence than under torso masking.
Temporal samplingSkip 1 / Skip 110.9033 / 0.6588Sparse sampling removes motion cues the representation cannot reconstruct.
Charts summarizing view, occlusion and temporal-ablation F1 scores
Reconstructed from the two project reports and poster tables · not a new rerun

Contribution and reproducibility

Architecture, implementation, and evidence.

The public repositories include model definitions, single- and multi-view data loaders, preprocessing, structured occlusion, configuration, training, and evaluation code.

I developed the core multimodal sequential architecture used across the studies.

My contribution centered on connecting graph-based spatial encoding, temporal transformers, and view-specific aggregation, and on framing experiments that expose viewpoint and occlusion behavior rather than reporting only one aggregate score.

The results and artifacts are collaborative. Co-authors are therefore named above, and this page does not claim sole ownership of the papers, datasets, or every experiment.

External-validity boundary. The experiments use curated pose sequences and controlled splits. Robustness to real-time pose-estimator failures, unseen environments, demographic variation, clinical use, and deployment latency was not established.

PoseFusion citationAbhijay S., Tharun P., Aniruddh B., Sathish V., Sai S. “Pose Fusion — Multi-View Pose Integration for Comprehensive Action Recognition.” University of Maryland, 2023. DOI: 10.13140/RG.2.2.32418.20169.
Fall-recognition artifactAbhijay Singh and Tharun V. Puthanveettil. “Striking the Balance: Human Pose Estimation Based Optimal Fall Recognition.” Project preprint and open implementation, 2023.
Related line of work

What changes when multimodal representation learning must detect attacks on a moving robot?

Open anomaly detection →