01 · Standing and applying material to a wall
“The person is standing and using a trowel to apply something to a wall.”
K700-M ID: AHzKaqaOx6A_000074_000084_p0
Extracting human motion from large-scale web videos offers a scalable solution to the data scarcity issue in character animation. However, some human parts in many video frames cannot be seen due to off-screen captures or occlusions. It brings a dilemma: discarding the data missing any part limits scale and diversity, while retaining it compromises data quality and model performance.
To address this problem, we propose leveraging credible part-level data extracted from videos to enhance motion generation via a robust part-aware masked autoregression model. First, we decompose a human body into five parts and detect the parts clearly seen in a video frame as "credible". Second, the credible parts are encoded into latent tokens by our proposed part-aware variational autoencoder. Third, we propose a robust part-level masked generation model to predict masked credible parts, while ignoring those noisy parts.
In addition, we contribute K700-M, a challenging new benchmark comprising approximately 200k real-world motion sequences, for evaluation. Experimental results indicate that our method successfully outperforms baselines on both clean and noisy datasets in terms of motion quality, semantic consistency and diversity.
Web videos contain rich human motion data, but many frames suffer from occlusions or off-screen captures. Traditional approaches face a dilemma: discarding incomplete data limits dataset scale and diversity, while keeping it compromises quality.
Detecting credible parts using joint confidence from ViTPose
Key Insight: Not all body parts are equally noisy in a given frame. By identifying and leveraging only the reliable (credible) parts, we can dramatically expand usable training data without compromising model quality.
We use ViTPose to obtain per-joint confidence scores, indicating visibility and detection accuracy. Body parts (torso, left/right arm, left/right leg) with average confidence above threshold τ are marked as "credible".
We compress credible part-level motion into a compact latent space using a shared-parameter VAE. This prevents noisy parts from corrupting the learned representations while maintaining spatial consistency across body parts.
Our masked autoregressive model selectively learns from credible tokens while unconditionally masking noisy ones. A diffusion head further refines the output for high-quality, diverse motion generation.
We compare our method with state-of-the-art baselines on both HumanML3D (clean) and K700-M (noisy) datasets. Our method achieves superior performance especially on the challenging noisy dataset.
Table 1: Comparison with state-of-the-art methods on HumanML3D and K700-M datasets. Our method (RoPAR) achieves the best performance on the noisy K700-M dataset across all metrics.
Our method maintains stable performance across different noise levels, while baselines degrade significantly.
Visual comparison shows our method generates motions with richer details and more natural transitions, especially for actions typically filmed in close-ups (e.g., "sitting and playing the drum").
Each pair compares two motions associated with the same text prompt, both retargeted to the Y Bot: Generated is RoPAR's output, while Recovered is the motion recovered from the K700-M video. The two clips in a pair share the same camera and playback length.
“The person is standing and using a trowel to apply something to a wall.”
K700-M ID: AHzKaqaOx6A_000074_000084_p0
“The person is standing and holding a microphone while seemingly in conversation.”
K700-M ID: O8V2pqZV9B4_000106_000116_p0
“The person is standing and scooping a granular substance.”
K700-M ID: l3U-XNqSolc_000153_000163_p0
“The person is standing relatively still, slightly leaning forward.”
K700-M ID: 7CMZSorheKE_000162_000172_p1
“The person is standing, initially stretching their right arm across their body, and then transitions into stretching both arms overhead.”
K700-M ID: yNI-myyQnig_000038_000048_p9
“The person is initially standing still, then engages in a brief interaction and finally raises their arms in celebration.”
K700-M ID: L6XiiIJXGJQ_000000_000010_p11
For cases 01–04, both legs in the recovered K700-M motion were masked as unreliable in every preview frame. The upper-body examples show a visual text/motion mismatch, but their inconsistent arm motion was only partly masked (case 05) or almost entirely retained (case 06); they should not be interpreted as proof that arm masking removed the mismatch.
@inproceedings{li2026ropar,
title={Robust Motion Generation using Part-level Reliable Data from Videos},
author={Li, Boyuan and Zheng, Sipeng and Cao, Bin and Song, Ruihua and Lu, Zongqing},
year={2026}
}