Reconstructing animatable 3D human avatars from monocular video is a fundamental problem in computer vision with broad applications in AR/VR and digital content creation. Existing approaches typically couple parametric body models with neural rendering or 3D Gaussian Splatting and optimize all body regions jointly from short videos, which often degrades fidelity in the visible areas. To overcome this limitation, we introduce FlexiAvatar, a unified framework that explicitly optimizes only the visible body regions, effectively eliminating artifacts arising from unobserved limbs.
Our method integrates occlusion-robust SMPL-X tracking with part-specific residual refinement to capture high-frequency geometric and appearance details. To complete entirely unseen regions (e.g., back views), we leverage a diffusion-based approach to generate texture consistent with the observed appearance. Experiments on full-body (NeuMan, ZJU-MoCap, WildAvatar), upper/half-body (talk-show clips), and head-only (INSTA) inputs show that FlexiAvatar delivers consistently higher reconstruction quality, outperforming state-of-the-art methods by an average PSNR improvement of approximately 3% across datasets. Finally, by restricting optimization to observed regions, our method reduces the effective number of Gaussians that must be optimized and rendered, leading to reduced runtime and memory overhead in partial-visibility scenarios.
Animatable avatars reconstructed by FlexiAvatar, driven by novel poses and expressions.
@inproceedings{tiruneh2026flexiavatar,
author = {Tiruneh, Yihalem Yimolal and Ali, Muhammad Salman and Jeong, Uyoung and
Khan, Muneeb A. and Sayem, MD Khalequzzaman Chowdhury and
Bayramgeldiyev, Allanur and Bhattarai, Binod and Baek, Seungryul},
title = {FlexiAvatar: Unified 3D Gaussian Human Avatars Under Arbitrary Body Visibility},
booktitle = {European Conference on Computer Vision (ECCV)},
year = {2026},
}