ECCV 2026

FlexiAvatar

Unified 3D Gaussian Human Avatars Under Arbitrary Body Visibility

Yihalem Yimolal Tiruneh1, Muhammad Salman Ali1, Uyoung Jeong1, Muneeb A. Khan1, MD Khalequzzaman Chowdhury Sayem1,
Allanur Bayramgeldiyev1, Binod Bhattarai2,3,4, Seungryul Baek1
1UNIST    2University of Aberdeen    3University College London    4Fogsphere (Redev.AI Ltd)

Paper, arXiv & code links coming soon.

Monocular InputDriving PoseAvatar

FlexiAvatar reconstructs an animatable 3D Gaussian avatar from a single monocular video — browse the full body, upper body, and head only inputs above, all handled by one unified pipeline.

Abstract

Reconstructing animatable 3D human avatars from monocular video is a fundamental problem in computer vision with broad applications in AR/VR and digital content creation. Existing approaches typically couple parametric body models with neural rendering or 3D Gaussian Splatting and optimize all body regions jointly from short videos, which often degrades fidelity in the visible areas. To overcome this limitation, we introduce FlexiAvatar, a unified framework that explicitly optimizes only the visible body regions, effectively eliminating artifacts arising from unobserved limbs.

Our method integrates occlusion-robust SMPL-X tracking with part-specific residual refinement to capture high-frequency geometric and appearance details. To complete entirely unseen regions (e.g., back views), we leverage a diffusion-based approach to generate texture consistent with the observed appearance. Experiments on full-body (NeuMan, ZJU-MoCap, WildAvatar), upper/half-body (talk-show clips), and head-only (INSTA) inputs show that FlexiAvatar delivers consistently higher reconstruction quality, outperforming state-of-the-art methods by an average PSNR improvement of approximately 3% across datasets. Finally, by restricting optimization to observed regions, our method reduces the effective number of Gaussians that must be optimized and rendered, leading to reduced runtime and memory overhead in partial-visibility scenarios.

FlexiAvatar results across full-body, upper-body and head-only inputs.
A single FlexiAvatar pipeline handles the entire visibility spectrum. Across full-body (Vid2AvatarPro), upper-body (GUAVA) and head-only (RGBAvatar) settings, FlexiAvatar reconstructs a clean canonical avatar and renders sharper, artifact-free results.

Method

FlexiAvatar overview pipeline.
Overview of the FlexiAvatar pipeline. From a monocular input we perform visibility-aware SMPL-X registration, augment with diffusion-generated multi-view videos, decode a triplane-conditioned hybrid mesh–Gaussian representation, apply visibility-aware optimization and part-specific residual refinement, and finally animate and render via 3D Gaussian Splatting.

Video

Animatable avatars reconstructed by FlexiAvatar, driven by novel poses and expressions.

Driving PoseAvatar

Qualitative Comparisons

Full Body — NeuMan

Qualitative comparison on NeuMan.
Against Vid2Avatar, Vid2AvatarPro and ExAvatar, FlexiAvatar recovers sharper clothing textures and more accurate hand structure.

Upper Body — TalkShow

Qualitative comparison on TalkShow.
Full-body-assuming methods degrade under partial input: GART produces severe artifacts and ExAvatar blurs high-frequency detail. FlexiAvatar preserves facial sharpness and fine hand structure, outperforming the upper-body specialist GUAVA.

Full Body — ZJU-MoCap

Qualitative comparison on ZJU-MoCap.
On ZJU-MoCap, FlexiAvatar yields clearer skin textures and more consistent limb geometry than GauHuman and ToMiE across multiple subjects.

BibTeX

@inproceedings{tiruneh2026flexiavatar,
  author    = {Tiruneh, Yihalem Yimolal and Ali, Muhammad Salman and Jeong, Uyoung and
               Khan, Muneeb A. and Sayem, MD Khalequzzaman Chowdhury and
               Bayramgeldiyev, Allanur and Bhattarai, Binod and Baek, Seungryul},
  title     = {FlexiAvatar: Unified 3D Gaussian Human Avatars Under Arbitrary Body Visibility},
  booktitle = {European Conference on Computer Vision (ECCV)},
  year      = {2026},
}