DCHM

Abstract

Multiview pedestrian detection typically involves two stages: human modeling and pedestrian localization. Human modeling represents pedestrians in 3D space by fusing multiview information, making its quality crucial for detection accuracy. However, existing methods often introduce noise and have low precision. While some approaches reduce noise by fitting on costly multiview 3D annotations, they often struggle to generalize across diverse scenes. To eliminate reliance on human-labeled annotations and accurately model humans, we propose Depth-Consistent Human Modeling (DCHM), a framework designed for consistent depth estimation and multiview fusion in global coordinates. Specifically, our proposed pipeline with superpixel-wise Gaussian Splatting achieves multiview depth consistency in sparse-view, large-scaled, and crowded scenarios, producing precise point clouds for pedestrian localization. Extensive validations demonstrate that our method significantly reduces noise during human modeling, outperforming previous state-of-the-art baselines. Additionally, to our knowledge, DCHM is the first to reconstruct pedestrians and perform multiview segmentation in such a challenging setting.

Method

Human Modeling. We represent pedestrians as collections of segmented Gaussian primitives to enable multiview detection. Our pipeline reconstructs and segments pedestrians in challenging sparse-view, large-scale, and occluded environments.

Overview of the framework. The proposed multiview detection pipeline consists of separate training (left) and inference (right) stages. During training, human modeling optimization refines mono-depth estimation for multiview consistent depth prediction, via pseudo-depth generation, mono-depth fine-tuning, and detection compensation. Specifically, we leverage superpixel to improve the Gaussians optimization in sparse-view setting. During inference, the optimized mono-depth produces Gaussians to model humans that are segmented and clustered to detect pedestrians, shown as blue points on the BEV plane.

BibTeX

@article{ma2025dchm, title={DCHM: Depth-Consistent Human Modeling for Multiview Detection}, author={Ma, Jiahao and Wang, Tianyu and Liu, Miaomiao and Ahmedt-Aristizabal, David and Nguyen, Chuong}, journal={arXiv preprint arXiv:2507.14505}, year={2025} }

DCHM: Depth-Consistent Human Modeling

for Multiview Detection

DCHM fuses sparse multiview images with depth-consistent, superpixel-wise Gaussian splatting to create accurate, label-free 3D pedestrian models that surpass prior methods.

Abstract

Method

Comparison

Results

BibTeX