Teaching a computer to find the joints of a human body is hard enough in a well-lit photograph. In a near-black image, it can be close to impossible. Extreme darkness brings severe sensor noise, crushed shadows and a near-total loss of the visual cues that pose estimation networks depend on, such as limb contours, joint shading and texture. A new study from researchers at VNU University of Engineering and Technology in Hanoi, Vietnam, published in Neural Computing and Applications, presents a method called WGPose-LL that tackles this problem head-on, and its results suggest that the key to seeing in the dark may lie in learning from images that are already bright.
Human pose estimation, or HPE, is the task of locating anatomical keypoints such as shoulders, elbows, hips and knees in an image, and then connecting them into a skeletal representation of the body. Modern systems based on deep neural networks, from convolutional architectures like stacked hourglass networks and high-resolution representation networks to transformer-based models such as ViTPose, achieve remarkable accuracy on standard benchmarks captured in good lighting. But the authors of the new study point out a persistent and uncomfortable gap: when the lights go out, performance collapses. The reasons are physical as much as computational. Under extreme low light, cameras amplify sensor signals, which multiplies noise; dynamic range is lost; and the fine gradients that delineate a wrist from a background object simply vanish.
Researchers have tried two broad strategies to close this gap. The first is to enhance the dark image before feeding it to a pose estimator, using low-light image enhancement techniques ranging from classical illumination map estimation to modern Retinex-based transformers and diffusion models. The second is domain adaptation, in which a network trained on well-lit data is gradually adapted to the dark domain, often with adversarial training or dual-teacher schemes. Both approaches have limitations that the Vietnamese team identifies clearly. Enhancement pipelines optimize for visual quality or perceptual metrics, not for the specific details a pose model needs, so they frequently fail to recover pose-critical structure in extremely dark scenes. Domain adaptation, meanwhile, struggles to bridge what is an enormous distribution shift between bright and dark imagery, leaving a large residual gap in accuracy.
WGPose-LL departs from both strategies with a different philosophy: instead of fixing the image, let the well-lit data guide the network throughout training. The method rests on two complementary training strategies and one architectural innovation, all designed around the observation that recent datasets now provide paired well-lit and low-light images of the same scenes, yet existing methods still fail to exploit this paired supervision effectively.
The first strategy, called Well-Lit Guided Dual Training, reformulates conventional knowledge distillation. In classic distillation, a large or well-trained teacher network transfers knowledge to a smaller student, typically in a one-way, sequential fashion. WGPose-LL instead trains a well-lit branch and a low-light branch of the network jointly and in parallel. Because the two branches process paired images of the same person under different lighting, the well-lit branch can act as a continuous guide rather than a one-off teacher. Guidance happens at two levels simultaneously: feature-level alignment encourages intermediate representations extracted from dark images to resemble those extracted from their bright counterparts, while output-level alignment pushes the pose predictions from dark images toward the more reliable predictions made on the lit versions. The result is a training signal that teaches the dark branch not just what the final answer should look like, but what the internal reasoning should look like at every stage of the network.
The second strategy, Difficulty-Progressive Training, addresses a subtler problem: optimization on extremely dark data is simply harder. If a network is asked from the first epoch to make sense of nearly black inputs, gradients are noisy and learning can stall or converge to poor solutions. Drawing on the idea of curriculum learning, the authors gradually increase the difficulty of training by constructing composite inputs that blend information from the well-lit and low-light domains. Early in training, the network sees mixtures that are relatively easy to interpret; as training proceeds, the proportion of genuinely dark content increases until the network handles the hardest, most underexposed cases. This progressive schedule eases the optimizer into the low-light domain rather than throwing it into the deep end, which the authors report improves robustness on the most extreme scenes.
The third component is architectural rather than a training trick. The Hierarchical Refinement Module performs iterative multi-scale feature refinement inside the backbone network itself. Inspired by encoder-decoder designs such as U-Net, the module revisits features at multiple scales and refines them iteratively, allowing the network to suppress noise and compensate for underexposure directly within its own representations. This matters because low-light artifacts are not uniform: noise dominates fine, high-frequency details, while underexposure destroys broad, low-frequency structure. A module that refines across scales can address both problems in the same pass, making the network inherently more robust to dark conditions rather than relying on external preprocessing.
The team evaluated WGPose-LL on the ExLPose and ExLPose-OCN benchmarks, datasets specifically designed for human pose estimation in extremely low-light conditions, containing paired well-lit and dark images. Across these benchmarks, the method consistently outperformed existing state-of-the-art approaches, including prior work that used enhancement pipelines, sensor fusion with infrared cameras, and domain-adaptive dual-teacher training. According to the paper, the combined approach establishes new performance baselines for pose estimation in extremely low-light scenarios, substantially narrowing the gap between dark and well-lit performance that has long plagued the field.
The significance of this work extends beyond a leaderboard. Reliable pose estimation in the dark is a prerequisite for a range of safety-critical applications: night-time surveillance and search-and-rescue, autonomous driving in unlit streets, gesture interfaces in dim rooms, sports and medical motion analysis without intrusive lighting, and robotics operating around humans after dark. Previous attempts often required additional hardware, such as the RGB-infrared sensor fusion approaches explored by other groups, which adds cost and complexity. WGPose-LL requires nothing more than standard RGB cameras and paired training data, making it far easier to deploy. The dual-training framework is also, in principle, agnostic to the underlying pose network, suggesting it could be bolted onto a variety of backbone architectures.
There are, of course, caveats. The method depends on paired well-lit and low-light data, which remains scarce; the authors themselves note that the scarcity of dedicated benchmarks is one of the field’s core challenges. How well the approach transfers to domains where no paired bright image exists, such as arbitrary night-time video, is a question for future work. The computing resources for the research were sponsored by Intelligent Integration Co., Ltd. of Vietnam, and the authors, Hoang Hai Pham, Tien Du Pham and Quoc Long Tran, declare no competing interests. Still, the conceptual contribution is clear and likely to influence the field: rather than treating darkness as an image-quality problem to be fixed before inference, WGPose-LL treats it as a learning problem to be solved with guidance, curriculum and architecture working together. As cameras and AI systems are increasingly asked to work around the clock, techniques that let machines see people in the dark may soon move from the research bench into the systems that watch over streets, vehicles and workplaces at night.
Subject of Research: Human pose estimation in extremely low-light conditions using well-lit guided dual training
Article Title: WGPose-LL: a robust method to leverage well-lit images for guiding human pose estimation in low-light conditions
Article References: Pham, H. H., Pham, T. D., & Tran, Q. L. (2026). WGPose-LL: a robust method to leverage well-lit images for guiding human pose estimation in low-light conditions. Neural Computing and Applications, 38(19), Article 786. https://doi.org/10.1007/s00521-026-12481-6
Image Credits: AI Generated
DOI: 10.1007/s00521-026-12481-6
Keywords: human pose estimation, low-light imaging, computer vision, deep learning, knowledge distillation, curriculum learning, domain adaptation, image enhancement, ExLPose dataset, feature refinement, neural networks, surveillance
Cite Scienmag News
Blake Davidson. (October 9, 2026). AI Learns to See Human Bodies in the Dark by Borrowing Light from Bright Images. Scienmag. https://scienmag.com/ai-learns-to-see-human-bodies-in-the-dark-by-borrowing-light-from-bright-images/
Blake Davidson. "AI Learns to See Human Bodies in the Dark by Borrowing Light from Bright Images." Scienmag, 9 October 2026, https://scienmag.com/ai-learns-to-see-human-bodies-in-the-dark-by-borrowing-light-from-bright-images/. Accessed 9 October 2026.
Blake Davidson. "AI Learns to See Human Bodies in the Dark by Borrowing Light from Bright Images." Scienmag. October 9, 2026. https://scienmag.com/ai-learns-to-see-human-bodies-in-the-dark-by-borrowing-light-from-bright-images/

