Artificial intelligence systems that recognize people in images—whether tracking individuals across surveillance cameras, detecting pedestrians for self-driving cars, or mapping the joints of a human body—share a stubborn weakness: they are voracious consumers of data, and the data they need is expensive, privacy-sensitive, and often scarce. A comprehensive new survey published in the open-access journal Vicinagearth argues that the solution lies in a family of techniques known as data augmentation, and for the first time maps the entire landscape of these methods as they apply specifically to human-centered computer vision tasks.
The review, led by Wentao Jiang of Beihang University together with colleagues including Shuicheng Yan of Skywork AI, focuses on four pillars of human-centric vision: person re-identification, human parsing, human pose estimation, and pedestrian detection. All four rely on deep neural networks that can excel on their training data yet stumble on unseen images—a phenomenon known as overfitting. The problem is compounded by the nature of the data itself. Annotating human bodies, with their articulated joints, fine-grained clothing boundaries, and occlusions in crowded scenes, is labor-intensive and costly, while privacy concerns often limit how much real footage of people can be collected and shared in the first place.
To organize a sprawling literature, the authors divide augmentation techniques into two broad families. The first, data perturbation, takes existing images and modifies them to create new training examples. The second, data generation, manufactures entirely new samples, either with graphics engines or with generative models. Within each family, the survey draws a further distinction that turns out to be crucial: some operations act on the whole image, while others target the human figure itself, exploiting knowledge of body structure that generic augmentation methods ignore.
Image-level perturbation is the simplest and most widely used branch. Global transformations such as rotation, scaling, noise injection, and style transfer alter the overall appearance of a picture. Style transfer, which blends the content of one image with the visual style of another, has proven particularly valuable for person re-identification, where images captured by different surveillance cameras exhibit systematic color and lighting differences. By stylizing training images across camera styles, models learn to focus on structural features of a person rather than the color cast of a particular camera. The survey’s experiments on the Market-1501 benchmark, which contains more than 32,000 images of over 1,500 identities filmed by multiple cameras, show that such camera-style augmentation lifts performance well above a baseline SVDNet model, with the strongest results coming from methods that combine generative and discriminative training signals.
Region-level perturbations operate more surgically. Techniques such as random erasing, Cutout, and GridMask deliberately remove or obscure rectangular patches of an image, forcing networks to cope with the partial occlusions that pervade real scenes. A related approach replaces selected regions with grayscale patches, encouraging models to become less dependent on color information, which varies wildly with lighting conditions. The authors caution, however, that these methods walk a fine line: too little occlusion provides no benefit, while excessive or unrealistic erasure can destroy the very features a model needs to learn, degrading performance instead of improving it.
The more distinctive contribution of the survey lies in its treatment of human-level perturbation, where augmentations are guided by the anatomy of the body itself. One technique masks body keypoints with background patches to simulate occluded joints, directly training pose estimation networks to recover from missing evidence. Another uses human parsing to segment a person into semantic parts—head, torso, arms, legs—and then recombines them into a pool from which complex scenes of occlusion and interaction can be synthesized. A third method overlays crops of one person onto another to mimic the crowding typical of pedestrian scenes. At the skeletal level, methods such as PoseTrans apply affine transformations to individual limbs after erasing them from the image, generating a rich variety of plausible poses, while 3D approaches like PoseAug adjust posture, body size, viewpoint, and even split-and-recombine upper and lower body configurations in a differentiable framework that is optimized jointly with the pose estimator. On the MS-COCO benchmark, these human-aware methods, particularly PoseTrans, outperformed generic image-level augmentations when applied to an HRNet-W32 backbone, and on 3D datasets such as Human3.6M and MPI-INF-3DHP, augmentation methods like PoseAug and DH-AUG reduced mean per-joint position error relative to unaugmented baselines.
When perturbation is not enough, researchers turn to outright generation. Graphics-engine approaches render synthetic humans into real backgrounds: MixedPeds, for example, automatically calibrates a virtual camera using the vanishing point of a real dataset, estimates pedestrian scales, and spawns synthetic human agents in unannotated images to train pedestrian detectors. Other systems sample the 3D pose space, deform parametric body models, map on clothing textures, and render the results from varied viewpoints and lighting conditions, producing not only RGB images but also ground-truth 2D and 3D poses, depth maps, surface normals, and body-part segmentation—annotations that would be prohibitively expensive to collect by hand. The survey notes that this is especially important for 3D pose estimation, where ground truth is largely confined to indoor motion-capture studios, causing models trained on such data to generalize poorly to the wild.
Generative models offer a second route to new data. Pose-transfer GANs extract skeletal poses and re-pair them with different appearances, synthesizing images of the same identity in novel poses and clothing—a direct boon for person re-identification, where each identity is typically represented by only a handful of images. Frameworks such as PTGAN and pose-transferring ReID systems add similarity-measurement modules or auxiliary guidance networks to keep generated samples realistic and useful for the downstream task. The survey’s comparison on Market-1501 found that DG-Net, which jointly learns discriminative and generative objectives, achieved the highest mean average precision among the augmentation methods tested. Yet GANs carry well-known liabilities, including training instability and mode collapse, in which the generator produces only a narrow slice of the possible variations.
This is where the survey’s forward-looking analysis becomes most striking. The authors identify pre-trained Latent Diffusion Models, exemplified by Stable Diffusion, as the most promising direction for the field. Diffusion models work by iteratively denoising random noise into coherent images, guided by learned priors, and they sidestep the adversarial discriminator entirely, simplifying training while producing diverse, high-quality samples. The survey sketches concrete applications: controllable diffusion models could generate the same person across outfits, poses, lighting conditions, and camera angles for re-identification; pose-guided synthesis systems such as ControlNet and HyperHuman could manufacture human figures in rare or difficult poses for pose estimation; parsing maps could condition the generation of images with varied clothing and body types for human parsing; and pedestrians could be rendered in diverse urban settings, weather conditions, and crowded or partially obscured scenarios for detection. Beyond full generation, diffusion models could also upgrade perturbation and recombination—seamlessly compositing foreground people into new backgrounds with correct lighting and perspective, swapping one human subject for another while preserving scene realism, and applying subtle changes to clothing or surroundings without the artifacts that plague simple copy-paste and inpainting.
The evidence assembled in the survey suggests that augmentation choices matter as much as architecture choices in human-centric vision. On the CrowdHuman pedestrian detection benchmark, methods such as CrowdAug, SimCP, and SAutoAug consistently improved average precision and Jaccard Index over Faster R-CNN and RetinaNet baselines, and qualitative examples show augmented models detecting pedestrians that baseline systems miss entirely. In human parsing, background replacement and copy-paste techniques sharpen instance segmentation, though the authors note that this subfield still lacks standardized datasets, making comparisons difficult. What emerges overall is a clear principle: methods that respect the structure of the human body—its joints, parts, and the way people occlude one another—consistently outperform generic image tricks. As generative models grow more powerful and controllable, the authors argue, the field stands on the verge of training data that is at once abundant, diverse, and realistic, promising vision systems that see people more robustly in the messy, crowded, and unpredictable environments where they actually operate.
Subject of Research: Data augmentation techniques for human-centric computer vision tasks
Article Title: Data augmentation in human-centric vision
Article References: Jiang, W., Zhang, Y., Zheng, S., Liu, S., & Yan, S. (2024). Data augmentation in human-centric vision. Vicinagearth, 1(1), Article 8. https://doi.org/10.1007/s44336-024-00002-9
Image Credits: AI Generated
DOI: 10.1007/s44336-024-00002-9
Keywords: data augmentation, computer vision, person re-identification, human pose estimation, pedestrian detection, human parsing, generative adversarial networks, diffusion models, synthetic data, deep learning, overfitting, Stable Diffusion
Cite Scienmag News
Blake Davidson. (October 1, 2026). How Synthetic Humans Are Teaching AI to See People Better. Scienmag. https://scienmag.com/how-synthetic-humans-are-teaching-ai-to-see-people-better/
Blake Davidson. "How Synthetic Humans Are Teaching AI to See People Better." Scienmag, 1 October 2026, https://scienmag.com/how-synthetic-humans-are-teaching-ai-to-see-people-better/. Accessed 1 October 2026.
Blake Davidson. "How Synthetic Humans Are Teaching AI to See People Better." Scienmag. October 1, 2026. https://scienmag.com/how-synthetic-humans-are-teaching-ai-to-see-people-better/

