A quiet revolution is unfolding in computer vision, and it is happening in three dimensions. While the public imagination has been captured by chatbots and image generators, researchers have been racing to solve a harder problem: teaching machines to understand the physical world as it truly exists, in full 3D. A comprehensive new survey published in the journal Vicinagearth by Yuenan Hou, Xiaoshui Huang, Shixiang Tang, Tong He and Wanli Ouyang of the Shanghai AI Laboratory maps this rapidly expanding territory, offering the most complete picture yet of how machines learn from 3D data and what that means for everything from self-driving cars to augmented reality.
The appeal of 3D data is easy to grasp. A photograph flattens the world onto a grid of pixels, discarding the very depth information that lets humans judge distance, volume and shape. Three-dimensional data, by contrast, carries precise spatial measurements, allowing an algorithm to know not just that a pedestrian is present but exactly how far away they stand and how much space they occupy. The survey identifies five dominant ways of representing this information: point clouds, voxel grids, depth and normal maps, polygonal meshes and neural fields. Each comes with trade-offs in computational cost, storage demands and expressive power, and choosing the right representation is often the first critical decision in any 3D pipeline.
Point clouds deserve particular attention because they dominate real-world sensing. Produced by LiDAR scanners and depth cameras, a point cloud is simply a set of points floating in space, each carrying a 3D coordinate and sometimes color or intensity. The trouble is that these points are irregular and unordered, which makes them fundamentally incompatible with the convolutional neural networks that revolutionized 2D image analysis. The field’s answer is sparse convolution, a clever trick that performs computation only on non-empty regions of a voxelized space. Because most of a scanned scene is empty air, sparse convolution slashes the computational burden dramatically. Its cousin, sub-manifold sparse convolution, goes further by restricting computation to kernel centers that already contain data, preventing the dilation that would otherwise destroy the input’s sparsity, though the dilation property itself can be useful for giving models contextual awareness of their surroundings.
With these foundations in place, the survey turns to the heart of modern 3D research: pre-training. The idea borrows directly from the playbook that made 2D vision so successful. Models pre-trained on ImageNet’s fourteen million labeled images learned general-purpose features that transferred effortlessly to new tasks. But 3D data lacks an equivalent of ImageNet, largely because annotating millions of 3D scans is prohibitively expensive. The solution has been self-supervised learning, in which models generate their own training signals from raw data. The survey organizes these efforts into three families: contrastive methods, masked autoencoder approaches and rendering-based techniques.
Contrastive learning teaches a network to pull similar data points together in an embedding space while pushing dissimilar ones apart, typically through an InfoNCE loss. In the 3D world this takes several forms. PointContrast, a landmark 3D-to-3D method, encourages networks to learn representations that remain stable under different viewpoints and noise, directly boosting downstream detection and segmentation. Other approaches reach across modalities: CrossPoint and SLidR bridge point clouds and rendered images, transferring knowledge from 2D backbones trained on vast image collections into 3D networks. Perhaps most strikingly, methods like PointCLIP and its successor PointCLIPV2 connect 3D perception to language, projecting point clouds into multi-view images and aligning them with the text embeddings of CLIP-style models. This enables zero-shot recognition, where a model classifies objects it has never been trained on simply by matching visual features to textual descriptions.
Masked autoencoder methods take a different route, inspired by the masked image modeling that powered recent breakthroughs in 2D vision. The core idea is to corrupt or hide parts of the input and force the model to reconstruct them, thereby learning the underlying structure of 3D shapes. VoxelMAE converts unstructured points into voxel grids and predicts the occupancy of masked voxels. PointMAE, borrowing a page from BERT, tokenizes point clouds using farthest point sampling and K-nearest-neighbor grouping, then trains with an L2 loss on masked token features; its descendant PointM2AE adds pyramid architectures that capture both fine geometric detail and high-level semantics, achieving state-of-the-art linear classification on the ModelNet40 benchmark. GD-MAE extends the paradigm to convolutional backbones favored in autonomous driving, with a masking strategy designed to prevent knowledge leakage during downsampling.
The third family is the most visually intuitive: rendering-based pre-training. Differentiable rendering makes the process of drawing an image from a 3D model mathematically transparent, so gradients can flow backward from the rendered image to the 3D representation itself. The Ponder framework describes 3D surfaces within a volume and optimizes them against 2D image projections using a NeRF-like structure, while PonderV2 scales the idea to outdoor scenes and can pre-train both 2D and 3D backbones. According to the survey, this rendering-based approach outperforms both contrastive and masked-autoencoder methods, suggesting that grounding 3D learning in the physics of image formation may be the field’s most promising direction.
These pre-trained representations then feed a rich ecosystem of downstream tasks. Classification, the field’s founding problem, began with PointNet’s pioneering use of shared multi-layer perceptrons on raw points and matured through architectures like DGCNN’s EdgeConv, which builds local graphs to learn edge embeddings. Segmentation assigns a category to every point in a scan and has spawned a fierce architectural competition: MinkowskiUNet applies U-Net-style designs to voxel data, Cylinder3D replaces cubic partitions with cylindrical ones to handle LiDAR’s varying point density, SphereFormer exploits the transformer’s global receptive field, and RangeFormer brings powerful 2D-style transformers to range-view representations. Detection, crucial for autonomous vehicles, spans LiDAR-based methods like PointRCNN and PVRCNN, camera-based systems like BEVFormer that follow Tesla’s vision-centric pipeline, and fusion approaches like PointPainting and LoGoNet that marry the precise geometry of LiDAR with the rich texture of images. Tracking, matching and registration round out the toolkit, with methods like GeoTransformer and robust variants of the classic ICP algorithm enabling machines to align separate 3D scans into coherent wholes.
Progress on all these fronts is measured against a demanding set of benchmarks. ModelNet40 and ShapeNet test synthetic shape understanding, ScanNetV2 and S3DIS probe indoor scene parsing, while KITTI, nuScenes, Waymo and ONCE push algorithms to their limits in real-world driving scenarios containing millions of annotated frames. Evaluation relies on metrics such as mean average precision, computed from intersection-over-union thresholds between predicted and ground-truth boxes, and mean intersection-over-union for segmentation, with registration tasks judged by rotation error, translation error and registration recall. The survey’s careful cataloging of these standards provides newcomers with an immediate map of how the field keeps score.
Looking forward, the authors identify several frontiers that could define the next decade. Large language models, brimming with world knowledge distilled from vast text corpora, are only beginning to be connected to 3D perception through systems like Uni3D-LLM. Knowledge transfer from the data-rich 2D and language domains into the data-starved 3D world remains a wide-open opportunity. Synthetic data could relieve the crushing cost of manual annotation, though bridging the gap between simulated and real scenes will demand advances in domain adaptation and neural rendering. Most ambitiously, the field awaits its own foundation model: while 2D vision has been reshaped by systems like Segment Anything, no equivalent exists for 3D, largely because large-scale 3D benchmarks are still lacking. Building one, the survey argues, could dramatically cut the design and deployment costs of countless 3D applications and pave the way toward a genuine era of 3D artificial general intelligence. For a field that measures the world in three dimensions, the trajectory is unmistakably upward.
Subject of Research: Self-supervised pre-training methods and downstream tasks for 3D point cloud perception
Article Title: Advances in 3D pre-training and downstream tasks: a survey
Article References: Hou, Y., Huang, X., Tang, S., He, T., & Ouyang, W. (2024). Advances in 3D pre-training and downstream tasks: a survey. Vicinagearth, 1(1), Article 6. https://doi.org/10.1007/s44336-024-00007-4
Image Credits: AI Generated
DOI: 10.1007/s44336-024-00007-4
Keywords: 3D perception, point clouds, self-supervised learning, pre-training, masked autoencoders, contrastive learning, sparse convolution, LiDAR, autonomous driving, neural rendering, foundation models, benchmarks
Cite Scienmag News
Blake Davidson. (October 2, 2026). How 3D Pre-Training Is Teaching Machines to See the World in Depth. Scienmag. https://scienmag.com/how-3d-pre-training-is-teaching-machines-to-see-the-world-in-depth/
Blake Davidson. "How 3D Pre-Training Is Teaching Machines to See the World in Depth." Scienmag, 2 October 2026, https://scienmag.com/how-3d-pre-training-is-teaching-machines-to-see-the-world-in-depth/. Accessed 2 October 2026.
Blake Davidson. "How 3D Pre-Training Is Teaching Machines to See the World in Depth." Scienmag. October 2, 2026. https://scienmag.com/how-3d-pre-training-is-teaching-machines-to-see-the-world-in-depth/

