<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>SemanticKITTI &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/semantickitti/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Wed, 30 Sep 2026 17:57:37 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>SemanticKITTI &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New Dual-Stream Training Strategy Sharpens 3D LiDAR Segmentation for Autonomous Driving</title>
		<link>https://scienmag.com/new-dual-stream-training-strategy-sharpens-3d-lidar-segmentation-for-autonomous-driving/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 17:57:37 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[3D LiDAR data processing]]></category>
		<category><![CDATA[3D scene understanding]]></category>
		<category><![CDATA[advanced deep learning for autonomous vehicles]]></category>
		<category><![CDATA[autonomous driving]]></category>
		<category><![CDATA[autonomous driving perception]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[cross-attention]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[dual-stream training strategy]]></category>
		<category><![CDATA[knowledge distillation]]></category>
		<category><![CDATA[large sparse convolutional neural networks]]></category>
		<category><![CDATA[LiDAR point cloud classification]]></category>
		<category><![CDATA[LiDAR segmentation]]></category>
		<category><![CDATA[LiDAR segmentation benchmark performance]]></category>
		<category><![CDATA[LiDAR semantic segmentation]]></category>
		<category><![CDATA[multi-modal sensor data fusion]]></category>
		<category><![CDATA[nuScenes]]></category>
		<category><![CDATA[point clouds]]></category>
		<category><![CDATA[real-time 3D environment mapping]]></category>
		<category><![CDATA[SemanticKITTI]]></category>
		<category><![CDATA[semanticKITTI and nuScenes datasets]]></category>
		<category><![CDATA[sparse convolution]]></category>
		<category><![CDATA[voxelization]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=217770</guid>

					<description><![CDATA[Researchers at Shenzhen Technology University report a LiDAR segmentation framework that reaches 74.4 percent mean IoU on SemanticKITTI and 83.5 percent on nuScenes by combining large sparse multi-directional convolutions with training-time dual-stream knowledge distillation from images and serialized point clouds.]]></description>
										<content:encoded><![CDATA[<p>Self-driving cars, delivery robots, and urban mapping platforms all depend on the same fundamental perceptual act: making sense of the millions of laser reflections that a LiDAR sensor scatters across the world every second. Turning that raw swarm of three-dimensional points into labeled categories—road, car, pedestrian, vegetation—is the task of semantic segmentation, and it remains one of the most demanding problems in 3D scene understanding. A team of researchers at Shenzhen Technology University, led by Ye Gu and Lixin Liang of the School of Artificial Intelligence together with Gengliang Chen of the Sino-German College of Intelligent Manufacturing, now reports a new framework that pushes the accuracy of LiDAR segmentation to a mean intersection-over-union of 74.4 percent on the SemanticKITTI benchmark and 83.5 percent on nuScenes, two of the most widely used datasets in the field. The work, published open access in Complex &amp; Intelligent Systems, combines an unusually expressive convolutional backbone with a training-time trick that borrows knowledge from other sensor modalities without slowing the network down at inference.</p>
<p>The first half of the innovation lies in the backbone architecture, which the authors call a 3D large sparse multi-directional convolution. To understand why this matters, it helps to recall how modern LiDAR networks process point clouds. Raw LiDAR data is irregular: each laser return is a point floating in space, with no grid structure that a conventional convolutional neural network can slide over. The standard solution is voxelization—dividing space into small cubic cells and treating occupied cells as sparse entries in a three-dimensional tensor. Sparse convolutions then operate only on the occupied cells, which keeps computation tractable even when a single scan contains well over a hundred thousand points.</p>
<p>The catch is that ordinary sparse convolutions use small kernels, typically limited to a few neighboring voxels. That gives each layer a narrow receptive field, so the network must stack many layers before any single voxel can &#8220;see&#8221; enough surrounding context to decide whether it belongs to a curb, a wall, or a distant vehicle. The Shenzhen team attacks this bottleneck with dynamic large sparse kernels, which allow the convolutional footprint to grow dramatically while still operating only on the sparse set of occupied voxels. Because the kernel is dynamic, its shape can adapt to the local geometry of the scene rather than remaining fixed, letting the network concentrate its receptive field along structures such as elongated roads or the flat planes of building facades.</p>
<p>Large kernels alone, however, can still miss context that does not align with the kernel&#8217;s orientation. The second ingredient of the backbone, multi-directional convolution, addresses this by sweeping the receptive field along several directions at once. The combination creates what the authors describe as large effective receptive fields: regions of the scene over which a single computational unit can integrate evidence. In practice, this means that a voxel at the edge of a scan can draw on information from far across the street, which is precisely the kind of long-range context needed to resolve ambiguous geometry in outdoor driving scenes.</p>
<p>Before any of this convolution happens, the raw points must be organized, and here the framework introduces a further refinement called radially non-uniform cylindrical voxelization. Standard cubic voxelization treats all regions of space equally, but LiDAR sensors do not sample the world equally. Points near the sensor are dense, while points at the edge of the sensor&#8217;s range are spread thinly, so a uniform grid leaves near-field cells crowded and far-field cells nearly empty. Cylindrical voxelization, organized around the sensor&#8217;s axis, naturally matches the geometry of the scan, and making the radial divisions non-uniform—finer near the sensor, coarser farther out—balances the number of points per voxel across the entire scene. The result is a more even representation that prevents the network from being dominated by nearby structures and helps it learn features that transfer across distances.</p>
<p>The second half of the framework is where the approach becomes genuinely distinctive. During training, the 3D backbone is accompanied by two auxiliary streams: one that processes 2D RGB images, and one that processes the point cloud after it has been serialized into a 1D sequence using a space-filling curve. Space-filling curves, such as the curves that wind through a volume while visiting neighboring cells consecutively, flatten a 3D structure into an ordered sequence that preserves local adjacency, allowing lightweight 1D networks to extract patterns from the point cloud. The RGB stream, meanwhile, captures the appearance information that cameras provide but LiDAR cannot—texture, color, and fine visual boundaries.</p>
<p>Knowledge from these two auxiliary modalities is then distilled into the 3D backbone through a multi-layer bidirectional polarity-aware linear cross-attention mechanism. Cross-attention lets each stream query the others, so features from the image stream and the serialized point stream can inform the features being learned by the 3D convolutional network. The bidirectional design allows information to flow in both directions between streams, and the polarity-aware component distinguishes the sign or orientation of the features being aligned, which helps the mechanism match corresponding structures across modalities that represent the world in very different ways. Crucially, all of this exchange happens only during training. Once the network has been trained, the auxiliary streams are discarded, and the deployed model is the 3D backbone alone—meaning the extra accuracy gained from images and serialized point sequences comes at zero additional inference cost.</p>
<p>This zero-cost distillation is the aspect of the work with the clearest practical significance. Real-time perception systems on autonomous vehicles operate under strict computational budgets, and any module that must run on board—however accurate—must also be fast. Techniques that fuse camera and LiDAR features at inference time can improve accuracy but add latency and hardware complexity. By confining the multi-modal interaction to the training phase, the Shenzhen framework captures much of the benefit of multi-modal learning while keeping the deployed model as lean as a single-modality network. For engineering teams weighing the cost of additional sensors and compute on a production vehicle, that distinction could be decisive.</p>
<p>The reported benchmark numbers place the framework among the strongest results on both datasets. SemanticKITTI, collected in Karlsruhe, Germany, is the canonical benchmark for outdoor LiDAR segmentation, with sequences spanning tens of kilometers of urban and highway driving; nuScenes, developed for autonomous driving research, pairs LiDAR with camera, radar, and other sensors across a thousand scenes in Singapore and Boston. Mean intersection-over-union, the standard accuracy metric, measures the overlap between predicted and ground-truth labels for each class and averages across them, so gains of even a few points are meaningful. Achieving 74.4 percent on SemanticKITTI and 83.5 percent on nuScenes reflects the combined effect of the large receptive fields, the balanced voxelization, and the dual-stream distillation.</p>
<p>The research was supported by the Shenzhen Science and Technology Program under grants JCYJ20220818102215034 and 20231129112637001, and in part by the National Key Research and Development Program of China under grant 2024YFB4709503. The article was received on 15 June 2026, accepted on 7 September 2026, and published on 28 September 2026 under a Creative Commons license that permits non-commercial sharing with attribution. As with any benchmark-driven advance, the ultimate test will be how the framework transfers to the messier conditions of production deployments—different sensor configurations, adverse weather, and novel cities—but the paper offers a clear demonstration that a 3D network can be taught by its richer siblings during training and still run solo, and fast, once it hits the road.</p>
<p><strong>Subject of Research:</strong> Knowledge distillation and large sparse multi-directional convolution for 3D LiDAR point cloud semantic segmentation</p>
<p><strong>Article Title:</strong> Dual stream knowledge distillation for 3D LiDAR segmentation model with large sparse multi-directional convolution</p>
<p><strong>Article References:</strong> Gu, Y., Xiao, H., Chen, G., &amp; Liang, L. (2026). Dual stream knowledge distillation for 3D LiDAR segmentation model with large sparse multi-directional convolution. <em>Complex &amp;amp; Intelligent Systems</em>. <a href="https://doi.org/10.1007/s40747-026-02523-w" rel="noopener noreferrer">https://doi.org/10.1007/s40747-026-02523-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s40747-026-02523-w" rel="noopener noreferrer">10.1007/s40747-026-02523-w</a></p>
<p><strong>Keywords:</strong> LiDAR segmentation, point clouds, knowledge distillation, sparse convolution, autonomous driving, SemanticKITTI, nuScenes, 3D scene understanding, cross-attention, voxelization, computer vision, deep learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">217770</post-id>	</item>
		<item>
		<title>Object-Based Semantic Descriptors Push Robot Loop Closure Beyond Close Quarters</title>
		<link>https://scienmag.com/object-based-semantic-descriptors-push-robot-loop-closure-beyond-close-quarters/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 23:40:21 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced robot localization techniques]]></category>
		<category><![CDATA[error correction in robot positioning]]></category>
		<category><![CDATA[improving navigation in complex environments]]></category>
		<category><![CDATA[LiDAR]]></category>
		<category><![CDATA[localization]]></category>
		<category><![CDATA[loop closure detection]]></category>
		<category><![CDATA[loop closure detection in mobile robotics]]></category>
		<category><![CDATA[mapping drift prevention]]></category>
		<category><![CDATA[mobile robotics]]></category>
		<category><![CDATA[object semantic scan context (OSSC)]]></category>
		<category><![CDATA[object semantics]]></category>
		<category><![CDATA[Object-based semantic descriptors]]></category>
		<category><![CDATA[place recognition]]></category>
		<category><![CDATA[place recognition challenges]]></category>
		<category><![CDATA[point cloud]]></category>
		<category><![CDATA[RELLIS-3D]]></category>
		<category><![CDATA[research in Singapore for robotic mapping]]></category>
		<category><![CDATA[scan context]]></category>
		<category><![CDATA[semantic scene understanding for robots]]></category>
		<category><![CDATA[semantic segmentation]]></category>
		<category><![CDATA[SemanticKITTI]]></category>
		<category><![CDATA[simultaneous localization and mapping (SLAM)]]></category>
		<category><![CDATA[SLAM]]></category>
		<category><![CDATA[urban and off-road navigation accuracy]]></category>
		<category><![CDATA[visual and object recognition in robotics]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=204004</guid>

					<description><![CDATA[Researchers in Singapore have developed a semantic object-based descriptor called OSSC that improves loop closure detection for robots, enabling accurate place recognition between spatially separated lidar scans on urban and off-road benchmarks.]]></description>
										<content:encoded><![CDATA[<p>One of the most stubborn problems in mobile robotics has just received a promising new solution. When a robot drives through a city, a warehouse, or an off-road trail, it must constantly ask itself a deceptively simple question: have I been here before? Answering that question correctly is the essence of loop closure detection, the process by which a robot recognizes a previously visited place and uses that recognition to correct the accumulated errors in its estimated position. Without reliable loop closure, even the most sophisticated simultaneous localization and mapping systems drift slowly but inevitably away from reality, producing warped maps that become useless for navigation. A team of researchers working in Singapore has now introduced a descriptor called Object Semantic Scan Context, or OSSC, which promises to make loop closure detection dramatically more accurate, especially in the difficult situations where conventional methods tend to fail.</p>
<p>The new approach, described in a paper published in the journal Autonomous Robots by Dhruv Kumarjiguda of Nanyang Technological University and colleagues at the Institute for Infocomm Research, part of the Agency for Science, Technology and Research in Singapore, tackles a specific and costly weakness of existing techniques. Most state-of-the-art loop closure methods depend on the robot physically revisiting a location in close proximity to where it was before. In other words, the robot must essentially travel back to nearly the exact same spot before the system can confidently declare that a loop has been closed. That requirement forces robots to perform unnecessary traversals of their environments, wasting time and energy, and it leaves a wide band of scenarios in which two scans of the same neighborhood, taken from moderately different vantage points, are simply not recognized as describing the same place.</p>
<p>OSSC departs from the conventional recipe in a fundamental way. Instead of encoding only the geometric structure of the environment, the raw shapes and distances captured by a lidar sensor as a three-dimensional point cloud, the new descriptor layers semantic information into the representation. Modern perception systems can label individual points in a lidar scan according to the object they belong to: this cluster is a car, that one is a tree, another is a building, a pedestrian, or a traffic sign. OSSC exploits these labels by organizing the description of a scene around prominent external reference points that the authors call Main Objects. Rather than treating the environment as an undifferentiated field of geometry, the descriptor builds a rich local representation of everything surrounding each Main Object, capturing not just where things are but what kinds of things they are.</p>
<p>The technical machinery behind the descriptor draws on the successful lineage of scan context methods. The original Scan Context, introduced in 2018, divides the space around a robot into a polar grid and encodes the maximum height of points in each cell, producing a compact two-dimensional matrix that can be compared rapidly against other scans. Scan Context++ and numerous successors refined this idea to handle rotation and lateral shifts in urban environments, and subsequent variants incorporated intensity information, deep learning, and other cues. OSSC extends this family by filling the grid not with geometric summaries alone but with weighted semantic labels, so that the pattern of object types in a neighborhood becomes a fingerprint of the place. Because objects such as buildings, poles, and vegetation tend to be arranged in stable configurations, two scans of the same area will encode similar semantic distributions even when the sensor viewpoints differ substantially.</p>
<p>A crucial design decision distinguishes OSSC from earlier attempts to inject semantics into place recognition. Some prior methods, such as the Semantic Scan Context approach, relied on a limited set of dominant or sparse semantic features, which made them fragile when the expected objects were missing, occluded, or poorly detected. The Singapore team instead chose to capture the semantic patterns and distributions of all objects around the Main Objects, not merely a handful of the most salient ones. This wholesale encoding of the semantic landscape gives the descriptor a resilience that sparse approaches lack. If one car moves between visits, or a pedestrian walks out of frame, the overall semantic composition of the scene remains recognizable, and the comparison between scans still yields a confident match.</p>
<p>The researchers also developed careful strategies for choosing which objects serve as Main Objects and for weighting different semantic labels according to their discriminative power. Not all object categories are equally useful for identifying a place. Buildings and poles persist and stay put, whereas cars and people come and go, so the system learns to emphasize the categories that reliably distinguish one location from another while downweighting the transient ones. These weighting strategies become especially important in challenging scenarios where the geometric structure of the environment is repetitive, such as corridors of similar-looking buildings or stretches of tree-lined road, and where semantic composition provides the only reliable signal of identity.</p>
<p>To test the approach, the team evaluated OSSC on two demanding public benchmarks. The first, SemanticKITTI, provides dense lidar point clouds with semantic annotations collected in structured urban and residential environments, and has become a standard proving ground for semantic perception research. The second, RELLIS-3D, offers point cloud data from unstructured, off-road terrain, a setting in which the tidy geometry of city streets gives way to irregular vegetation, uneven ground, and far less predictable scene composition. Performing well on both benchmarks is a meaningful achievement, because methods that thrive on the regular structure of urban scenes frequently collapse when confronted with the visual chaos of natural terrain.</p>
<p>The results showed high accuracy across a variety of scenarios, and, most significantly, the descriptor maintained its performance on scans that were spatially separated from one another. This is precisely the capability that matters most for practical deployment. A robot equipped with OSSC can recognize a previously visited region even from a moderately distant vantage point, which means it does not have to drive all the way back to the same spot before its mapping system can correct itself. The reduction in unnecessary traversals translates directly into operational savings: less energy consumed, less time wasted, and faster map convergence, benefits that compound over long autonomous missions in warehouses, campuses, agricultural fields, and city streets alike.</p>
<p>The significance of this work extends beyond any single algorithm. Loop closure detection sits at the heart of the growing mobile robotics sector, underpinning autonomous vehicles, delivery robots, inspection drones, and agricultural machinery, all of which must build and maintain accurate maps to function safely. As the industry scales, the robustness of place recognition in diverse environments, from structured cities to unstructured wild terrain, becomes a bottleneck for deployment. A descriptor that fuses geometry with semantics, anchored on stable objects and tolerant of viewpoint change, addresses the problem at its conceptual root: places are identified not just by their shapes but by the meaningful things they contain. The research also highlights the value of rich semantic segmentation, since the entire approach depends on accurately labeling points in the point cloud, and improvements in perception models will feed directly into better loop closure.</p>
<p>For the robotics community, OSSC offers a demonstration that the long-standing trade-off between the strictness of place recognition and the flexibility of robot behavior can be loosened. By encoding the full semantic distribution around carefully selected reference objects, and by weighting semantic labels to maximize discriminative power, the method achieves robustness in exactly the regimes, spatially apart scans, dynamic scenes, and unstructured terrain, where geometric descriptors stumble. The work, supported by the Robotics and Machine Intellection departments at A*STAR&#8217;s Institute for Infocomm Research and tested on openly available datasets, points toward a generation of robots that can navigate the world with a more human-like sense of place, one that recognizes a street corner not because the laser rangefinder sees identical geometry, but because the same distinctive assembly of buildings, poles, and vegetation stands sentinel there.</p>
<p><strong>Subject of Research:</strong> Semantic object-based loop closure detection for robot SLAM</p>
<p><strong>Article Title:</strong> Enhancing loop closure detection with object semantic scan context</p>
<p><strong>Article References:</strong> Kumarjiguda, D., Verma, S., Dutta, R., Ahmed, S. Z., &amp; Kun, Z. (2026). Enhancing loop closure detection with object semantic scan context. <em>Autonomous Robots, 50</em>(4), Article 39. <a href="https://doi.org/10.1007/s10514-026-10270-7" rel="noopener noreferrer">https://doi.org/10.1007/s10514-026-10270-7</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10514-026-10270-7" rel="noopener noreferrer">10.1007/s10514-026-10270-7</a></p>
<p><strong>Keywords:</strong> loop closure detection, SLAM, place recognition, lidar, point cloud, semantic segmentation, scan context, object semantics, mobile robotics, localization, SemanticKITTI, RELLIS-3D</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">204004</post-id>	</item>
	</channel>
</rss>
