Every river tells a story through the stones it carries. The size of the pebbles scattered along a stream bed reveals how fast the water flows, where the sediment came from, and how the landscape is evolving over thousands of years. For decades, however, reading that story has required geoscientists to crouch on riverbanks with metric rulers, painstakingly measuring pebble after pebble in what is known as the Wolman pebble count. Now, a team at the University of Potsdam has unveiled a tool that could render that tradition obsolete: an artificial intelligence workflow called OrthoSAM that automatically identifies and outlines every individual pebble in enormous, high-resolution aerial images.
The research, published in the journal Earth Surface Dynamics, tackles a surprisingly stubborn problem in geomorphology. Grain-size analysis is fundamental to understanding natural hazards, hydrologic conditions, and ecosystems, but traditional field methods are costly, labor-intensive, and slow. A trained observer recording sizes along a stream can typically measure only a few hundred pebbles, yet robust statistical analysis of grain-size distributions often requires far more measurements than that. Photo-based alternatives, from image texture analysis to machine-learning classifiers, have tried to close the gap, but each comes with compromises that limit their usefulness in the messy, shadow-strewn environment of a real mountain river.
OrthoSAM builds on the Segment Anything Model, or SAM, a foundation model released by Meta AI researchers in 2023 that was trained on more than one billion masks across eleven million images. SAM’s remarkable capability is zero-shot inference: it can segment objects in entirely new image domains, from cell microscopy to pebble photographs, without any additional training. Its architecture consists of three decoupled components, an image encoder, a prompt encoder, and a mask decoder, which together allow the model to generate precise segmentation masks from simple input prompts such as points or polygons. For scientific imagery, this flexibility is a genuine advantage over custom-trained neural networks, which tend to perform well only within the narrow conditions they were trained on.
But SAM was never designed for what river scientists need. The model rescales every input image to 1024 by 1024 pixels to fit its transformer architecture, which means a 24-megapixel photograph is effectively compressed down to roughly 0.7 megapixels. For a sprawling orthomosaic containing thousands of densely packed pebbles, that compression is catastrophic: small stones shrink into invisibility and larger objects lose mask quality. The standard automated scheme also applies a grid of 32 by 32 equally spaced input points, optimized for images containing a limited number of objects, not the hundreds to thousands of grains found on a single riverbed image. Hardware constraints compound the difficulty, since each input point generates three candidate masks that must all be held in GPU memory.
The Potsdam team, led by Vito Chan together with Aljoscha Rheinwalt and Bodo Bookhagen, engineered a workflow of three interlocking components to overcome these limits. First, a tiling scheme divides large orthomosaics into 1024 by 1024 pixel patches with a definable overlap, so that SAM processes each patch at full effective resolution. Masks that touch a tile border are discarded to avoid artificially over-segmented pebbles along the seams, and a filtering box ensures only one mask is kept per object. Second, an improved input point generator determines how densely the prompt grid must be spaced so that every object, even the smallest pebble of interest, receives at least one input point, while a centroid-based refinement step reduces duplicate and overlapping masks. Third, a multi-scale resampling scheme runs the segmentation at several resolutions and merges the results, allowing boulders too large to fit within a single tile to be captured in a coarser pass and stitched back into the final labeled map.
Validating such a system posed its own challenge, because reliable ground-truth datasets with thousands of individually delineated pebbles simply do not exist. The researchers therefore built a synthetic pebble generator, producing images of 10,000 by 10,000 pixels filled with up to 5,000 non-overlapping solid circles of random sizes, rendered in black and white, in color, with Gaussian noise, and with simulated shadows cast by hemispherical domes under a directional light source. Because the true positions and sizes of every circle are known, the team could rigorously quantify detection quality, mask accuracy, and the fidelity of the resulting size distributions.
The results were striking. On 1,872 synthetic black-and-white pebbles, OrthoSAM achieved a precision of 1.0, a recall of 0.87, and a mean Intersection over Union, the standard measure of mask accuracy, of 0.98. On 27,528 colored pebbles with shadows, the workflow reached a precision of 0.98, a recall of 0.94, and a mean IoU of 0.91. A two-sample Kolmogorov-Smirnov test confirmed that the predicted grain-size distributions were statistically indistinguishable from the ground truth, with p-values exceeding 0.25 in every synthetic image. The team also identified a crucial detection limit: pebbles with a diameter below 30 pixels are not reliably detected, a threshold that directly informs how field crews should plan camera distance and image resolution.
To prove the concept on real terrain, the researchers applied OrthoSAM to three orthomosaics from the Ravi River in the western Himalaya, images assembled from hundreds of photographs taken with a Sony camera and processed to a spatial resolution of 0.2 millimeters per pixel. The workflow delineated 6,087 pebbles across the three scenes, with manual verification of every predicted mask yielding a precision of 0.93 and a recall of 0.94. For each pebble, the software computes a rich table of measurements, including projected area, the lengths of the a- and b-axes, perimeter, color statistics, and a normalized isoperimetric ratio that serves as a proxy for roundness, all exported in a format ready for further geomorphic analysis.
The study is candid about the method’s remaining weaknesses. Shadows proved to be a subtle adversary: pebbles completely covered by shadow were segmented as accurately as those in uniform light, but partially shadowed pebbles, where stark contrast divides a single object into lit and dark halves, were responsible for most of the observed drop in precision and mask quality. Strong image noise also degrades accuracy, prompting the authors to recommend keeping camera ISO settings moderate during data collection. Because SAM performs only instance segmentation without any classification component, non-pebble objects such as marker panels, wooden branches, and structure-from-motion artifacts can also be delineated and must be filtered out before statistics are computed. The team calls for a community-based reference dataset for granular material segmentation, analogous to the benchmark datasets that transformed lidar research.
The implications reach well beyond one Himalayan river. Grain-size distributions govern how sediment moves through drainage basins, how flood hazards develop, and how aquatic ecosystems function, and the ability to sample more than ten thousand objects from a single virtual outcrop promises a new scale of observation for Earth scientists. OrthoSAM’s Python-based pipeline is openly available on GitHub and archived on Zenodo, and the Ravi River orthomosaics are published under an open license, meaning any research group with a capable GPU can begin counting stones. A 10,000 by 10,000 pixel synthetic image currently takes about four hours to process on a 16-gigabyte NVIDIA Quadro RTX 5000, so wider adoption will depend on computing resources, but the direction is clear. The humble pebble count, a field technique essentially unchanged since 1954, is being reborn as an automated, reproducible, and scalable digital science.
Subject of Research: Automated delineation of river pebbles and grain-size analysis from high-resolution orthophotos using an extended Segment Anything Model
Article Title: OrthoSAM: multi-scale extension of the Segment Anything Model for river pebble delineation from large orthophotos
Article References: Chan, V., Rheinwalt, A., & Bookhagen, B. (2026). OrthoSAM: multi-scale extension of the Segment Anything Model for river pebble delineation from large orthophotos. Earth Surface Dynamics, 14(3), 391-416. https://doi.org/10.5194/esurf-14-391-2026
Image Credits: AI Generated
DOI: 10.5194/esurf-14-391-2026
Keywords: OrthoSAM, Segment Anything Model, grain-size analysis, geomorphology, river pebbles, orthomosaics, deep learning, image segmentation, remote sensing, Himalaya, Ravi River, sediment transport
Cite Scienmag News
Violet Maxwell. (October 10, 2026). AI Learns to Count Every Pebble in a River, Transforming Sediment Science. Scienmag. https://scienmag.com/ai-learns-to-count-every-pebble-in-a-river-transforming-sediment-science/
Violet Maxwell. "AI Learns to Count Every Pebble in a River, Transforming Sediment Science." Scienmag, 10 October 2026, https://scienmag.com/ai-learns-to-count-every-pebble-in-a-river-transforming-sediment-science/. Accessed 10 October 2026.
Violet Maxwell. "AI Learns to Count Every Pebble in a River, Transforming Sediment Science." Scienmag. October 10, 2026. https://scienmag.com/ai-learns-to-count-every-pebble-in-a-river-transforming-sediment-science/

