The explosive growth of smartphone-based location data has transformed how scientists study human mobility, how cities plan transportation networks, and how public health officials track the movements of populations during crises. From estimating commute patterns to modeling epidemic spread, datasets harvested from mobile applications now underpin countless research findings and policy decisions. Yet a new multiscale analysis of the United States, published in Scientific Reports, delivers a sobering reminder that these data are far from a neutral window onto society. The study systematically characterizes the spatial and sociodemographic sampling biases embedded in mobile location data, showing that the people who appear in these datasets are not a representative sample of the population at any of the geographic scales examined.
The research team set out to answer a deceptively simple question: who is actually represented in mobile location data, and how does that representation vary across space and demographic groups? To do so, the investigators compared the demographic and geographic composition of large-scale mobile location panels against authoritative benchmark populations drawn from the United States Census. By aligning the observed device counts with census figures across multiple spatial resolutions, from the national level down to counties and finer census geographies, the researchers were able to quantify where and for whom mobile location data over-represents or under-represents the underlying population. This benchmarking approach transforms a vague concern about bias into a measurable, mappable quantity.
The multiscale design is central to the study’s contribution. Many previous audits of mobility data focused on a single geographic level, typically the nation as a whole, where biases can average out and appear modest. But the new analysis demonstrates that national aggregates conceal dramatic local variation. A dataset that looks reasonably balanced at the country level may severely distort the picture in individual counties, particularly in rural areas, on tribal lands, and in communities with distinctive age or income profiles. Because many downstream applications, from retail site selection to emergency response planning, operate at precisely these local scales, the discrepancy between national fairness and local distortion carries real consequences.
Technically, the analysis relied on comparing device-level residence estimates with census counts of population by age, sex, race and ethnicity, income, and urbanicity. The researchers computed representation ratios, expressing the share of mobile devices attributed to a group or area relative to that group’s share of the benchmark population. Values above one indicate over-representation and values below one indicate under-representation. By computing these ratios repeatedly across scales and demographic dimensions, the team built a comprehensive portrait of sampling bias that could be decomposed into spatial components, sociodemographic components, and their interactions. This decomposition matters because spatial and demographic biases are intertwined: certain places are home to certain populations, so a geographic gap is often simultaneously a demographic one.
The findings confirm and sharpen long-standing suspicions about who is missing from mobile location data. Older adults, particularly those in the oldest age brackets, appear substantially less often than their share of the population would predict, reflecting lower smartphone adoption, different app ecosystems, and greater privacy sensitivity. Racial and ethnic minorities show uneven representation that varies by region, with some communities under-represented and others appearing at rates closer to parity. Income gradients emerge as well, with lower-income households less likely to contribute devices to commercial location panels, a pattern consistent with disparities in smartphone access, data plan affordability, and the types of applications that collect and monetize location information.
Geographically, the study finds that rural counties are systematically under-represented relative to urban cores, where dense populations, heavy app usage, and the economics of the mobile advertising ecosystem concentrate data collection. The magnitude of the rural gap is not uniform: it widens in regions with older populations and lower broadband penetration, and it narrows in peri-urban counties that blend rural character with metropolitan labor markets. The authors also document variation across states and metropolitan areas, suggesting that regional culture, state-level demographics, and the spatial distribution of technology industries all leave fingerprints on the composition of location panels. For analysts, the practical implication is that bias correction cannot rely on a single national adjustment factor.
Why do these biases arise in the first place? The study situates the problem in the mechanics of how commercial location data are produced. Devices enter these datasets only when users grant location permissions to applications that partner with data aggregators, and only when the resulting location pings pass quality filters designed to remove noise and protect privacy. Each step in this pipeline, from app adoption to permission granting to data brokerage, filters the population in ways that correlate with age, income, race, and place. The result is a dataset that is enormous, in the tens of millions of devices, yet still systematically skewed. Scale, the authors emphasize, is not the same as representativeness; a bigger biased sample is still a biased sample.
The implications ripple across the sciences that depend on mobility data. Epidemiological models that used location data to estimate contact patterns during the COVID-19 pandemic, for example, may have under-counted the movements of older and lower-income individuals, precisely the groups whose vulnerability and exposure dynamics differed most from the smartphone-owning mainstream. Transportation planners who calibrate travel demand models with mobile data risk designing infrastructure around the behavior of a demographically narrow public. Studies of segregation, access to services, and environmental exposure that infer activity spaces from location pings may mischaracterize the daily geographies of communities that are thinly sampled. The new bias maps offer these fields a diagnostic tool: before using location data for a given region or population, researchers can now check how well that region and population are actually represented.
The authors also point toward remedies. Weighting and reweighting schemes, in which devices are assigned statistical weights so that the weighted sample matches census benchmarks, can reduce aggregate bias, though the study cautions that such corrections struggle when entire groups are nearly absent from the data. Hybrid approaches that blend mobile data with traditional surveys, administrative records, or census microdata can anchor estimates to known population structure. Perhaps most importantly, the researchers advocate routine bias auditing as a standard step in any analysis using commercial location data, analogous to the data-quality checks that are customary in survey research. Transparency about representation, they argue, should become a baseline expectation for both data vendors and the scientists who consume their products.
As mobile location data continue to expand into nearly every corner of computational social science, this multiscale audit of the United States serves as both a warning and a roadmap. The warning is that the digital traces left by smartphones reflect the inequalities of digital access and technology adoption, and that these inequalities are written into the data at every geographic scale. The roadmap is a practical methodology for measuring those biases, mapping them across the country, and correcting for them where possible. In an era when decisions about public health, transportation, and resource allocation increasingly rest on algorithmic interpretations of human movement, knowing exactly who is missing from the data is the first step toward ensuring that data-driven decisions serve everyone, not just those who carry a trackable device.
Subject of Research: Spatial and sociodemographic sampling biases in mobile location data across the United States
Article Title: Characterizing the spatial and sociodemographic sampling biases found in mobile location data through a multiscale analysis of the United States
Article References: Embury, J., Dodge, S., & Nara, A. (2026). Characterizing the spatial and sociodemographic sampling biases found in mobile location data through a multiscale analysis of the United States. Scientific Reports. https://doi.org/10.1038/s41598-026-70982-9
Image Credits: AI Generated
DOI: 10.1038/s41598-026-70982-9
Keywords: mobile location data, sampling bias, human mobility, demographics, census benchmarking, spatial analysis, rural under-representation, digital divide, data representativeness, computational social science, Scientific Reports, United States
Cite Scienmag News
Denise Maddox. (September 22, 2026). Mobile Location Data Carries Hidden Biases, Multiscale US Analysis Reveals. Scienmag. https://scienmag.com/mobile-location-data-carries-hidden-biases-multiscale-us-analysis-reveals/
Denise Maddox. "Mobile Location Data Carries Hidden Biases, Multiscale US Analysis Reveals." Scienmag, 22 September 2026, https://scienmag.com/mobile-location-data-carries-hidden-biases-multiscale-us-analysis-reveals/. Accessed 22 September 2026.
Denise Maddox. "Mobile Location Data Carries Hidden Biases, Multiscale US Analysis Reveals." Scienmag. September 22, 2026. https://scienmag.com/mobile-location-data-carries-hidden-biases-multiscale-us-analysis-reveals/








