The rapid global spread of multimodal large language models has raised a deceptively simple question: when these systems look at a photograph of a city street, do they see it the same way regardless of where in the world that street happens to be? A new study published in npj Urban Sustainability suggests that the answer is a resounding no. Researchers led by H. Kim, X. Li and M. Quintana have assembled one of the most geographically comprehensive test beds to date for probing the visual and perceptual capacities of multimodal artificial intelligence, drawing on imagery from more than 200 cities across every inhabited continent. Their findings reveal systematic geographic and perceptual biases that mirror, and in some cases amplify, longstanding inequities in the digital data on which these models are trained.
The research team’s central concern was not merely whether multimodal large language models, or MLLMs, can identify a landmark or read a street sign. Rather, they set out to determine whether the perceptual judgments these systems make about urban environments — assessments of safety, walkability, beauty, wealth and order — are applied consistently across cities in different countries, income brackets and cultural contexts. Such perceptual scores increasingly matter in the real world. Urban planners, technology companies and civic institutions have begun experimenting with AI-assisted tools that evaluate neighborhoods, rank streetscapes and inform investment decisions. If the underlying models carry hidden geographic biases, those biases could quietly shape policy and perception at scale.
To investigate, the authors curated a global dataset spanning more than 200 cities, deliberately reaching well beyond the heavily documented urban cores of North America and Western Europe. The dataset incorporates street-level and urban imagery paired with perception labels, allowing the researchers to compare how state-of-the-art multimodal models rate visual attributes of city environments against established human benchmarks derived from perception studies. The design enabled a kind of controlled stress test: hold the perceptual question constant, vary the geography, and observe where model performance holds and where it fractures.
The results were striking. The team documented consistent disparities in how accurately the models perceived urban attributes depending on where a city is located. In many regions of the world — particularly in parts of Africa, Latin America and South and Southeast Asia — the models’ perceptual judgments diverged more sharply from human ground truth than they did for cities in North America, Europe and East Asia’s wealthiest urban centers. The pattern constitutes what the researchers describe as a geographic bias: a systematic degradation of perceptual reliability that is correlated with location rather than with any inherent property of the places themselves.
Beneath that headline pattern, the study unpacked a more granular form of distortion the authors characterize as perceptual bias. Even when a model could roughly identify what it was looking at, its judgments along specific perceptual dimensions were skewed. Some attributes were systematically overestimated in certain regions and underestimated in others. A street in one city might be rated as cleaner, safer or more affluent than an objectively comparable street elsewhere, purely because of the visual and cultural priors embedded in the model’s training data. These are not random errors; they are directional errors, and directional errors at scale become a form of algorithmic stereotyping.
The technical roots of the problem trace back to how multimodal models are built. Contemporary MLLMs combine a vision encoder, typically a transformer-based architecture trained on enormous collections of web-scraped images, with a language model that interprets and verbalizes the visual representations. Both components inherit the statistical fingerprints of their training corpora. Web imagery is not uniformly distributed across the globe: images from wealthier, English-speaking and heavily digitized countries are dramatically overrepresented, while many cities in the Global South appear only sparsely. The result is a vision system that has effectively “seen” far more of some parts of the world than others, and a language layer that has absorbed uneven cultural narratives about what different places look like and mean.
The new study’s contribution is to quantify these effects with global scope and methodological rigor. By correlating model error with indicators such as geographic region, economic development and data availability, the authors were able to dissociate several candidate explanations. They found that the biases are not simply a function of image quality or resolution, nor can they be dismissed as artifacts of labeling noise. Instead, the evidence points to deep-seated representational gaps: for underrepresented cities, the models appear to substitute learned priors — generalized impressions of what a “developing” or “non-Western” urban environment supposedly looks like — for accurate perception of the actual scene.
The implications extend well beyond the laboratory. Multimodal AI is increasingly positioned as an engine for automated assessment of the built world, from satellite- and street-level analytics platforms to smart-city dashboards that promise real-time monitoring of infrastructure and quality of life. If a model systematically underestimates the safety or aesthetic quality of streets in Lagos, Jakarta or Lima relative to comparable streets in Amsterdam or Toronto, any downstream application built on those scores inherits the distortion. Real-estate analytics, insurance pricing, tourism ranking and urban planning support tools could all, without anyone intending it, channel attention and resources toward places the model already favors, reinforcing existing cycles of advantage and neglect.
The study also speaks to a broader tension in contemporary AI governance. Efforts to audit large models for bias have concentrated heavily on demographic attributes such as race and gender, often measured within Western contexts. Geographic bias is more diffuse and harder to operationalize, because the relevant variable — where a place is — entangles everything from lighting conditions and vegetation to signage, architecture and cultural conventions of visual order. The global dataset assembled by Kim, Li, Quintana and colleagues offers a template for making this variable tractable, and demonstrates that perceptual fairness is a measurable property, not merely an intuition.
What can be done? The authors point toward several complementary remedies. Training and fine-tuning data can be rebalanced to include far denser coverage of underrepresented regions, ideally collected with local participation and local contextual grounding. Evaluation practices should become geographically stratified, with performance reported not only as a global average but as a breakdown across regions and city types, so that a model cannot pass an audit on the strength of high scores in familiar territory alone. And downstream users should treat perceptual outputs for unfamiliar geographies with explicit caution, since the models’ confidence in these regions is not matched by their accuracy.
There is also a subtler lesson embedded in the findings. Human perception of urban environments is itself culturally situated — people from different backgrounds rate the same street differently — and any perception benchmark encodes some cultural standpoint. The challenge for AI developers is not to eliminate standpoint altogether, an impossible goal, but to make the standpoints plural and the uncertainty visible. A model that says “this street feels unsafe, but I am far less certain for cities in this region” is a fundamentally more trustworthy instrument than one that projects confident, homogeneous judgments across an unevenly understood planet.
As multimodal AI systems continue their rapid diffusion into everyday tools and professional workflows, the study stands as a timely warning wrapped in a rigorous empirical package. More than 200 cities’ worth of evidence shows that the artificial eyes now increasingly used to interpret the urban world do not see all of it equally. Making them see more fairly is not a niche technical concern; it is a prerequisite for any serious claim that AI can help build smarter, more equitable cities.
Cite Scienmag News
Courtney Benton. (September 11, 2026). Multimodal LLMs show geographic and perceptual bias across 200 cities. Scienmag. https://scienmag.com/multimodal-llms-show-geographic-and-perceptual-bias-across-200-cities/
Courtney Benton. "Multimodal LLMs show geographic and perceptual bias across 200 cities." Scienmag, 11 September 2026, https://scienmag.com/multimodal-llms-show-geographic-and-perceptual-bias-across-200-cities/. Accessed 11 September 2026.
Courtney Benton. "Multimodal LLMs show geographic and perceptual bias across 200 cities." Scienmag. September 11, 2026. https://scienmag.com/multimodal-llms-show-geographic-and-perceptual-bias-across-200-cities/

