Air pollution remains one of the deadliest and most stubborn environmental hazards of urban life, yet many of the cities that need air quality forecasts most urgently are precisely the ones that cannot afford dense networks of expensive monitoring stations. A new study published in Theoretical and Applied Climatology offers a striking demonstration that this gap can be bridged with software rather than hardware. Researchers Muhammed Ernur Akiner of Akdeniz University and Mehdi Ghasri of the University of Sistan and Baluchestan developed a hybrid machine learning framework that predicts the Air Quality Index in Kağıthane, a densely populated district on the European side of Istanbul, with an overall accuracy of approximately 98 percent, using roughly a decade of data and relying heavily on meteorological variables that are widely available almost everywhere.
The significance of the result lies less in the headline accuracy figure than in what it implies for the millions of people living in rapidly urbanizing regions where air quality monitoring infrastructure is thin. The Air Quality Index, or AQI, condenses the concentrations of several regulated pollutants into a single number that public health authorities use to issue warnings, close schools, and restrict traffic. Producing that number reliably normally requires continuous measurements of particulate matter, sulfur dioxide, carbon monoxide, and nitrogen dioxide from calibrated instruments. When a district has only sparse or intermittent monitoring, the index becomes unreliable exactly when residents need it most, and the health consequences of missed pollution episodes can be severe.
Akiner and Ghasri attacked the problem with a two-stage design that reflects a key insight about atmospheric data: pollution does not behave the same way under all weather conditions. In the first stage, they applied K-Means clustering, an unsupervised algorithm that groups observations by similarity without being told what the groups should be, to a combination of meteorological variables and pollutant concentrations. The algorithm partitioned the ten-year record into distinct atmospheric regimes, essentially discovering recurring weather-and-pollution patterns on its own. Each cluster corresponds to a characteristic set of conditions, such as stagnant, cold episodes in which pollutants accumulate near the surface, or windy, well-mixed periods in which concentrations disperse rapidly.
This clustering step matters because it gives the downstream model a structured picture of the atmosphere instead of forcing it to learn every relationship from raw, undifferentiated data. In the second stage, the researchers employed XGBoost, a gradient-boosted decision tree algorithm that has become a workhorse of tabular machine learning. XGBoost builds an ensemble of decision trees sequentially, with each new tree trained to correct the residual errors of the ensemble so far, and it includes regularization terms that discourage overfitting. The model was tasked with estimating AQI values from meteorological parameters together with the key pollutants PM10, PM2.5, SO2, CO, and NO2, conditioned on the regime structure identified by the clustering stage.
The crucial design choice is the framework’s reliance on widely available meteorological data to compensate for scarce pollutant measurements. Weather variables such as temperature, wind speed, humidity, and pressure are collected routinely by national meteorological services almost everywhere on Earth, and Türkiye’s General Directorate of Meteorology, together with the Ministry of Environment, Urbanization, and Climate Change, provided the database underpinning the study. Because these inputs are cheap and abundant, the framework reduces dependence on dense sensor infrastructure. In effect, the model learns how the local climate converts emissions into pollution episodes, and can then reconstruct or forecast air quality even where direct pollutant observations are limited.
The reported performance is remarkable for a real-world urban dataset. An overall accuracy of approximately 98 percent means the model reproduces the observed AQI with enough fidelity to support timely health advisories, giving authorities lead time to warn vulnerable populations, adjust traffic, or deploy inspection resources. Beyond raw accuracy, the authors used feature-importance analysis to identify which meteorological drivers dominate AQI variability in Kağıthane. This kind of interpretability is increasingly demanded of machine learning in environmental applications, because a black-box prediction is of limited use to a planner who needs to know whether ventilation, temperature inversions, or precipitation scavenging is controlling pollution on a given day.
Kağıthane itself is a fitting laboratory for the method. The district sits in a valley that channels both traffic and terrain-driven airflow, and previous research has documented elevated particulate levels in the Kağıthane Valley and the wider Istanbul metropolitan area. Istanbul combines intense vehicular emissions, a population exceeding fifteen million, complex topography, and a coastal climate whose winds and humidity strongly modulate pollutant dispersion. A model that performs well in such a demanding setting is more likely to transfer to other challenging environments than one tuned on a flat, well-monitored city.
The authors emphasize that the framework is modular and readily transferable across different geographic and temporal contexts. Both K-Means and XGBoost are standard, well-documented algorithms, and the input data are commonly available, which means a municipality does not need supercomputing resources or proprietary technology to adopt the approach. The pipeline can be retrained on local data, the clustering stage will rediscover the atmospheric regimes specific to a new location, and the boosted trees will re-weight the meteorological drivers accordingly. That scalability is the study’s central claim: a practical answer for regions facing environmental monitoring constraints, from fast-growing cities in the Global South to districts of wealthy cities where sensor coverage is uneven.
The work also arrives amid a broader shift in how air pollution is measured and modeled. Researchers worldwide are experimenting with low-cost sensor networks, satellite retrievals of aerosol optical depth, and deep learning architectures for spatiotemporal forecasting, and systematic reviews have catalogued both the promise and the pitfalls of machine learning in air pollution epidemiology. Against that backdrop, the Istanbul study makes a deliberately pragmatic argument: rather than waiting for perfect instrumentation, cities can extract far more value from the data they already possess by combining unsupervised structure discovery with powerful supervised regression. The approach complements, rather than replaces, physical dispersion models, which remain essential for source attribution and regulatory analysis.
For public health, the implications are immediate. Air pollution is linked to cardiovascular and respiratory disease, adverse birth outcomes, and growing evidence of effects on mental health and children’s cognitive development, and the burden falls disproportionately on neighborhoods with the least monitoring. A framework that turns routine weather data into near-accurate AQI forecasts can level that playing field, enabling evidence-based environmental policy in places that have historically flown blind. As urbanization accelerates and climate change alters the meteorological conditions that govern pollutant dispersion, tools like this hybrid model may become as standard a piece of civic infrastructure as the weather forecast itself, quietly translating the atmosphere’s behavior into actionable warnings for the people who live beneath it.
Subject of Research: Hybrid machine learning for air quality index forecasting in data-scarce urban environments
Article Title: Overcoming data shortages through hybrid machine learning in air quality forecasting: a case study of Kağıthane, Istanbul
Article References: Akiner, M. E., & Ghasri, M. (2026). Overcoming data shortages through hybrid machine learning in air quality forecasting: a case study of Kağıthane, Istanbul. Theoretical and Applied Climatology, 157(10), Article 609. https://doi.org/10.1007/s00704-026-06545-9
Image Credits: AI Generated
DOI: 10.1007/s00704-026-06545-9
Keywords: air quality, machine learning, XGBoost, K-Means clustering, Air Quality Index, Istanbul, Kağıthane, meteorology, urban pollution, data scarcity, public health, environmental forecasting
Cite Scienmag News
Russell Cooper. (October 8, 2026). Hybrid AI Reaches 98% Accuracy Forecasting Air Quality in Data-Poor Istanbul District. Scienmag. https://scienmag.com/hybrid-ai-reaches-98-accuracy-forecasting-air-quality-in-data-poor-istanbul-district/
Russell Cooper. "Hybrid AI Reaches 98% Accuracy Forecasting Air Quality in Data-Poor Istanbul District." Scienmag, 8 October 2026, https://scienmag.com/hybrid-ai-reaches-98-accuracy-forecasting-air-quality-in-data-poor-istanbul-district/. Accessed 8 October 2026.
Russell Cooper. "Hybrid AI Reaches 98% Accuracy Forecasting Air Quality in Data-Poor Istanbul District." Scienmag. October 8, 2026. https://scienmag.com/hybrid-ai-reaches-98-accuracy-forecasting-air-quality-in-data-poor-istanbul-district/

