The global garment industry has long struggled with a stubborn problem: how do you know whether the factories stitching your clothes are actually safe, fair, and environmentally responsible, rather than simply claiming to be? A new study published in Discover Sustainability by Kazi Md. Tanvir Anzum of Khulna University of Engineering & Technology in Bangladesh offers a computational answer. The research builds a reinforcement learning system that learns, through trial and error, how to score suppliers on sustainability, adjust sourcing shares, and target audits in a stylized model of a Bangladesh ready-made garment sourcing portfolio. In simulated head-to-head competitions, the learning agent outperformed a suite of conventional governance policies, achieving higher true compliance with fewer audits and nearly half as many severe incidents.
The technical heart of the study is a safety-gated form of tabular Q-learning, a classic reinforcement learning algorithm in which an agent maintains a table of expected long-term rewards for each combination of state and action. What makes this implementation distinctive is the set of hard constraints wrapped around the learning process. The simulation represents safety, labour, and environmental performance as separate dimensions for each supplier, and the agent is never allowed to collect rewards when verified safety or labour performance falls below a threshold of 0.60. A minimum annual audit is imposed on every supplier regardless of what the agent would prefer. These gates are evaluated before any policy action is executed, and crucially, an unverified self-disclosure from a supplier cannot lift the gate. Only the last successfully verified safety and labour values count.
This design choice reflects a deep insight about the problem structure. Supplier sustainability is a partially observable process: the agent cannot see the latent true sustainability of a factory, only noisy proxies such as the most recent verified score, current disclosure, performance trend, audit recency, and any verified critical flags. The mathematical framework carefully separates these latent variables from observed ones, and the observable state is constructed as an approximate discrete representation of the underlying reality. By preventing disclosure alone from making a supplier reward-eligible, the model encodes a principle that governance experts have long advocated: self-reported claims are not evidence.
To test the approach, the researcher trained three independent Q-learning agents and evaluated them in 75 paired out-of-sample replications against six competing policies: a contextual bandit, a risk-based adaptive heuristic, a static threshold rule, a combined scheduled-audit-and-threshold policy, scheduled auditing alone, and random action selection. The paired replication design matters for statistical rigor. Rather than treating each of the 1,440 supplier-month transitions as an independent observation, the study aggregates incidents and actions at the replication level before running paired tests, avoiding the artificial precision that inflates significance in many simulation studies.
The results are striking. The Q-learning agents achieved a mean true compliance of 0.862, compared with 0.820 under the static threshold policy that represents a common industry baseline. The share of suppliers scoring at or above 0.80 rose from 58.2 percent under the baseline to 81.3 percent under the learning agent. Perhaps most dramatically, the mean number of severe incidents fell from 3.29 to 1.71, a reduction of nearly half. All of this was accomplished with 167 audits, against 362 for the combined scheduled-audit-and-threshold policy, a 53.8 percent reduction in audit burden. The trade-off was a modest one: mean procurement cost ran about 1.1 percent higher than under the static threshold.
The comparison with the contextual bandit is particularly instructive for machine learning practitioners. A contextual bandit is a simpler learning algorithm that chooses actions based on immediate observed context without modeling delayed consequences. In these experiments, the bandit produced lower compliance, at 0.742, and more severe incidents, at 2.96, than the full Q-learning agent. The gap supports a central argument of the paper: supplier governance is a delayed-action problem. The consequences of shifting sourcing share away from a factory, or of skipping an audit, unfold over months, and an algorithm that only optimizes immediate reward cannot capture those dynamics. Q-learning, which propagates future value backward through its state-action table, can.
The risk-based adaptive heuristic, a hand-crafted policy that allocates audits according to assessed risk, fared better than most alternatives. It produced statistically similar mean compliance to the Q-learning agent, but it required 55 additional audits and yielded a narrower share of compliant suppliers. In other words, the learned policy matched the performance of a well-designed human heuristic while using fewer resources and spreading compliance gains more broadly across the supplier base. For procurement organizations weighing the cost of sophisticated analytics against the cost of audits, that efficiency margin could be decisive.
The study also reports an unusually thorough set of robustness checks. Paired statistical tests are accompanied by effect sizes, convergence diagnostics confirm that the Q-tables stabilized during training, and results are broken down by supplier stratum, distinguishing direct exporters, subcontractors, and small or informal suppliers, the three tiers that characterize real garment sourcing portfolios. A credit-assignment ablation examines how the reinforcement signal drives learning, and time-state testing probes whether the policy generalizes across simulation periods. Multi-parameter robustness analyses vary key model settings to confirm the findings are not artifacts of a single configuration. The sourcing-share adjustment itself is formulated as a convex quadratic projection with linear equality and bound constraints, which guarantees a unique allocation that is independent of the order in which suppliers are processed, a subtle but important fairness property.
Several implementation details reveal how carefully the model guards against gaming and error. The annual audit floor is computed from elapsed months since the last successful verification, not since an attempted audit, so a failed inspection does not reset the clock and buy a supplier extra time. Audit recency and the supplier-specific reinforcement signal are stored as separate fields in the implementation, preventing conceptual conflation. The critical gate relies exclusively on verified values, closing the loophole through which a struggling factory might talk its way back into good standing. These choices make the simulation a meaningful test of adaptive governance rather than a naive optimization exercise.
The author is candid about the limits of the work. This is simulation evidence, built from a stochastic model calibrated with published aggregate statistics, not factory-level validation, and no real procurement manager has yet handed sourcing decisions to an algorithm. The study involved no human participants and received no external funding. Still, the implications are considerable. Ready-made garment supply chains face mounting legal, commercial, and stakeholder pressure to verify sustainability rather than trust supplier claims, and audit budgets are finite. A framework that lifts compliance, halves severe incidents, and cuts audit volume by more than half, even in a stylized setting, suggests that reinforcement learning could become a practical instrument of supply chain governance. The open-access paper, published on 5 October 2026 with DOI 10.1007/s43621-026-04745-x, provides a reproducible foundation for the field trials that must come next, and it signals a future in which the clothes we wear are policed not just by clipboard-wielding inspectors, but by algorithms that learn where the real risks hide.
Subject of Research: Reinforcement learning for adaptive supplier sustainability scoring and audit targeting in ready-made garment supply chains
Article Title: Reinforcement learning for adaptive supplier sustainability scoring and sourcing share adjustment in ready made garment supply chains
Article References: Anzum, K. M. T. (2026). Reinforcement learning for adaptive supplier sustainability scoring and sourcing share adjustment in ready made garment supply chains. Discover Sustainability. https://doi.org/10.1007/s43621-026-04745-x
Image Credits: AI Generated
DOI: 10.1007/s43621-026-04745-x
Keywords: reinforcement learning, Q-learning, supply chain management, sustainability, garment industry, supplier governance, audit targeting, sourcing share adjustment, Bangladesh, machine learning, simulation, compliance
Cite Scienmag News
Violet Maxwell. (October 5, 2026). AI Learns to Police Garment Suppliers, Cutting Severe Incidents Nearly in Half in Simulation. Scienmag. https://scienmag.com/ai-learns-to-police-garment-suppliers-cutting-severe-incidents-nearly-in-half-in-simulation/
Violet Maxwell. "AI Learns to Police Garment Suppliers, Cutting Severe Incidents Nearly in Half in Simulation." Scienmag, 5 October 2026, https://scienmag.com/ai-learns-to-police-garment-suppliers-cutting-severe-incidents-nearly-in-half-in-simulation/. Accessed 5 October 2026.
Violet Maxwell. "AI Learns to Police Garment Suppliers, Cutting Severe Incidents Nearly in Half in Simulation." Scienmag. October 5, 2026. https://scienmag.com/ai-learns-to-police-garment-suppliers-cutting-severe-incidents-nearly-in-half-in-simulation/

