Thursday, October 8, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

AI Agents Match Human Averages but Diverge in How They Respond to Rewards and Penalties

October 8, 2026
in Technology and Engineering
Denise Maddox
By Denise Maddox Scienmag Editorial Profile - Mechanical Engineering
Reading Time: 5 mins read
0
AI Agents Match Human Averages but Diverge in How They Respond to Rewards and Penalties

AI Agents Match Human Averages but Diverge in How They Respond to Rewards and Penalties

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Large language models are increasingly being drafted as stand-ins for human subjects, populating virtual societies and simulating policy responses that would be too slow, costly, or impractical to test on real people. But a new study published in iScience delivers a cautionary finding for this fast-growing practice: when humans and AI agents face the same sequential reward-and-penalty incentives, they can land on nearly identical average outcomes while getting there in strikingly different ways. The research, led by Sihan Tao, Linghao Wang, Zheng Zhu, and Der-Horng Lee, suggests that matching aggregate behavior is not enough to certify an AI as a faithful proxy for human decision-making.

The team built their experiment around a real-world policy problem: greener last-mile delivery in e-commerce logistics. Online retail has made urban delivery a significant source of carbon emissions, and platforms such as Amazon have experimented with one-shot incentives like vouchers to nudge shoppers toward slower, lower-emission shipping. The researchers instead tested a credit-charge-cum-reward mechanism, in which participants accumulate carbon credits across repeated decisions and face a periodic settlement. Choosing standard, faster delivery deducts credits; choosing slow delivery earns them. Positive balances convert into cash rewards, negative balances trigger penalties, and the account resets after each 20-day round. The design forces an intertemporal trade-off between immediate speed and long-term credit, much like the cumulative incentive structures found in real environmental policy.

To anchor the comparison, the researchers first ran a baseline scenario with no incentives at all. Participants, whether human or artificial, simply chose between standard delivery, which took three days but could stretch under simulated congestion, and slow delivery, fixed at seven days. Each participant was assigned a value of time, a concept borrowed from transportation engineering that quantifies how much money a person will pay to save a unit of time, set at high, medium, or low levels. Theory predicted that everyone should pick the fast option, and that is precisely what happened. Humans chose standard delivery 98.9 percent of the time, DeepSeek-v.3.2 agents 99.7 percent, and Qwen3-Max agents 98.1 percent, with no statistically detectable differences among the three groups after correction for multiple comparisons.

Then came the carbon credit mechanism, and with it the study’s central test. The theoretical equilibrium under the new rules called for an overall standard delivery proportion of 51.39 percent, stratified sharply by value of time: 100 percent for high-value participants, 54.17 percent for the medium group, and zero for the low group. In aggregate, all three participant types responded dramatically and almost identically. Humans cut their standard delivery proportion from 0.989 to 0.463, DeepSeek-v.3.2 from 0.997 to 0.513, and Qwen3-Max from 0.981 to 0.450. The adjusted effect sizes ranged from roughly 48 to 53 percentage points, none of the pairwise differences between participant types was statistically significant, and all pooled estimates were compatible with the theoretical benchmark. On the surface, the AI agents looked like perfect miniature humans.

The illusion dissolved once the researchers stratified the results by value of time. Humans retained a stubborn residual of standard delivery choices even in the low-value group, where theory predicted none, holding at 9.2 percent. DeepSeek-v.3.2 produced a much smaller residual of 2.3 percent, while every Qwen3-Max system converged to exactly zero. At the high end, DeepSeek-v.3.2 hugged the theoretical ceiling most closely, whereas humans showed the largest deviation. Dispersion patterns diverged too: human systems remained heterogeneous across all strata, DeepSeek-v.3.2 systems clustered tightly around the interior benchmark in the medium group, and Qwen3-Max systems collapsed to the boundary in the low group. The same average masked three different behavioral signatures.

Longitudinal analysis over 120 virtual decision days exposed further fault lines. Humans adjusted gradually and sometimes incompletely, with their best response rate, a measure of how often a choice minimized one-step generalized cost given available information, improving significantly only in the high-value group. DeepSeek-v.3.2 stayed near its theoretical benchmarks from the outset and maintained the highest overall best response rate at 0.960, compared with 0.666 for humans and 0.679 for Qwen3-Max. Qwen3-Max displayed delayed adjustment, with its standard delivery probability in the medium-value group climbing from 0.018 in the first round to 0.493 in the sixth, a convergence that ultimately exceeded that of the other participants but followed a qualitatively different path. Because the language models generate responses without parameter updating, these trajectories reflect in-context conditioning and stochastic generation rather than the experience-based learning that shapes human behavior.

Conditional choice persistence revealed perhaps the strangest divergence. In the high-value group, whether an agent had chosen standard delivery on the previous purchase day strongly predicted its current choice for both AI models, with adjusted probability increases of 0.218 for DeepSeek-v.3.2 and 0.405 for Qwen3-Max, while humans showed essentially no such carryover at 0.010. In the medium-value group the pattern inverted, with DeepSeek-v.3.2 showing a strong negative persistence of −0.471 while humans and Qwen3-Max showed none. In other words, the AI agents’ decisions were structured by their recent histories in ways that no aggregate statistic would ever reveal.

An exploratory clustering analysis added a final layer of evidence. Using a Gaussian mixture model over allocation frequencies on one-item and two-item purchase days, the researchers identified four strategy modes: standard dominant, slow dominant, mixed allocation, and variable allocation. Human portraits spread across multiple modes within every value-of-time group, reflecting genuine individual diversity. DeepSeek-v.3.2 followed a clean ordered transition from standard dominant at high value of time, to mixed allocation at medium, to slow dominant at low. Qwen3-Max instead concentrated overwhelmingly in the variable allocation mode at high and medium value of time before shifting entirely to slow dominant at low. Notably, when decision-state features such as credit balances were added to the clustering, most of Qwen3-Max’s variable-mode assignments dissolved, indicating that frequency-based portraits can compress distinct state-conditioned behaviors into a single misleading category.

The authors are careful about scope. Only two models were tested, no prompt or parameter sensitivity analysis was performed, and the human sample consisted primarily of Zhejiang University students, so the findings speak to the specific configurations examined rather than to language models in general. The experimental environment also simplifies real logistics, omitting promotional discounts, product heterogeneity, and seasonal demand spikes. Still, the pattern is consistent with a growing body of work showing that surface similarity between human and machine outputs can conceal deep differences in the strategies and cognitive structures that produce them.

The implications reach well beyond delivery apps. As generative agents become a scalable substitute for human experiments in policy simulation, this study provides a concrete methodological warning: validating an AI proxy by its aggregate response alone can certify a simulator that gets the subgroup patterns, the temporal dynamics, and the underlying choice architecture wrong. For policymakers designing cumulative incentive mechanisms, whether carbon credits, congestion charges, or energy-use rewards, the practical effectiveness of a policy depends not just on average uptake but on how individuals perceive, learn, and adapt over time. The study suggests LLM simulations can reliably reproduce the direction of an aggregate policy response under well-defined rules, but researchers and regulators should demand disaggregate evidence, stratified trajectories, and strategy-level portraits before trusting an artificial crowd to stand in for a real one.

Subject of Research: Comparing human and large language model decision-making under sequential reward-penalty incentives in e-commerce delivery choices

Article Title: Similar aggregate but divergent disaggregate responses in human and AI under sequential reward-penalty incentives

Article References: Tao, S., Wang, L., Zhu, Z., & Lee, D.-H. (2026). Similar aggregate but divergent disaggregate responses in human and AI under sequential reward-penalty incentives. iScience, 29(11), Article 117800. https://doi.org/10.1016/j.isci.2026.117800

Image Credits: AI Generated

DOI: 10.1016/j.isci.2026.117800

Keywords: large language models, behavioral simulation, carbon credits, e-commerce logistics, decision-making, incentive design, generative agents, value of time, policy simulation, last-mile delivery, behavioral economics, DeepSeek

Cite Scienmag News

Denise Maddox. (October 8, 2026). AI Agents Match Human Averages but Diverge in How They Respond to Rewards and Penalties. Scienmag. https://scienmag.com/ai-agents-match-human-averages-but-diverge-in-how-they-respond-to-rewards-and-penalties/

Denise Maddox. "AI Agents Match Human Averages but Diverge in How They Respond to Rewards and Penalties." Scienmag, 8 October 2026, https://scienmag.com/ai-agents-match-human-averages-but-diverge-in-how-they-respond-to-rewards-and-penalties/. Accessed 8 October 2026.

Denise Maddox. "AI Agents Match Human Averages but Diverge in How They Respond to Rewards and Penalties." Scienmag. October 8, 2026. https://scienmag.com/ai-agents-match-human-averages-but-diverge-in-how-they-respond-to-rewards-and-penalties/

Tags: AI decision-making behavior comparisonAI ethics and decision-making fidelityAI modeling of human decision processesAI proxy for human policy responsesbehavioral economicsbehavioral simulationcarbon creditsdecision-makingDeepSeekdifferences in AI and human responses to incentivese-commerce logisticse-commerce logistics and carbon emission reductiongenerative agentshuman versus AI response to reward and penalty incentivesimpact of incentives on AI and human behaviorimplications for AI policy simulation validityincentive designlarge language modelslast-mile deliverylast-mile delivery incentive mechanismspolicy simulationsequential reward-and-penalty effect on decision-makingvalue of timevirtual societies simulation with language models
Share26Tweet16
Previous Post

Why Promising Immune Biomarkers Collapse in the Clinic: New Framework Predicts Failure Before Validation

Next Post

Housework Hours Linked to Depression in Single Mothers, Korean Study Finds

Related Posts

Privacy-First AI Learns to Spot DDoS Attacks Across IoT Networks Without Sharing Raw Data
Technology and Engineering

Privacy-First AI Learns to Spot DDoS Attacks Across IoT Networks Without Sharing Raw Data

October 8, 2026
Molybdenum Boosts Catalyst That Scrubs Two Pollutants at Once and Shrugs Off Poisoning
Technology and Engineering

Molybdenum Boosts Catalyst That Scrubs Two Pollutants at Once and Shrugs Off Poisoning

October 8, 2026
New LoRA-Based Method Steers Language Models Toward Specific Human Values
Technology and Engineering

New LoRA-Based Method Steers Language Models Toward Specific Human Values

October 8, 2026
Contrastive Learning Tames Out-of-Distribution Actions in Offline Reinforcement Learning
Technology and Engineering

Contrastive Learning Tames Out-of-Distribution Actions in Offline Reinforcement Learning

October 8, 2026
New AI Network Tackles Missing Data in Spatio-Temporal Forecasting
Technology and Engineering

New AI Network Tackles Missing Data in Spatio-Temporal Forecasting

October 8, 2026
Cloud AI Framework Maps Flood Danger Where Gauges and Models Are Missing
Technology and Engineering

Cloud AI Framework Maps Flood Danger Where Gauges and Models Are Missing

October 8, 2026
Next Post
Housework Hours Linked to Depression in Single Mothers, Korean Study Finds

Housework Hours Linked to Depression in Single Mothers, Korean Study Finds

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Two Endophytic Fungi Show Powerful Genome-Backed Defense Against Southern Corn Rust
  • Senolytic Drugs Show Promise Against Bone Loss Caused by Chemotherapy
  • Housework Hours Linked to Depression in Single Mothers, Korean Study Finds
  • AI Agents Match Human Averages but Diverge in How They Respond to Rewards and Penalties

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,150 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading