When millions of people ask a chatbot whether they should lend money to a friend with a troubled past, or whether they owe help to a colleague known for selfish behavior, the answer they receive is not a neutral reflection of universal morality. It is the output of a statistical system whose implicit ethical code has been shaped by training data, alignment procedures, and design choices that few users ever see. A new study published in PNAS Nexus suggests that those hidden codes vary dramatically across leading artificial intelligence systems, and that the differences could matter not just for individual decisions but for the cooperative fabric of society itself.
Alexandre S. Pires and colleagues set out to map what might be called the social norms of large language models. The research team examined 21 different LLMs, probing each system with a deceptively simple task: judge fictional people as good or bad after learning how those people behaved in a series of interpersonal dilemmas. The scenarios covered familiar moral terrain, such as whether a person chose to spend time and energy helping someone with a problem, or whether they gave food or money to another person in need. By varying the reputations of the people involved, the researchers could determine whether a model rewarded generosity universally or conditioned its approval on the character of the recipient.
The results revealed a broad consensus on one point and deep disagreement on another. Across the board, the models approved of helping and sharing with people who had already demonstrated themselves to be good. Cooperation with well-reputed individuals was, in the eyes of every system tested, virtuous conduct. The trouble began when the fictional actors interacted with individuals of ill repute. Here the models split into distinct camps, each implicitly endorsing a different philosophy about how decent people should treat those who have behaved badly.
Most, but not all, of the LLMs tested assigned a good reputation to people who cooperated with bad recipients. In this view, extending help even to a known defector is a laudable act, perhaps reflecting an ethic of unconditional kindness or a hope that generosity might reform the wrongdoer. But two models stood apart. Gemma 2 27B IT and Llama 3.1 8B tended to endorse shunning bad individuals, judging that cooperating with them was itself bad behavior. For these systems, morality is not just about what you give but about who deserves to receive it, and associating with the wrong people can stain your own standing.
The disagreements grew sharper when the researchers examined the opposite case: refusing to help or share with bad people. GPT-4o tended to penalize people for not cooperating with bad individuals, effectively treating withholding aid as a moral failing regardless of the recipient’s history. Llama 3.3 70B took the opposite stance, typically arguing that bad people should be punished for their prior lack of cooperation, so refusing to help them is entirely fine. Meanwhile, Gemini 1.5-Pro and Grok 2 proved inconsistent in their judgements, sometimes approving and sometimes condemning the same pattern of conduct depending on details of the scenario. For users seeking moral guidance, this means that identical questions posed to different popular chatbots can yield substantively different ethical verdicts.
Despite the noise, the researchers identified a discernible trend. Most LLM model families appear to be evolving toward a norm that social scientists call Simple Standing. Under this norm, cooperating is always good, and defecting against bad individuals is also deemed good. Simple Standing is permissive: it never requires anyone to punish or exclude wrongdoers, but it gives everyone license to walk away from them without guilt. It stands in contrast to stricter reputational codes found in human societies, such as norms under which failing to sanction a defector makes you complicit in their defection.
Why does this taxonomy matter beyond the philosophy seminar? Because norms are not merely descriptions of behavior; they are engines of it. The authors modeled what would happen if the various philosophies underpinning the LLMs’ judgements were universalized across an entire population. If everyone operated under the rules of Simple Standing, cooperation would indeed ensue, but it would not reach the high levels achieved under a sterner framework, one in which you should always fail to help bad people or be judged as bad yourself. That harsher regime, sometimes associated with norms of punitive standing, sustains near-universal cooperation by making association with defectors costly. In other words, the gentle, forgiving norm that most AI systems seem to be converging on produces a decent but less cooperative world than the unforgiving alternative.
The stakes rise when one considers how people actually use these systems. As more and more individuals turn to LLMs for advice on interpersonal matters, the ethical foundations behind the tools’ replies increasingly have the power to shape individual lives and, at scale, the social fabric itself. A person deciding whether to cut off a difficult relative, whether to give a second chance to an unreliable business partner, or whether to help a stranger with a bad reputation may be quietly absorbing the normative leanings of whichever model they consult. If a generation of users internalizes Simple Standing because their assistants endorse it, patterns of human cooperation, punishment, and reconciliation could shift in ways that no democratic process ever debated.
The study also uncovered subtler biases in how models render their judgements. LLM-based assessments depended on the gender and perceived cultural background of the recipients as well as the overall context of the fictional situation. The same act of helping or withholding could receive different moral evaluations depending on who was being helped, a finding that raises questions about consistency and fairness in AI moral reasoning. If a model’s verdict on your behavior shifts with the demographics of the people involved, the advice it gives may encode stereotypes that users never consented to and may not notice.
Can the problem be fixed with a simple instruction? Apparently not easily. The authors report that prompting interventions, such as instructing a model to adopt norms that promote overall cooperation, have a limited and inconsistent effect across LLMs. A nudge that steers one system toward a cooperative norm may leave another untouched or produce erratic behavior. This fragility suggests that the moral dispositions of these models are baked deeply into their training and alignment, and that shaping them deliberately will require more than clever phrasing in a user prompt. As AI systems become ubiquitous advisers on human relationships, the study implies, society will need to pay attention not only to what these machines know but to what, implicitly, they believe is right.
Subject of Research: Social norms and moral judgements of large language models regarding human cooperation
Article Title: AI’s social norms and their implications for society
Article References: AI’s social norms and their implications for society. (n.d.). Original publication
Image Credits: AI Generated
DOI: Not provided
Keywords: large language models, social norms, human cooperation, AI ethics, PNAS Nexus, reputation, Simple Standing, moral judgement, GPT-4o, Llama, Gemini, prompting interventions
Cite Scienmag News
Courtney Benton. (October 3, 2026). AI Models Disagree on the Ethics of Shunning Bad Actors, Study Finds. Scienmag. https://scienmag.com/ai-models-disagree-on-the-ethics-of-shunning-bad-actors-study-finds/
Courtney Benton. "AI Models Disagree on the Ethics of Shunning Bad Actors, Study Finds." Scienmag, 3 October 2026, https://scienmag.com/ai-models-disagree-on-the-ethics-of-shunning-bad-actors-study-finds/. Accessed 3 October 2026.
Courtney Benton. "AI Models Disagree on the Ethics of Shunning Bad Actors, Study Finds." Scienmag. October 3, 2026. https://scienmag.com/ai-models-disagree-on-the-ethics-of-shunning-bad-actors-study-finds/

