When ChatGPT burst onto the scene, critics quickly reached for a blunt but philosophically loaded word: bullshit. Not in the casual sense, but in the precise meaning given to the term by the philosopher Harry Frankfurt, who defined bullshit as language produced with complete indifference to truth—a mode of talk in which the speaker does not care whether what they say is accurate, only that it sounds convincing. A team of researchers at the University of Cambridge, led by Alessandro Trevisan, Harry Giddens, Sarah Dillon and Alan Blackwell, has now taken that insult seriously enough to measure it. In a study published in the journal AI & Society, they built what they call a Masterman Semantic Detector, or MSD—a statistical instrument they present rhetorically as a BS meter—and used it to show that the language of ChatGPT shares measurable features with two famously dysfunctional forms of human communication: political speech and the output of what the anthropologist David Graeber called bullshit jobs.
The theoretical backbone of the study is an unexpected figure: Margaret Masterman, founder of the Cambridge Language Research Unit and one of Wittgenstein’s students, who compiled his lecture notes into The Blue Book. More than sixty years ago, Masterman called for a semantic detector that could identify the actual message of a text rather than merely translate its words. Her technical work on conceptual networks in thesauri anticipated the corpus-driven methods that later fed into vector semantics, word embeddings and the TF*IDF measure attributed to her student Karen Spärck Jones. The Cambridge team argues that Masterman’s ambition, unrealisable with the computers of her era, can finally be attempted with modern machine learning—and that her Wittgensteinian framing of language as a set of games embedded in social forms of life is exactly the right lens for understanding what chatbots do.
The researchers dissect a product like ChatGPT into two anatomical components. The first is the underlying large language model, a probabilistic system trained on vast corpora of human writing that encodes many kinds of language game—giving orders, cracking jokes, reporting events, reasoning logically, and yes, bullshitting. The second is the dialogue management system, the hidden layer of prompts, guardrails, instruction-tuning and reinforcement learning from human feedback that turns the raw model into a friendly assistant. The authors interpret this dialogue management system through the literary concept of the paratext: the framing apparatus—titles, prefaces, blurbs—that controls how a text is received. Unlike a book’s paratext, however, the chatbot’s paratext actively conceals the statistical nature of the underlying model, presenting predicted word sequences as if they came from a knowledgeable, intentional interlocutor. This, they argue, is precisely where Frankfurtian bullshit enters: the deception lies not in the facts themselves but in misrepresenting what the system is.
This framing also explains why the popular term hallucination irritates the authors. A hallucination, as the Oxford English Dictionary defines it, requires a mind entertaining unfounded notions—and a language model has neither. The anthropomorphic vocabulary, they contend, feeds what Douglas Hofstadter named the Eliza effect back in 1995: the human tendency to read far more understanding into computer-generated strings of symbols than is warranted. The first-person voice of chatbot replies, the sycophantic tone, the apologetic contrition when errors are pointed out—all are conventions of the dialogue management system, engineered to sustain the illusion of a conversational partner rather than a sequence predictor.
To build the detector itself, the team borrowed a strategy from computer vision. When Microsoft developed the Kinect motion sensor, its neural networks were trained not on filmed humans but on artificially generated CGI animations of characteristic movements, allowing the recogniser to learn patterns from unlimited synthetic examples. The Cambridge group applied the same logic to language: instead of filming real bodies, they used ChatGPT to animate simulated human writing. They prompted the then-latest version of the model, GPT-4o, to write articles for Nature magazine, complete with the journal’s instructions to authors, on topics matching real papers. The results were, by any reader’s judgement, obvious bullshit—fabricated data tables, outright lies, disconnected arguments and long stretches of vague but science-flavoured prose. These were contrasted with 1,000 genuine Nature articles, chosen as an exemplar of precision, factual clarity and concision.
Two classifiers were then trained on this corpus. The first, an XGBoost model, relied on TF*IDF word frequencies to identify which terms best distinguished real science from machine-generated imitation, after removing stop-words and formatting artefacts such as the tell-tale references to figures that appear in journal articles but not in plain chatbot text. The second, a fine-tuned RoBERTa transformer, judged texts by their contextual embeddings—multidimensional representations of how words are used in surrounding context—rather than by the words themselves. Both classifiers achieved 100 percent accuracy in cross-validation, with confidence values of 99.84 and 99.97 percent respectively. Because the two models correlated only weakly with each other (r = 0.282), they were clearly detecting different linguistic features, so the researchers combined their outputs into a single score on a scale from minus five to plus five, later converted linearly to a 0–100 BS-meter percentage for public discussion.
The first experiment tested whether this meter, trained only on scientific and chatbot text, could detect bullshit in an entirely different domain: politics. Building on George Orwell’s critique of political language as intentionally uninformative, the team scored 45 UK party manifestos from 1945 to 2005, drawn from the Manifesto Project Database, against 45 transcripts of everyday spoken English from the British National Corpus—conversations, lessons, tutorials and casual chats that Orwell called demotic speech. The result was striking. The manifestos averaged 49.36 on the BS meter, while ordinary speech averaged just 9.40, a difference of enormous statistical significance (t(54) = 18.18, p far below 0.001). The authors note the irony: one might expect the educated register of scientists to resemble the political class and the conversational chatbot to resemble everyday talk. The opposite holds. Ordinary people speak like scientists, while politicians talk like ChatGPT.
The second experiment turned from politics to labour, testing a parallel between chatbot rhetoric and David Graeber’s account of bullshit jobs—paid employment so pointless that even the employee cannot justify its existence, yet must pretend otherwise. Drawing on Graeber’s typology of flunkies, goons, duct tapers, box tickers and taskmasters, the researchers collected 100 online texts: 50 from professions matching those categories and 50 from contrasting, demonstrably useful work such as laundry, cleaning or road repair. A repeated-measures analysis of variance revealed a highly significant main effect (F(99,1) = 43.73, p far below 0.001), with bullshit-job texts averaging 52.47 on the BS meter against 28.87 for the control sample. Text from pointless employment, in other words, statistically resembles the output of ChatGPT’s paratext far more than it resembles careful scientific writing. A weaker effect across Graeber’s five categories was also observed, though the authors caution that variability in their sample—some corporate advice texts scored high, and one wall-building guide was flagged by the detector GPTZero as 84 percent likely to be AI-generated—means future replications should sample more carefully.
The team is careful about what these results do and do not prove. They do not claim that large language models can only ever simulate the language game of bullshit; the underlying models encode many games, and a differently designed dialogue management system could emphasise others. Nor do they claim that correlation implies causation, or that their detector identifies the meaning of a text rather than a statistical signature of how words are used. What they have shown is that a property of language—whatever we ultimately call it—is reliably shared between the engineered persona of modern chatbots, the political misuse of English castigated by Orwell more than half a century ago, and the professional writing produced in jobs that even their holders cannot justify. The coincidence carries obvious social weight as chatbots are increasingly deployed in workplaces and public communication.
The study also closes a historical loop. Masterman’s student Spärck Jones, celebrated for the observation that computing is too important to be left to men, authored one of the earliest critiques of computational language models, raising questions about their semantic content that echo in today’s stochastic parrots debate. By naming their instrument after Masterman, the Cambridge team frames the current anxiety about AI slop not as a novel crisis but as the latest chapter in a long inquiry into how language, labour and machinery intertwine. Their conclusion is blunt: the statistical methods pioneered in 1950s Cambridge can now measure bullshit, and the measurements suggest we really are experiencing it—in our machines, our politics and, sometimes, our jobs.
Subject of Research: A statistical detector of Frankfurtian bullshit language, trained on ChatGPT and scientific text, applied to political speech and bullshit jobs
Article Title: The BS meter: detecting politics and labour through ChatGPT’s language
Article References: Trevisan, A., Giddens, H., Dillon, S., & Blackwell, A. F. (2026). The BS meter: detecting politics and labour through ChatGPT’s language. AI & SOCIETY. https://doi.org/10.1007/s00146-026-03238-9
Image Credits: AI Generated
DOI: 10.1007/s00146-026-03238-9
Keywords: ChatGPT, large language models, bullshit, Harry Frankfurt, Margaret Masterman, George Orwell, David Graeber, bullshit jobs, political language, TF-IDF, RoBERTa, AI & Society
Cite Scienmag News
Denise Maddox. (October 1, 2026). Scientists build a BS meter that reads ChatGPT, politicians and pointless jobs. Scienmag. https://scienmag.com/scientists-build-a-bs-meter-that-reads-chatgpt-politicians-and-pointless-jobs/
Denise Maddox. "Scientists build a BS meter that reads ChatGPT, politicians and pointless jobs." Scienmag, 1 October 2026, https://scienmag.com/scientists-build-a-bs-meter-that-reads-chatgpt-politicians-and-pointless-jobs/. Accessed 2 October 2026.
Denise Maddox. "Scientists build a BS meter that reads ChatGPT, politicians and pointless jobs." Scienmag. October 1, 2026. https://scienmag.com/scientists-build-a-bs-meter-that-reads-chatgpt-politicians-and-pointless-jobs/








