<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>AI extraction crisis in Wikipedia &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/ai-extraction-crisis-in-wikipedia/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 12 Sep 2026 13:03:16 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>AI extraction crisis in Wikipedia &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Two Decades of Wikipedia Research Reveal a Fractured Field Shaped by Big Data and AI</title>
		<link>https://scienmag.com/two-decades-of-wikipedia-research-reveal-a-fractured-field-shaped-by-big-data-and-ai/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 13:03:16 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI extraction crisis in Wikipedia]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[bibliometric analysis of Wikipedia studies]]></category>
		<category><![CDATA[bibliometrics]]></category>
		<category><![CDATA[big data]]></category>
		<category><![CDATA[computational and data-driven studies of Wikipedia]]></category>
		<category><![CDATA[consequences of research siloing in Wikipedia studies]]></category>
		<category><![CDATA[digital commons]]></category>
		<category><![CDATA[evolution of Wikipedia research over two decades]]></category>
		<category><![CDATA[history of Wikipedia as a research laboratory]]></category>
		<category><![CDATA[impact of AI and large language models on Wikipedia]]></category>
		<category><![CDATA[interdisciplinary fragmentation in Wikipedia research]]></category>
		<category><![CDATA[interdisciplinary research]]></category>
		<category><![CDATA[knowledge graphs]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[methodological approaches in bibliometric studies]]></category>
		<category><![CDATA[open data and Wikipedia research accessibility]]></category>
		<category><![CDATA[peer production]]></category>
		<category><![CDATA[science mapping]]></category>
		<category><![CDATA[social science research on Wikipedia]]></category>
		<category><![CDATA[Wikidata]]></category>
		<category><![CDATA[Wikimedia]]></category>
		<category><![CDATA[Wikipedia]]></category>
		<category><![CDATA[Wikipedia research]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=194643</guid>

					<description><![CDATA[A new bibliometric study maps twenty years of Wikipedia research, revealing a fragmented field whose corporate AI lineage helped normalize the extraction of community-produced knowledge.]]></description>
										<content:encoded><![CDATA[<p>Wikipedia has spent more than twenty years as one of the most intensively studied organizations on the planet, with over 6,000 scholarly studies published by 2019 alone. Researchers have called it the most important laboratory for social scientific and computing research in history, a status it earned partly by making its entire database freely downloadable as early as 2002. Yet a new bibliometric investigation argues that the enormous body of work built around the encyclopedia is not a coherent field at all, but a fragmented collection of disciplinary silos whose findings rarely reach one another. That fragmentation, the study suggests, has real consequences at a moment when Wikipedia faces what researchers describe as an AI extraction crisis, with large language models consuming its content while simultaneously drawing readers away from the site itself.</p>
<p>The research, conducted by Steve Jankowski of the Media Studies Department at the University of Amsterdam and published in the journal AI &amp; Society, takes an unusual methodological approach to this problem. Rather than conducting a conventional systematic literature review, Jankowski performed a bibliometric science mapping of the most popularly cited research about Wikipedia and other Wikimedia projects between 2005 and 2025. Starting from 524 documents collected through Google Scholar, the study purposively sampled 70 of the most heavily cited works and subjected them to a reflexive thematic analysis, a qualitative technique in which themes are defined by conceptual coherence rather than statistical clustering. The goal was to excavate the intellectual lineages that have shaped how scholars understand Wikimedia projects and to identify the epistemic boundaries that keep the field divided.</p>
<p>The analysis surfaced seven recurring matters of concern that have animated two decades of Wikimedia scholarship. These include the status of Wikipedia as digital property or commons, the development of corporate AI software from Wikipedia content, the social and cultural consequences of social media platforms, the optimization of human-computer interaction systems, the shifting cultural significance of the encyclopedia, Wikipedia&#8217;s reliability as an information source, and the sociotechnical governance of distributed authority. Jankowski organizes these concerns into three magnitudes of citation impact, described as major, minor, and diminished, and traces how each cluster rose, plateaued, or declined across the twenty-year window.</p>
<p>The most heavily cited concern centers on the political economy of digital property and commons, a debate launched in the mid-2000s by Yochai Benkler&#8217;s concept of commons-based peer production in The Wealth of Networks and Don Tapscott&#8217;s market-oriented vision in Wikinomics. These foundational texts, alongside work by Clay Shirky, Lawrence Lessig, and Axel Bruns, refashioned Wikipedia from an online curiosity into a serious research object of economic innovation. Their citation curve follows what Jankowski calls a marathon pattern, peaking around 2015 but never falling below 1,500 annual citations, largely sustained by the enduring influence of Benkler&#8217;s book. Notably, these works relied on evaluative and theory-building methods, using eclectic case comparisons rather than statistical analysis to make claims about the social value of collaborative production.</p>
<p>The second major concern is the one with the most consequential implications for Wikipedia&#8217;s present troubles: the extraction of semantic knowledge from Wikimedia content to build corporate artificial intelligence. Jankowski traces two waves of this research. The first, exemplified by work from Yahoo! Research scientists on deriving semantic relatedness from Wikipedia, plateaued between 2013 and 2017. The second wave arrived with the 2012 launch of Wikidata, a project that converted Wikimedia content into structured knowledge graphs and received substantial funding from Microsoft co-founder Paul Allen&#8217;s Institute for Artificial Intelligence, Google, and the Gordon and Betty Moore Foundation. The study documents tight corporate entanglement throughout this research lineage, including co-authors from Facebook AI Research and Wikidata&#8217;s lead developer working as an ontologist at Google. Citation counts for this concern climbed steadily to a peak in 2021 and 2022, then dropped sharply following the release of OpenAI&#8217;s chatbot, the moment the study identifies as the beginning of Wikipedia&#8217;s AI extraction crisis.</p>
<p>This crisis is not merely academic. The Wikimedia Foundation has warned that large language models are trained on Wikipedia content, often weighing it more heavily than any other dataset, while answer-based interfaces reduce direct human visits to the encyclopedia. If knowledge seekers obtain Wikipedia content through chatbots and search results rather than the site itself, both readers and potential contributors become disintermediated from the community, severing the feedback loop that sustains peer production. Jankowski notes that these warnings echo concerns raised more than fifteen years ago, when scholars of peer production cautioned that packaging and distributing collaborative content disconnects it from the norms, protocols, and community structures that created it. Yet the fragmented nature of Wikimedia research has prevented these warnings from circulating as a unified debate within the field.</p>
<p>That fragmentation, the study argues, is rooted in deep methodological divisions. Drawing on Hans-Georg Gadamer&#8217;s philosophy of the human sciences and on science and technology studies concepts such as trading zones and boundary objects, Jankowski characterizes Wikimedia research as a fractionated trading zone: a meeting place of communities with incommensurable epistemic traditions who nonetheless organize their work around a shared object. Mapping the seven concerns by methodological composition reveals stark contrasts. Research on optimizing Wikipedia and on corporate AI software is overwhelmingly quantitative, grounded in statistical analysis and computational modeling. Research on Wikipedia&#8217;s cultural significance is almost entirely evaluative, consisting of ethnographies, histories, and cultural critiques that seek to understand how the encyclopedia came to be what it is, in line with Gadamer&#8217;s insistence that human science aims at historical understanding rather than predictive regularity.</p>
<p>Between these poles, the study identifies a methodological inter-language that could mediate across the divide. Three concerns, those addressing social media platforms, distributed authority, and Wikipedia as a reliable source, employ nearly equal mixtures of quantitative, qualitative, and evaluative methods. Jankowski describes these as an inter-language translation network, positions from which insights can be converted and communicated across otherwise distant research communities. For example, the gulf between scholars studying Wikipedia&#8217;s cultural meaning and engineers building AI from its content could be bridged through conversations about social media platforms and information reliability, concerns that both camps can recognize. The study also finds, through co-citation analysis, that the two most influential concerns are ironically the least cited within Wikimedia research itself: only 40 percent of citations to the digital property literature and just 24 percent of citations to the social media platforms literature come from papers with Wikipedia as a keyword, suggesting these insights circulate mainly outside the field&#8217;s core.</p>
<p>The historical arc that emerges from the citation curves tells a story in four waves. A first wave from 2005 to 2007 established Wikipedia as a valid research object, driven by researchers drawn to both its accessible data and its utopian promise of peer production. A second wave, beginning in 2007, theorized ownership and generated the vocabulary of commons-based peer production, wikinomics, and produsage, giving the wider world a language for a medium it did not yet know how to discuss. A third wave broke out in 2011, when machine learning and knowledge base research departed from the peer production frame and pursued semantic extraction as an end in itself. A final wave, arriving late with the mid-2010s epistemic crisis of networked propaganda, mapped the sociotechnics of social media platforms and continues to rise, now standing nearly equal in citations to the AI extraction literature.</p>
<p>Jankowski concludes that fragmentation need not arrest collaboration, pointing out that decentralized, contested conditions are part and parcel of Wikimedia itself. The remedy he proposes is neither another literature review nor a forced unification, but deliberate cross-disciplinary citation practices that put AI researchers and cultural scholars into direct conversation, building on a recent manifesto calling to unite and reignite critical Wikimedia research. The deeper lesson of the bibliometric history is a recursive one: the methodological choices of researchers have not merely described Wikipedia but have changed its cultural meaning and sociotechnical function, from commons to commodity to training corpus. As the encyclopedia confronts an answer-based web increasingly mediated by the very AI systems its content helped build, the study argues that understanding how we came to know Wikipedia may be the first step toward deciding what it should become.</p>
<p><strong>Subject of Research:</strong> A bibliometric history of Wikimedia research methodologies, peer production, and artificial intelligence from 2005 to 2025.</p>
<p><strong>Article Title:</strong> Wikimedia research methodologies: a bibliometric history of peer production and AI (2005–2025)</p>
<p><strong>Article References:</strong> Jankowski, S. (2026). Wikimedia research methodologies: a bibliometric history of peer production and AI (2005–2025). <em>AI &amp;amp; SOCIETY</em>. <a href="https://doi.org/10.1007/s00146-026-03359-1" rel="noopener noreferrer">https://doi.org/10.1007/s00146-026-03359-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s00146-026-03359-1" rel="noopener noreferrer">10.1007/s00146-026-03359-1</a></p>
<p><strong>Keywords:</strong> Wikipedia, Wikimedia, bibliometrics, peer production, artificial intelligence, large language models, Wikidata, science mapping, big data, digital commons, knowledge graphs, interdisciplinary research</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">194643</post-id>	</item>
	</channel>
</rss>
