<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>source verification &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/source-verification/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 11 Oct 2026 00:16:30 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>source verification &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Open-Source AI Chatbot Brings Verifiable Answers to Greek Newsrooms</title>
		<link>https://scienmag.com/open-source-ai-chatbot-brings-verifiable-answers-to-greek-newsrooms/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 11 Oct 2026 00:16:30 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[addressing misinformation with AI]]></category>
		<category><![CDATA[AI-based question answering with source attribution]]></category>
		<category><![CDATA[Chroma]]></category>
		<category><![CDATA[citation-enabled AI chatbots]]></category>
		<category><![CDATA[combating fake news with retrieval-augmented generation]]></category>
		<category><![CDATA[Greek language AI news solutions]]></category>
		<category><![CDATA[Greek language NLP]]></category>
		<category><![CDATA[journalism]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[Llama-KriKri]]></category>
		<category><![CDATA[locally deployable AI for media organizations]]></category>
		<category><![CDATA[misinformation]]></category>
		<category><![CDATA[newsroom AI]]></category>
		<category><![CDATA[Open-source AI chatbot for Greek newsrooms]]></category>
		<category><![CDATA[open-source NLP models for news verification]]></category>
		<category><![CDATA[open-source software]]></category>
		<category><![CDATA[privacy-focused AI tools for journalism]]></category>
		<category><![CDATA[query expansion]]></category>
		<category><![CDATA[retrieval-augmented generation]]></category>
		<category><![CDATA[retrieval-augmented generation in journalism]]></category>
		<category><![CDATA[source verification]]></category>
		<category><![CDATA[transparency in AI-generated news]]></category>
		<category><![CDATA[vector database]]></category>
		<category><![CDATA[verifiable fact-checking AI tools]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=260458</guid>

					<description><![CDATA[Greek researchers have built a fully open-source retrieval-augmented chatbot that answers journalists' questions using 73,000 archived Greek news articles, complete with citations, metadata filtering, and local deployment.]]></description>
										<content:encoded><![CDATA[<p>In an age when misinformation can travel around the world before a correction is even drafted, journalists are increasingly looking for artificial intelligence tools that do more than generate fluent text. They want systems that can prove where every claim came from. A team of Greek researchers has now built exactly that: a fully open-source, retrieval-augmented generation chatbot designed specifically for Greek newsrooms, described in a paper published in the journal SoftwareX. The system answers journalists&#8217; questions using only a curated archive of Greek news articles, attaches citations to every response, and runs entirely on locally deployable components, freeing media organizations from dependence on proprietary cloud platforms.</p>
<p>The motivation behind the project is straightforward. Large language models have demonstrated impressive performance on tasks that matter to journalists, including summarization, translation, question answering, and text classification. But general-purpose models rely on knowledge frozen during training, which means their answers can be outdated, unsupported by evidence, or impossible to trace back to a specific source. In a newsroom, where every published claim must be verifiable and attributable, those limitations are disqualifying. Retrieval-augmented generation, or RAG, offers a way around the problem: instead of answering from memory, the model first retrieves relevant documents from a trusted database and then generates a response grounded in that retrieved context, with the sources available for inspection.</p>
<p>The new system, developed by Nikolaos Armenakis, Panagiotis Germanakos, Christos Makris, and Constantinos Mourlas, is built entirely from open-source parts. At its core sits Llama-KriKri-8B-Instruct, an instruction-tuned large language model created specifically for Greek, paired with the Multilingual-E5-large-instruct embedding model, which ranks highly on multilingual embedding benchmarks and handles Greek particularly well. The knowledge base consists of roughly 73,000 articles scraped from four major Greek news outlets: kathimerini.gr, efsyn.gr, skai.gr, and zougla.gr. The choice of Greek is deliberate. Despite real progress, Greek natural language processing still lags behind English in available resources, tools, and task coverage, leaving Greek-language newsrooms with few dedicated AI options.</p>
<p>The technical pipeline begins with data acquisition and preprocessing. Articles were collected using Beautiful Soup and Selenium, then loaded from CSV files with normalized dates converted into Unix timestamps so that time-based filtering becomes possible. Each document is split into chunks of 500 tokens, the limit imposed by the embedding model, with a 20 percent token overlap between adjacent chunks to preserve semantic continuity and prevent important context from being lost at chunk boundaries. Every chunk receives a unique identifier combining the article&#8217;s URL, its row index in the original dataset, and the chunk&#8217;s position within the article, guaranteeing that any retrieved fragment can be traced back to its exact origin. The chunks and their metadata, including URL, title, author, website, date, and section, are stored in the open-source Chroma vector database, which was designed to support incremental updates so that newly published articles can be added without rebuilding the entire index.</p>
<p>Where the system genuinely departs from a standard RAG setup is in its two-stage retrieval process with query expansion. When a journalist submits a question, the system first retrieves only the single most relevant article, asks the language model to summarize it in exactly five words, and appends that summary to the original query. The expanded query then drives a second, more precise retrieval pass. The researchers tested four expansion strategies empirically, measuring precision, response quality, response time, and source correctness. They found that a five-word summary provided enough contextual enrichment without introducing semantic noise; an earlier ten-word version pulled in irrelevant material. If no article clears the minimum similarity threshold, the user is simply told so rather than being given a speculative answer.</p>
<p>Filtering is another key feature. Before submitting a query, users can restrict the search by author, source website, section, or date range through dropdown menus and date pickers in a Streamlit-based interface. These filters are converted into Chroma-compatible conditions, with author, website, and section matched through logical operators and dates applied as Unix timestamp ranges, all combined through logical conjunction when multiple filters are selected. A slider lets users set how many source documents, from one to twenty, should inform the final answer. Once retrieval is complete, the system assembles a composite prompt containing the full conversation history, an instruction line, the five-word summary, the retrieved document excerpts with their metadata, and the user&#8217;s current question.</p>
<p>The system prompt instructs the model to act as a journalistic assistant: to respond objectively, base its answers exclusively on the provided sources, avoid personal opinions, and write in an analytical but concise style. Generation settings are tuned for reliability rather than creativity. The model produces at most 1,024 new tokens, a value chosen empirically, and runs at a temperature of 0.1. That low temperature makes outputs more deterministic and closely aligned with the retrieved sources, which is exactly what a fact-checking context demands, while stopping short of zero to preserve a small degree of linguistic flexibility and avoid rigid, repetitive phrasing. Duplicate source URLs are stripped from the final answer, and the remaining references appear in expandable interface elements so the journalist can open the original articles and verify the evidence directly.</p>
<p>The team put the system in front of thirty senior journalism students in a controlled laboratory evaluation, asking them to assess responses to newsroom questions of varying cognitive complexity. The results were encouraging but revealing. The system scored well on relevance, accuracy, clarity, and unbiasedness for low- and medium-complexity questions involving fact retrieval and information synthesis, but performance dropped for high-complexity questions requiring logical reasoning and comparison, a limitation consistent with what is known about current RAG architectures. On the user-experience side, the numbers were strong across the board: perceived usefulness, search effectiveness, satisfaction with answers, and willingness to adopt each scored 8.2, ease of use reached 8.6, and the calculated Net Promoter Score of 50 indicated a solid tendency among participants to recommend the tool to others.</p>
<p>The researchers are careful to position the chatbot as an assistant rather than an autonomous article generator. Its purpose is to support journalistic research, archive exploration, and evidence verification, with editorial judgment remaining firmly in human hands. That framing aligns with a broader body of work on AI in newsrooms, including studies of GraphRAG for complex multi-source queries, on-premises retrieval frameworks emphasizing auditability, and recent Greek studies examining ChatGPT-assisted analysis of political rhetoric and the fragmented adoption of generative AI in Greek media. What distinguishes this system is its combination of Greek-language specialization, a fully open and locally deployable stack, and explicit source attribution as a first-class feature rather than an afterthought.</p>
<p>The implications reach beyond Greece. Because the architecture depends only on open components, Python, Streamlit, Chroma, PyTorch, Hugging Face Transformers, and suitably tuned open models, the same methodology can be adapted to other languages and domains wherever a suitable knowledge base exists. The authors plan to expand the dataset with more diverse and balanced sources, refine query expansion, ensure that retrieved passages are drawn from multiple outlets, optimize dialogue history handling, and benchmark the system against alternative open-source and proprietary language models. Planned extensions also include support for user-uploaded documents, multi-user deployment, and agent-based journalist profiles tailored to individual workflows. At a moment when public trust in media is under sustained pressure from propaganda, hate speech, and rapidly spreading falsehoods, a newsroom AI that shows its work, runs on infrastructure the newsroom controls, and refuses to answer without evidence offers a compelling template for responsible AI-assisted journalism anywhere in the world.</p>
<p><strong>Subject of Research:</strong> An open-source retrieval-augmented generation chatbot for verifiable question answering over Greek news archives in journalism</p>
<p><strong>Article Title:</strong> Leveraging open-source technologies for a retrieval-augmented generation chatbot in Greek newsrooms</p>
<p><strong>Article References:</strong> Leveraging open-source technologies for a retrieval-augmented generation chatbot in Greek newsrooms. (n.d.). <a href="https://doi.org/10.1016/j.softx.2026.103103" rel="noopener noreferrer">https://doi.org/10.1016/j.softx.2026.103103</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.softx.2026.103103" rel="noopener noreferrer">10.1016/j.softx.2026.103103</a></p>
<p><strong>Keywords:</strong> retrieval-augmented generation, large language models, journalism, Greek language NLP, open-source software, misinformation, vector database, newsroom AI, Llama-KriKri, Chroma, query expansion, source verification</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">260458</post-id>	</item>
	</channel>
</rss>
