<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>legal and ethical considerations in Telegram research &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/legal-and-ethical-considerations-in-telegram-research/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 09:36:54 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>legal and ethical considerations in Telegram research &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Open-Source Platform Atenea Turns Telegram Into a Searchable Research Archive</title>
		<link>https://scienmag.com/open-source-platform-atenea-turns-telegram-into-a-searchable-research-archive/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 09:36:54 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[challenges in Telegram community tracking]]></category>
		<category><![CDATA[combating disinformation through Telegram archives]]></category>
		<category><![CDATA[computational social science]]></category>
		<category><![CDATA[continuous monitoring of Telegram channels]]></category>
		<category><![CDATA[Crime-as-a-Service]]></category>
		<category><![CDATA[cyber threat intelligence]]></category>
		<category><![CDATA[data collection]]></category>
		<category><![CDATA[Digital Services Act]]></category>
		<category><![CDATA[disinformation]]></category>
		<category><![CDATA[legal and ethical considerations in Telegram research]]></category>
		<category><![CDATA[named entity recognition]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[natural language processing for Telegram analysis]]></category>
		<category><![CDATA[open-source platform for Telegram data collection]]></category>
		<category><![CDATA[open-source software]]></category>
		<category><![CDATA[political discourse and protest movement analysis on Telegram]]></category>
		<category><![CDATA[reproducible research with Telegram data]]></category>
		<category><![CDATA[self-hosted Telegram data analysis workflows]]></category>
		<category><![CDATA[semantic search]]></category>
		<category><![CDATA[semantic search in Telegram research]]></category>
		<category><![CDATA[Telegram]]></category>
		<category><![CDATA[Telegram research archive]]></category>
		<category><![CDATA[tracking cybercriminal ecosystems on Telegram]]></category>
		<category><![CDATA[vector embeddings]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=221778</guid>

					<description><![CDATA[Researchers have unveiled Atenea, an open-source platform that continuously archives public Telegram channels, enriches messages with NLP, and enables semantic search across hundreds of millions of messages.]]></description>
										<content:encoded><![CDATA[<p>Telegram has quietly become one of the most important research environments on the internet. With more than one billion monthly active users reported in 2025, the messaging platform hosts public channels and groups that function as open archives of political debate, disinformation, protest movements, and, in darker corners, entire criminal marketplaces. Its open API makes this content technically accessible, yet researchers who try to study it systematically have long faced a frustrating problem: the communities they want to track can vanish overnight, migrate to new channels, or delete their histories, leaving one-off data collections incomplete and irreproducible. A new open-source platform called Atenea, described in the journal SoftwareX by Alfonso de Paz and David Arroyo of the Spanish National Research Council, aims to solve this by turning continuous Telegram monitoring, natural language processing, and semantic search into a single self-hosted workflow.</p>
<p>The motivation goes beyond simple data collection. In cybercriminal ecosystems, low-level identifiers such as channel handles are easily replaced, but higher-level semantic structures, including the language, products, and networks that define a campaign, tend to persist even as the surface infrastructure changes. This is why the authors argue for continuous observation rather than snapshot sampling. Archiving even short-lived communities supports the identification of illegal content under the EU Digital Services Act, enables the extraction of Indicators of Compromise for characterising the Crime-as-a-Service economy, and can ultimately assist the prosecution of cybercriminal activity. At the same time, many researchers currently depend on commercial Software-as-a-Service intermediaries such as TGStat or Telemetrio, which offer basic metrics but introduce opaque filtering, undermine corpus representativeness and reproducibility, and lock data inside proprietary silos without multi-dimensional enrichment.</p>
<p>Existing open-source tools each cover only part of this workflow. pytopicgram focuses on local extraction, preprocessing, sentiment analysis, and topic modelling; 4CAT provides persistent datasets, queued processing, and modular analysis processors; and TeleCatch offers a remote interface for filtered collection and message extraction. A more comprehensive framework called MST has been proposed in the literature, but because no versioned, publicly accessible source repository could be identified, the authors treat it strictly as related work rather than a direct competitor. In a capability comparison verified against public repositories and documentation as of August 2026, Atenea stands out as the only tool of the four that combines continuous monitoring, a persistent server-side message repository, named entity recognition, sentiment analysis, embedding generation, and, critically, true semantic retrieval over stored messages.</p>
<p>Architecturally, Atenea separates collection, processing, storage, retrieval, and exploration behind a common REST interface, with heavy inference decoupled from the core backend so that monitoring continues uninterrupted while services scale or evolve independently. Its data model revolves around three concepts: rooms, which are monitorable Telegram channels or groups; users, whether human accounts or bots; and seeds, external references such as invite links that define which rooms should be ingested. A dedicated Telegram Auth Manager registers and manages multiple API credentials, increasing effective request capacity and handling temporary restrictions. Two load-balancing strategies, one that splits work across credentials and one that groups resources by credential, keep collection resilient, and credential-to-room associations allow access recovery when a particular credential expires.</p>
<p>The processing layer relies on asynchronous, distributed ETL pipelines built from Telethon, Redis, Celery, and Celery Beat, triggered manually or by an automated scheduler. PostgreSQL serves as the transactional system of record, Elasticsearch powers lexical search and exploration through Kibana, and Qdrant stores vector embeddings for semantic retrieval. Downloaded media live outside the relational database in S3-compatible object storage, deployable locally via Garagehq, while PostgreSQL retains their provenance, metadata, hashes, and processing states. Database-level locking and configurable task concurrency prevent worker collisions across pipeline stages. Once collected, messages flow into a service-oriented enrichment layer offering named entity recognition, sentiment analysis, and embedding generation through independent microservices, including custom APIs, vLLM, or Ollama backends.</p>
<p>The enrichment pipeline itself is notably hybrid and language-aware. During ingestion, Atenea cleans message text and assigns an ISO 639-1 language code using the Efficient Language Detector library, discarding detections below a confidence threshold of 0.15. Before any model-based processing runs, regular-expression detectors extract structured indicators, including email addresses, URLs, Telegram mentions, hashtags, and cryptocurrency wallet addresses for Bitcoin, Ethereum, Dash, and Monero, with emoji tokenisation handled separately. These detectors preserve character offsets and operate independently of language. Model-based named entity recognition currently runs for English and Spanish using spaCy 3.8.14 with its large English and Spanish pipelines, mapping model-specific labels onto a canonical taxonomy. Adding a new language requires only installing a compatible spaCy model, registering its code, and redeploying, without changes to the enrichment API, task pipeline, or database schema.</p>
<p>To demonstrate the platform at scale, the authors describe a 2.5-year operational deployment that, as of a July 2026 snapshot, held more than 191 million messages from 6,412 monitored channels and groups, with history reaching back to 2016 and nearly 40 million extracted entities alongside 7.1 million inferred embeddings. A case study built on the DarkGram dataset, a research catalogue of cybercriminal Telegram channels, shows the full workflow in action. From 329 imported seeds, the populate pipeline resolved 207 monitorable rooms. An initial historical-recovery scan stored over 510,000 messages, peaking at 177,357 messages in a single hour, before settling into recurrent four-hour scans averaging roughly 128 new messages per day. The result was a tagged, reproducible collection of 574,432 messages spanning nearly nine years of channel history.</p>
<p>The enrichment results offer a descriptive profile of what automated extraction can reveal. Of the DarkGram messages, 378,905 qualified for English or Spanish model-based processing, and the hybrid extractor produced 175,021 unique entities across the full snapshot. Comparisons against the rest of the repository proved revealing: URL labels appeared at nearly identical rates in both collections, around 27 percent of messages, but PRODUCT labels were overrepresented in the cybercriminal collection at 4.7 percent versus 0.7 percent in the baseline, a pattern consistent with illicit marketplaces. References to people, by contrast, were far rarer in DarkGram, reflecting the baseline repository&#8217;s focus on disinformation and political-news channels. Media cataloguing recorded metadata for 147,646 attachments totalling 6.41 tebibytes of potential content, with opaque .bin files and compressed archives accounting for roughly 6.09 tebibytes, a distribution the authors note is consistent with packaged software and resources circulating in such channels.</p>
<p>Perhaps the most striking demonstration is the platform&#8217;s semantic capability. On external hardware equipped with two NVIDIA L40 S GPUs, a vLLM server running the Qwen3-Embedding-4B model generated 2,560-dimensional vector embeddings for more than 563,000 messages in roughly five hours, averaging 107,400 messages per hour with a peak of 35,224 in a single fifteen-minute bin. Stored in Qdrant and linked to their source messages, these embeddings allow researchers to retrieve content matching natural-language descriptions, such as credential theft or illicit marketplace activity, rather than relying solely on dates or keyword matches. Combined with statistical anomaly detection, which flagged two high-activity windows in the DarkGram collection, including one week in late 2025 with 11,005 messages and a z-score of 8.17, the system turns raw monitoring into a queryable analytical environment.</p>
<p>The implications reach well beyond cybersecurity. The authors position Atenea as a flexible backend for computational social science, journalism, and cyber threat intelligence, supporting workflows from network analysis to narrative detection, social listening during protests and crises, and the automated recovery of indicators such as IP addresses, domains, URLs, wallet addresses, and cryptographic hashes. Because the platform is released under the MIT license, with full documentation, a reproducible deployment capsule, and modular microservices that can be extended without touching the core, it lowers the barrier for any research group to build a long-running, auditable observation post on one of the world&#8217;s most consequential yet least-studied communication platforms. In an era when online communities can disappear in an instant, tools that preserve both the messages and their meaning may prove as valuable as the analyses they enable.</p>
<p><strong>Subject of Research:</strong> An open-source platform for continuous Telegram monitoring, NLP enrichment, and semantic retrieval</p>
<p><strong>Article Title:</strong> Atenea: An open-source platform for continuous telegram monitoring, NLP enrichment, and semantic retrieval</p>
<p><strong>Article References:</strong> de Paz, A., &amp; Arroyo, D. (2026). Atenea: An open-source platform for continuous telegram monitoring, NLP enrichment, and semantic retrieval. <em>SoftwareX, 36</em>, Article 103067. <a href="https://doi.org/10.1016/j.softx.2026.103067" rel="noopener noreferrer">https://doi.org/10.1016/j.softx.2026.103067</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.softx.2026.103067" rel="noopener noreferrer">10.1016/j.softx.2026.103067</a></p>
<p><strong>Keywords:</strong> Telegram, open-source software, natural language processing, semantic search, cyber threat intelligence, named entity recognition, computational social science, disinformation, Crime-as-a-Service, vector embeddings, data collection, Digital Services Act</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">221778</post-id>	</item>
	</channel>
</rss>
