Thursday, October 1, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

Open-Source Platform Atenea Turns Telegram Into a Searchable Research Archive

October 1, 2026
in Technology and Engineering
Denise Maddox
By Denise Maddox Scienmag Editorial Profile - Mechanical Engineering
Reading Time: 5 mins read
0
Open-Source Platform Atenea Turns Telegram Into a Searchable Research Archive

Open-Source Platform Atenea Turns Telegram Into a Searchable Research Archive

Open-Source Platform Atenea Turns Telegram Into a Searchable Research Archive

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Telegram has quietly become one of the most important research environments on the internet. With more than one billion monthly active users reported in 2025, the messaging platform hosts public channels and groups that function as open archives of political debate, disinformation, protest movements, and, in darker corners, entire criminal marketplaces. Its open API makes this content technically accessible, yet researchers who try to study it systematically have long faced a frustrating problem: the communities they want to track can vanish overnight, migrate to new channels, or delete their histories, leaving one-off data collections incomplete and irreproducible. A new open-source platform called Atenea, described in the journal SoftwareX by Alfonso de Paz and David Arroyo of the Spanish National Research Council, aims to solve this by turning continuous Telegram monitoring, natural language processing, and semantic search into a single self-hosted workflow.

The motivation goes beyond simple data collection. In cybercriminal ecosystems, low-level identifiers such as channel handles are easily replaced, but higher-level semantic structures, including the language, products, and networks that define a campaign, tend to persist even as the surface infrastructure changes. This is why the authors argue for continuous observation rather than snapshot sampling. Archiving even short-lived communities supports the identification of illegal content under the EU Digital Services Act, enables the extraction of Indicators of Compromise for characterising the Crime-as-a-Service economy, and can ultimately assist the prosecution of cybercriminal activity. At the same time, many researchers currently depend on commercial Software-as-a-Service intermediaries such as TGStat or Telemetrio, which offer basic metrics but introduce opaque filtering, undermine corpus representativeness and reproducibility, and lock data inside proprietary silos without multi-dimensional enrichment.

Existing open-source tools each cover only part of this workflow. pytopicgram focuses on local extraction, preprocessing, sentiment analysis, and topic modelling; 4CAT provides persistent datasets, queued processing, and modular analysis processors; and TeleCatch offers a remote interface for filtered collection and message extraction. A more comprehensive framework called MST has been proposed in the literature, but because no versioned, publicly accessible source repository could be identified, the authors treat it strictly as related work rather than a direct competitor. In a capability comparison verified against public repositories and documentation as of August 2026, Atenea stands out as the only tool of the four that combines continuous monitoring, a persistent server-side message repository, named entity recognition, sentiment analysis, embedding generation, and, critically, true semantic retrieval over stored messages.

Architecturally, Atenea separates collection, processing, storage, retrieval, and exploration behind a common REST interface, with heavy inference decoupled from the core backend so that monitoring continues uninterrupted while services scale or evolve independently. Its data model revolves around three concepts: rooms, which are monitorable Telegram channels or groups; users, whether human accounts or bots; and seeds, external references such as invite links that define which rooms should be ingested. A dedicated Telegram Auth Manager registers and manages multiple API credentials, increasing effective request capacity and handling temporary restrictions. Two load-balancing strategies, one that splits work across credentials and one that groups resources by credential, keep collection resilient, and credential-to-room associations allow access recovery when a particular credential expires.

The processing layer relies on asynchronous, distributed ETL pipelines built from Telethon, Redis, Celery, and Celery Beat, triggered manually or by an automated scheduler. PostgreSQL serves as the transactional system of record, Elasticsearch powers lexical search and exploration through Kibana, and Qdrant stores vector embeddings for semantic retrieval. Downloaded media live outside the relational database in S3-compatible object storage, deployable locally via Garagehq, while PostgreSQL retains their provenance, metadata, hashes, and processing states. Database-level locking and configurable task concurrency prevent worker collisions across pipeline stages. Once collected, messages flow into a service-oriented enrichment layer offering named entity recognition, sentiment analysis, and embedding generation through independent microservices, including custom APIs, vLLM, or Ollama backends.

The enrichment pipeline itself is notably hybrid and language-aware. During ingestion, Atenea cleans message text and assigns an ISO 639-1 language code using the Efficient Language Detector library, discarding detections below a confidence threshold of 0.15. Before any model-based processing runs, regular-expression detectors extract structured indicators, including email addresses, URLs, Telegram mentions, hashtags, and cryptocurrency wallet addresses for Bitcoin, Ethereum, Dash, and Monero, with emoji tokenisation handled separately. These detectors preserve character offsets and operate independently of language. Model-based named entity recognition currently runs for English and Spanish using spaCy 3.8.14 with its large English and Spanish pipelines, mapping model-specific labels onto a canonical taxonomy. Adding a new language requires only installing a compatible spaCy model, registering its code, and redeploying, without changes to the enrichment API, task pipeline, or database schema.

To demonstrate the platform at scale, the authors describe a 2.5-year operational deployment that, as of a July 2026 snapshot, held more than 191 million messages from 6,412 monitored channels and groups, with history reaching back to 2016 and nearly 40 million extracted entities alongside 7.1 million inferred embeddings. A case study built on the DarkGram dataset, a research catalogue of cybercriminal Telegram channels, shows the full workflow in action. From 329 imported seeds, the populate pipeline resolved 207 monitorable rooms. An initial historical-recovery scan stored over 510,000 messages, peaking at 177,357 messages in a single hour, before settling into recurrent four-hour scans averaging roughly 128 new messages per day. The result was a tagged, reproducible collection of 574,432 messages spanning nearly nine years of channel history.

The enrichment results offer a descriptive profile of what automated extraction can reveal. Of the DarkGram messages, 378,905 qualified for English or Spanish model-based processing, and the hybrid extractor produced 175,021 unique entities across the full snapshot. Comparisons against the rest of the repository proved revealing: URL labels appeared at nearly identical rates in both collections, around 27 percent of messages, but PRODUCT labels were overrepresented in the cybercriminal collection at 4.7 percent versus 0.7 percent in the baseline, a pattern consistent with illicit marketplaces. References to people, by contrast, were far rarer in DarkGram, reflecting the baseline repository’s focus on disinformation and political-news channels. Media cataloguing recorded metadata for 147,646 attachments totalling 6.41 tebibytes of potential content, with opaque .bin files and compressed archives accounting for roughly 6.09 tebibytes, a distribution the authors note is consistent with packaged software and resources circulating in such channels.

Perhaps the most striking demonstration is the platform’s semantic capability. On external hardware equipped with two NVIDIA L40 S GPUs, a vLLM server running the Qwen3-Embedding-4B model generated 2,560-dimensional vector embeddings for more than 563,000 messages in roughly five hours, averaging 107,400 messages per hour with a peak of 35,224 in a single fifteen-minute bin. Stored in Qdrant and linked to their source messages, these embeddings allow researchers to retrieve content matching natural-language descriptions, such as credential theft or illicit marketplace activity, rather than relying solely on dates or keyword matches. Combined with statistical anomaly detection, which flagged two high-activity windows in the DarkGram collection, including one week in late 2025 with 11,005 messages and a z-score of 8.17, the system turns raw monitoring into a queryable analytical environment.

The implications reach well beyond cybersecurity. The authors position Atenea as a flexible backend for computational social science, journalism, and cyber threat intelligence, supporting workflows from network analysis to narrative detection, social listening during protests and crises, and the automated recovery of indicators such as IP addresses, domains, URLs, wallet addresses, and cryptographic hashes. Because the platform is released under the MIT license, with full documentation, a reproducible deployment capsule, and modular microservices that can be extended without touching the core, it lowers the barrier for any research group to build a long-running, auditable observation post on one of the world’s most consequential yet least-studied communication platforms. In an era when online communities can disappear in an instant, tools that preserve both the messages and their meaning may prove as valuable as the analyses they enable.

Subject of Research: An open-source platform for continuous Telegram monitoring, NLP enrichment, and semantic retrieval

Article Title: Atenea: An open-source platform for continuous telegram monitoring, NLP enrichment, and semantic retrieval

Article References: de Paz, A., & Arroyo, D. (2026). Atenea: An open-source platform for continuous telegram monitoring, NLP enrichment, and semantic retrieval. SoftwareX, 36, Article 103067. https://doi.org/10.1016/j.softx.2026.103067

Image Credits: AI Generated

DOI: 10.1016/j.softx.2026.103067

Keywords: Telegram, open-source software, natural language processing, semantic search, cyber threat intelligence, named entity recognition, computational social science, disinformation, Crime-as-a-Service, vector embeddings, data collection, Digital Services Act

Cite Scienmag News

Denise Maddox. (October 1, 2026). Open-Source Platform Atenea Turns Telegram Into a Searchable Research Archive. Scienmag. https://scienmag.com/open-source-platform-atenea-turns-telegram-into-a-searchable-research-archive/

Denise Maddox. "Open-Source Platform Atenea Turns Telegram Into a Searchable Research Archive." Scienmag, 1 October 2026, https://scienmag.com/open-source-platform-atenea-turns-telegram-into-a-searchable-research-archive/. Accessed 1 October 2026.

Denise Maddox. "Open-Source Platform Atenea Turns Telegram Into a Searchable Research Archive." Scienmag. October 1, 2026. https://scienmag.com/open-source-platform-atenea-turns-telegram-into-a-searchable-research-archive/

Tags: challenges in Telegram community trackingcombating disinformation through Telegram archivescomputational social sciencecontinuous monitoring of Telegram channelsCrime-as-a-Servicecyber threat intelligencedata collectionDigital Services Actdisinformationlegal and ethical considerations in Telegram researchnamed entity recognitionnatural language processingnatural language processing for Telegram analysisopen-source platform for Telegram data collectionopen-source softwarepolitical discourse and protest movement analysis on Telegramreproducible research with Telegram dataself-hosted Telegram data analysis workflowssemantic searchsemantic search in Telegram researchTelegramTelegram research archivetracking cybercriminal ecosystems on Telegramvector embeddings
Share26Tweet16
Previous Post

New Scoring Algorithm Turns Routine Surgical Ratings Into Fairer Resident Entrustability Scores

Next Post

Proteomic Aging Clocks Expose the Tangled Truth Behind Alcohol’s Healthy Drinking Paradox

Related Posts

Stress Test for AI Fairness Reveals Which Algorithms Break First Under Biased Labels
Technology and Engineering

Stress Test for AI Fairness Reveals Which Algorithms Break First Under Biased Labels

October 1, 2026
Self-Supervised Neighborhood Probing Shields Tabular AI From Poisoned Labels
Technology and Engineering

Self-Supervised Neighborhood Probing Shields Tabular AI From Poisoned Labels

October 1, 2026
Quantum Shield for the Cloud: Hybrid Cryptography Takes Aim at DDoS Attacks
Technology and Engineering

Quantum Shield for the Cloud: Hybrid Cryptography Takes Aim at DDoS Attacks

October 1, 2026
Frictional Interfaces Emit Strange Non-Local Waves That Defy Classical Rupture Mechanics
Technology and Engineering

Frictional Interfaces Emit Strange Non-Local Waves That Defy Classical Rupture Mechanics

October 1, 2026
Trustworthy AI Has a Toolkit Problem, Landmark Analysis of 938 Tools Reveals
Technology and Engineering

Trustworthy AI Has a Toolkit Problem, Landmark Analysis of 938 Tools Reveals

October 1, 2026
Volcanic Rock Fibers Are Poised to Reshape the Future of Thermoplastic Composites
Technology and Engineering

Volcanic Rock Fibers Are Poised to Reshape the Future of Thermoplastic Composites

October 1, 2026
Next Post
Proteomic Aging Clocks Expose the Tangled Truth Behind Alcohol’s Healthy Drinking Paradox

Proteomic Aging Clocks Expose the Tangled Truth Behind Alcohol's Healthy Drinking Paradox

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Game Theory Reveals When Shale Gas Firms Should Think Long Term
  • Proteomic Aging Clocks Expose the Tangled Truth Behind Alcohol’s Healthy Drinking Paradox
  • Open-Source Platform Atenea Turns Telegram Into a Searchable Research Archive
  • New Scoring Algorithm Turns Routine Surgical Ratings Into Fairer Resident Entrustability Scores

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading