<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>vector embeddings &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/vector-embeddings/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 09:36:54 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>vector embeddings &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Open-Source Platform Atenea Turns Telegram Into a Searchable Research Archive</title>
		<link>https://scienmag.com/open-source-platform-atenea-turns-telegram-into-a-searchable-research-archive/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 09:36:54 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[challenges in Telegram community tracking]]></category>
		<category><![CDATA[combating disinformation through Telegram archives]]></category>
		<category><![CDATA[computational social science]]></category>
		<category><![CDATA[continuous monitoring of Telegram channels]]></category>
		<category><![CDATA[Crime-as-a-Service]]></category>
		<category><![CDATA[cyber threat intelligence]]></category>
		<category><![CDATA[data collection]]></category>
		<category><![CDATA[Digital Services Act]]></category>
		<category><![CDATA[disinformation]]></category>
		<category><![CDATA[legal and ethical considerations in Telegram research]]></category>
		<category><![CDATA[named entity recognition]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[natural language processing for Telegram analysis]]></category>
		<category><![CDATA[open-source platform for Telegram data collection]]></category>
		<category><![CDATA[open-source software]]></category>
		<category><![CDATA[political discourse and protest movement analysis on Telegram]]></category>
		<category><![CDATA[reproducible research with Telegram data]]></category>
		<category><![CDATA[self-hosted Telegram data analysis workflows]]></category>
		<category><![CDATA[semantic search]]></category>
		<category><![CDATA[semantic search in Telegram research]]></category>
		<category><![CDATA[Telegram]]></category>
		<category><![CDATA[Telegram research archive]]></category>
		<category><![CDATA[tracking cybercriminal ecosystems on Telegram]]></category>
		<category><![CDATA[vector embeddings]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=221778</guid>

					<description><![CDATA[Researchers have unveiled Atenea, an open-source platform that continuously archives public Telegram channels, enriches messages with NLP, and enables semantic search across hundreds of millions of messages.]]></description>
										<content:encoded><![CDATA[<p>Telegram has quietly become one of the most important research environments on the internet. With more than one billion monthly active users reported in 2025, the messaging platform hosts public channels and groups that function as open archives of political debate, disinformation, protest movements, and, in darker corners, entire criminal marketplaces. Its open API makes this content technically accessible, yet researchers who try to study it systematically have long faced a frustrating problem: the communities they want to track can vanish overnight, migrate to new channels, or delete their histories, leaving one-off data collections incomplete and irreproducible. A new open-source platform called Atenea, described in the journal SoftwareX by Alfonso de Paz and David Arroyo of the Spanish National Research Council, aims to solve this by turning continuous Telegram monitoring, natural language processing, and semantic search into a single self-hosted workflow.</p>
<p>The motivation goes beyond simple data collection. In cybercriminal ecosystems, low-level identifiers such as channel handles are easily replaced, but higher-level semantic structures, including the language, products, and networks that define a campaign, tend to persist even as the surface infrastructure changes. This is why the authors argue for continuous observation rather than snapshot sampling. Archiving even short-lived communities supports the identification of illegal content under the EU Digital Services Act, enables the extraction of Indicators of Compromise for characterising the Crime-as-a-Service economy, and can ultimately assist the prosecution of cybercriminal activity. At the same time, many researchers currently depend on commercial Software-as-a-Service intermediaries such as TGStat or Telemetrio, which offer basic metrics but introduce opaque filtering, undermine corpus representativeness and reproducibility, and lock data inside proprietary silos without multi-dimensional enrichment.</p>
<p>Existing open-source tools each cover only part of this workflow. pytopicgram focuses on local extraction, preprocessing, sentiment analysis, and topic modelling; 4CAT provides persistent datasets, queued processing, and modular analysis processors; and TeleCatch offers a remote interface for filtered collection and message extraction. A more comprehensive framework called MST has been proposed in the literature, but because no versioned, publicly accessible source repository could be identified, the authors treat it strictly as related work rather than a direct competitor. In a capability comparison verified against public repositories and documentation as of August 2026, Atenea stands out as the only tool of the four that combines continuous monitoring, a persistent server-side message repository, named entity recognition, sentiment analysis, embedding generation, and, critically, true semantic retrieval over stored messages.</p>
<p>Architecturally, Atenea separates collection, processing, storage, retrieval, and exploration behind a common REST interface, with heavy inference decoupled from the core backend so that monitoring continues uninterrupted while services scale or evolve independently. Its data model revolves around three concepts: rooms, which are monitorable Telegram channels or groups; users, whether human accounts or bots; and seeds, external references such as invite links that define which rooms should be ingested. A dedicated Telegram Auth Manager registers and manages multiple API credentials, increasing effective request capacity and handling temporary restrictions. Two load-balancing strategies, one that splits work across credentials and one that groups resources by credential, keep collection resilient, and credential-to-room associations allow access recovery when a particular credential expires.</p>
<p>The processing layer relies on asynchronous, distributed ETL pipelines built from Telethon, Redis, Celery, and Celery Beat, triggered manually or by an automated scheduler. PostgreSQL serves as the transactional system of record, Elasticsearch powers lexical search and exploration through Kibana, and Qdrant stores vector embeddings for semantic retrieval. Downloaded media live outside the relational database in S3-compatible object storage, deployable locally via Garagehq, while PostgreSQL retains their provenance, metadata, hashes, and processing states. Database-level locking and configurable task concurrency prevent worker collisions across pipeline stages. Once collected, messages flow into a service-oriented enrichment layer offering named entity recognition, sentiment analysis, and embedding generation through independent microservices, including custom APIs, vLLM, or Ollama backends.</p>
<p>The enrichment pipeline itself is notably hybrid and language-aware. During ingestion, Atenea cleans message text and assigns an ISO 639-1 language code using the Efficient Language Detector library, discarding detections below a confidence threshold of 0.15. Before any model-based processing runs, regular-expression detectors extract structured indicators, including email addresses, URLs, Telegram mentions, hashtags, and cryptocurrency wallet addresses for Bitcoin, Ethereum, Dash, and Monero, with emoji tokenisation handled separately. These detectors preserve character offsets and operate independently of language. Model-based named entity recognition currently runs for English and Spanish using spaCy 3.8.14 with its large English and Spanish pipelines, mapping model-specific labels onto a canonical taxonomy. Adding a new language requires only installing a compatible spaCy model, registering its code, and redeploying, without changes to the enrichment API, task pipeline, or database schema.</p>
<p>To demonstrate the platform at scale, the authors describe a 2.5-year operational deployment that, as of a July 2026 snapshot, held more than 191 million messages from 6,412 monitored channels and groups, with history reaching back to 2016 and nearly 40 million extracted entities alongside 7.1 million inferred embeddings. A case study built on the DarkGram dataset, a research catalogue of cybercriminal Telegram channels, shows the full workflow in action. From 329 imported seeds, the populate pipeline resolved 207 monitorable rooms. An initial historical-recovery scan stored over 510,000 messages, peaking at 177,357 messages in a single hour, before settling into recurrent four-hour scans averaging roughly 128 new messages per day. The result was a tagged, reproducible collection of 574,432 messages spanning nearly nine years of channel history.</p>
<p>The enrichment results offer a descriptive profile of what automated extraction can reveal. Of the DarkGram messages, 378,905 qualified for English or Spanish model-based processing, and the hybrid extractor produced 175,021 unique entities across the full snapshot. Comparisons against the rest of the repository proved revealing: URL labels appeared at nearly identical rates in both collections, around 27 percent of messages, but PRODUCT labels were overrepresented in the cybercriminal collection at 4.7 percent versus 0.7 percent in the baseline, a pattern consistent with illicit marketplaces. References to people, by contrast, were far rarer in DarkGram, reflecting the baseline repository&#8217;s focus on disinformation and political-news channels. Media cataloguing recorded metadata for 147,646 attachments totalling 6.41 tebibytes of potential content, with opaque .bin files and compressed archives accounting for roughly 6.09 tebibytes, a distribution the authors note is consistent with packaged software and resources circulating in such channels.</p>
<p>Perhaps the most striking demonstration is the platform&#8217;s semantic capability. On external hardware equipped with two NVIDIA L40 S GPUs, a vLLM server running the Qwen3-Embedding-4B model generated 2,560-dimensional vector embeddings for more than 563,000 messages in roughly five hours, averaging 107,400 messages per hour with a peak of 35,224 in a single fifteen-minute bin. Stored in Qdrant and linked to their source messages, these embeddings allow researchers to retrieve content matching natural-language descriptions, such as credential theft or illicit marketplace activity, rather than relying solely on dates or keyword matches. Combined with statistical anomaly detection, which flagged two high-activity windows in the DarkGram collection, including one week in late 2025 with 11,005 messages and a z-score of 8.17, the system turns raw monitoring into a queryable analytical environment.</p>
<p>The implications reach well beyond cybersecurity. The authors position Atenea as a flexible backend for computational social science, journalism, and cyber threat intelligence, supporting workflows from network analysis to narrative detection, social listening during protests and crises, and the automated recovery of indicators such as IP addresses, domains, URLs, wallet addresses, and cryptographic hashes. Because the platform is released under the MIT license, with full documentation, a reproducible deployment capsule, and modular microservices that can be extended without touching the core, it lowers the barrier for any research group to build a long-running, auditable observation post on one of the world&#8217;s most consequential yet least-studied communication platforms. In an era when online communities can disappear in an instant, tools that preserve both the messages and their meaning may prove as valuable as the analyses they enable.</p>
<p><strong>Subject of Research:</strong> An open-source platform for continuous Telegram monitoring, NLP enrichment, and semantic retrieval</p>
<p><strong>Article Title:</strong> Atenea: An open-source platform for continuous telegram monitoring, NLP enrichment, and semantic retrieval</p>
<p><strong>Article References:</strong> de Paz, A., &amp; Arroyo, D. (2026). Atenea: An open-source platform for continuous telegram monitoring, NLP enrichment, and semantic retrieval. <em>SoftwareX, 36</em>, Article 103067. <a href="https://doi.org/10.1016/j.softx.2026.103067" rel="noopener noreferrer">https://doi.org/10.1016/j.softx.2026.103067</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.softx.2026.103067" rel="noopener noreferrer">10.1016/j.softx.2026.103067</a></p>
<p><strong>Keywords:</strong> Telegram, open-source software, natural language processing, semantic search, cyber threat intelligence, named entity recognition, computational social science, disinformation, Crime-as-a-Service, vector embeddings, data collection, Digital Services Act</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">221778</post-id>	</item>
		<item>
		<title>AI Platform Merges Patent Valuation, Market Analysis, and Prior-Art Search</title>
		<link>https://scienmag.com/ai-platform-merges-patent-valuation-market-analysis-and-prior-art-search/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 21:06:11 +0000</pubDate>
				<category><![CDATA[Policy]]></category>
		<category><![CDATA[AI-based patent infringement detection]]></category>
		<category><![CDATA[AI-driven patent analysis platform]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[comprehensive intellectual property evaluation]]></category>
		<category><![CDATA[innovative patent evaluation tools]]></category>
		<category><![CDATA[integrated patent valuation and market analysis]]></category>
		<category><![CDATA[intellectual property]]></category>
		<category><![CDATA[marketability assessment]]></category>
		<category><![CDATA[marketability assessment for patents]]></category>
		<category><![CDATA[patent licensing]]></category>
		<category><![CDATA[patent licensing and licensing potential analysis]]></category>
		<category><![CDATA[patent originality and commercial assessment]]></category>
		<category><![CDATA[patent valuation]]></category>
		<category><![CDATA[prior art]]></category>
		<category><![CDATA[prior-art search automation]]></category>
		<category><![CDATA[Product-Market Fit]]></category>
		<category><![CDATA[semantic patent analysis technology]]></category>
		<category><![CDATA[semantic search]]></category>
		<category><![CDATA[SUNY]]></category>
		<category><![CDATA[technology readiness level 3 patent platform]]></category>
		<category><![CDATA[technology transfer]]></category>
		<category><![CDATA[TRL 3]]></category>
		<category><![CDATA[unified patent workflow system]]></category>
		<category><![CDATA[vector embeddings]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=202500</guid>

					<description><![CDATA[A patent-pending AI system from SUNY combines semantic prior-art search, vector-based patent analysis, and Product-Market Fit scoring to evaluate patent originality and commercial potential in one workflow.]]></description>
										<content:encoded><![CDATA[<p>Evaluating a patent has never been a simple task. Inventors, patent agents, and attorneys must wade through enormous volumes of intellectual property data to determine whether an idea is genuinely novel, whether it infringes on earlier work, and whether anyone will actually want to buy, license, or build upon it. Traditional workflows treat these questions separately, relying on time-consuming manual searches for prior art and disconnected market reports that rarely speak to one another. A new AI-driven system developed within the State University of New York system aims to collapse that fragmented process into a single, integrated workflow, combining semantic patent analysis with market-driven commercial assessment in one platform.</p>
<p>The technology, announced by the Research Foundation for the State University of New York and now available for licensing, is described as a comprehensive solution for evaluating patent originality, marketability, and competitive positioning. It is currently at technology readiness level 3, meaning the core concepts have been demonstrated in principle, and the underlying intellectual property is patent pending. According to the developers, the system was motivated by a persistent gap: first-time inventors in particular struggle to navigate patent evaluation because existing tools are either too technical, too expensive, or too narrowly focused on legal novelty while ignoring economic value.</p>
<p>At the technical heart of the platform is a data processing pipeline that ingests large-scale patent databases and transforms each patent record into a vector embedding, a numerical representation of the document&#8217;s meaning rather than its raw text. This approach, drawn from modern natural language processing, allows the system to perform semantic similarity searches that go far beyond keyword matching. Two patents can use entirely different vocabulary to describe related concepts, and a keyword search would likely miss the connection. Vector embeddings capture conceptual proximity, so a query about a novel battery chemistry, for example, can surface earlier filings that describe comparable electrochemical principles in different language, even across technical domains.</p>
<p>Once the patent corpus has been embedded, a multi-stage retrieval engine takes over. Rather than executing a single search and returning a raw list of results, the engine systematically filters and refines candidate documents in successive passes, narrowing the field to the most relevant prior art and enhancing both the accuracy and the depth of patent comparison. This staged architecture matters because patent databases now contain tens of millions of documents, and naive similarity search at that scale tends to return noise. By layering filtering operations, the system can distinguish between patents that are superficially similar in wording and those that genuinely anticipate or overlap with the invention under review.</p>
<p>The component that most clearly distinguishes this platform from conventional prior-art tools is its Product-Market Fit, or PMF, scoring engine. This module evaluates the commercial viability of a patent by analyzing market signals and trends alongside the technical data extracted from the filing itself. In practice, that means the system does not simply ask whether an invention is new; it asks whether there is evidence of demand, competitive activity, or market movement that suggests the patented technology could be monetized. The output is a structured assessment of economic potential that inventors and licensing professionals can weigh alongside the legal novelty analysis, all within the same interface.</p>
<p>To keep its inputs current and its outputs actionable, the system incorporates third-party application programming interfaces that enrich the data pipeline and allow seamless integration into existing patent evaluation workflows. Firms and university technology transfer offices rarely abandon their established tools outright, so the ability to plug this platform into existing software environments is presented as a deliberate design choice. The developers emphasize scalability as well: the embedding and retrieval pipeline is built to handle industrial volumes of patent records, making the approach viable not just for a single invention review but for portfolio-level analysis across hundreds or thousands of filings.</p>
<p>The intended user base is deliberately broad. For first-time inventors, the platform promises an accessible entry point into a process that has traditionally required either legal counsel or years of experience to navigate. For patent agents and attorneys, it offers a faster, more thorough route to prior-art searches and freedom-to-operate style comparisons. For companies, it provides a way to assess the economic value and competitive positioning of both existing patents and new applications, informing decisions about where to invest research dollars and which assets to license, sell, or abandon. The system&#8217;s designers argue that by simplifying complex patent and market data into actionable insights, it reduces the time and complexity traditionally associated with intellectual property evaluation.</p>
<p>The application space extends across the full life cycle of an invention. The platform can support inventors in judging the originality and market potential of their innovations before committing to the cost of filing. It can assist legal professionals in conducting exhaustive prior-art analyses that are less likely to miss semantically distant but conceptually relevant references. It can help organizations streamline patent portfolio management by flagging assets with strong market alignment and those with weak commercial prospects. It can also serve research and development teams in a more strategic capacity: by mapping where patents cluster and where market signals point, the system can reveal technological gaps and untapped opportunities, guiding future invention rather than merely auditing past filings.</p>
<p>From a broader perspective, the technology reflects a growing trend in which artificial intelligence is applied not to generating inventions but to managing and valuing them. Patent analytics has long been a data-rich but insight-poor field; the raw information exists in abundance, yet translating it into decisions about filing, licensing, and commercialization has remained labor-intensive. By coupling semantic analysis of large patent corpora with market evaluation metrics, this system attempts to bridge the divide between technical patent assessment and market-driven decision-making, addressing what its developers identify as key gaps in traditional evaluation tools.</p>
<p>The technology is now being offered through SUNY&#8217;s technology licensing channels, with the Research Foundation positioning it as part of a broader portfolio of university innovations available for commercialization. Whether the platform achieves adoption will depend on how well its PMF scoring performs against real market outcomes and how gracefully it integrates into the daily routines of patent professionals, but its central premise is clear: the questions of whether an invention is new, whether it can be defended, and whether anyone will pay for it are best answered together, not in isolation. For a field where a single missed prior-art reference or misjudged market can cost years of effort and substantial investment, an integrated, AI-assisted evaluation workflow represents a meaningful shift in how intellectual property decisions may be made.</p>
<p><strong>Subject of Research:</strong> An AI-driven platform for integrated patent valuation, marketability assessment, and prior-art intelligence</p>
<p><strong>Article Title:</strong> AI-driven system and methods for integrated patent valuation, marketability assessment, and prior-art intelligence</p>
<p><strong>Article References:</strong> AI-driven system and methods for integrated patent valuation, marketability assessment, and prior-art intelligence. (n.d.). <a href="https://www.eurekalert.org/news-releases/1144574" rel="noopener noreferrer">Original publication</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> Not provided</p>
<p><strong>Keywords:</strong> artificial intelligence, patent valuation, prior art, vector embeddings, semantic search, Product-Market Fit, intellectual property, patent licensing, marketability assessment, technology transfer, SUNY, TRL 3</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">202500</post-id>	</item>
	</channel>
</rss>
