<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>stock market impact of patent networks &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/stock-market-impact-of-patent-networks/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 11 Oct 2026 10:57:15 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>stock market impact of patent networks &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Reads Patents to Map Which Firms Are Technological Rivals</title>
		<link>https://scienmag.com/ai-reads-patents-to-map-which-firms-are-technological-rivals/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 11 Oct 2026 10:57:15 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI in intellectual property analysis]]></category>
		<category><![CDATA[AI-powered patent analysis]]></category>
		<category><![CDATA[analyzing patent abstracts with AI]]></category>
		<category><![CDATA[asset pricing]]></category>
		<category><![CDATA[Chinese listed firms]]></category>
		<category><![CDATA[dense numerical vectors for patents]]></category>
		<category><![CDATA[firm networks]]></category>
		<category><![CDATA[firm-level technological similarity]]></category>
		<category><![CDATA[innovation analytics]]></category>
		<category><![CDATA[innovation economics measurement]]></category>
		<category><![CDATA[K-means clustering]]></category>
		<category><![CDATA[large-scale patent data processing]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[patent classification systems comparison]]></category>
		<category><![CDATA[patent text analysis]]></category>
		<category><![CDATA[patent text analysis using AI]]></category>
		<category><![CDATA[scalable innovation metrics]]></category>
		<category><![CDATA[semantic embeddings]]></category>
		<category><![CDATA[stock market impact of patent networks]]></category>
		<category><![CDATA[stock returns]]></category>
		<category><![CDATA[technological momentum]]></category>
		<category><![CDATA[technological rivalry mapping]]></category>
		<category><![CDATA[technological similarity]]></category>
		<category><![CDATA[transformer models]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=262006</guid>

					<description><![CDATA[Researchers converted patent abstracts into semantic embeddings to measure technological similarity between firms and found the resulting network predicts subsequent stock returns in China.]]></description>
										<content:encoded><![CDATA[<p>Every patent tells a story about what a company can do, and a new study shows that artificial intelligence can read millions of those stories at once to figure out which firms are truly working on the same technology. In a paper published in the International Journal of Data Science and Analytics, Chuan Shi of The Chinese University of Hong Kong, Shenzhen, together with Rui Luo, Shuang Zhao, Qingge Geng and Di Wu, presents a framework that converts the raw text of patent abstracts into dense numerical vectors, aggregates those vectors into firm-level measures of technological similarity, and then uses the resulting network of related firms to study stock market behavior. The work addresses a long-standing bottleneck in innovation economics: how to measure technological proximity between companies in a way that is both rich in meaning and scalable to entire markets.</p>
<p>For decades, researchers have relied on the International Patent Classification system, a hierarchical scheme established by the Strasbourg Agreement of 1971, to describe what a patent covers. The IPC uses language-independent symbols to sort inventions into technology areas, which makes it universal and easy to apply at scale. But the authors argue that this structure is relatively coarse. Two patents can sit in the same IPC class while describing fundamentally different ideas, and two closely related inventions can be filed under different codes. The paper illustrates the problem with examples drawn from the classification itself: code F16M covers frames, casings and beds of engines and machines, while G09F covers displaying, advertising, signs and labels. A firm&#8217;s portfolio of such codes says little about the finer semantic texture of its research program.</p>
<p>The new framework replaces coarse codes with semantic embeddings. The team uses a Chinese general-purpose embedding model from the C-Pack family of resources, developed by researchers associated with the Beijing Academy of Artificial Intelligence and collaborators. BGE is a pre-trained transformer encoder, the same architectural family that underlies modern large language models, and it maps an input text sequence into a fixed-length dense vector. In the study&#8217;s implementation, each patent abstract is encoded into a 1,024-dimensional normalized embedding through the standard model API. The similarity between any two patents is then measured by the cosine similarity of their two vectors, a number close to one when the texts describe similar technology and close to zero when they do not.</p>
<p>Turning patent-level vectors into firm-level measures requires a carefully documented pipeline. The authors first apply K-means clustering to a random sample of 100,000 patent abstracts, grouping semantically similar patents into 500 clusters. Each cluster centroid acts as a prototype of a technology theme. For every firm and month, the researchers then count the patents the firm added to its portfolio over the preceding twelve months, assign each of those patents to its nearest cluster, and build a firm-level distribution vector describing the composition of the firm&#8217;s recent innovation activity. The similarity between two firms is computed from these distribution vectors, producing a large network in which linked firms are those whose recent patenting draws on the same technological themes.</p>
<p>A crucial design question is how many clusters to use, and the paper confronts it with standard clustering diagnostics rather than arbitrary choice. The within-cluster sum of squares declines monotonically as K increases, with diminishing incremental gains, while the Davies-Bouldin index improves gradually without revealing a sharp optimum. Most tellingly, the silhouette coefficient, which measures how well each patent fits its assigned cluster relative to neighboring clusters, remains highly stable across the full range of values considered. This indicates that partition quality is broadly similar over an intermediate-to-high range of granularities, and the authors conclude that K equal to 500 lies comfortably within that region, making it a reasonable level of semantic granularity for constructing firm-level technology profiles.</p>
<p>To test whether the embedding-based measure genuinely captures more than existing methods, the team benchmarks it against two alternatives. One benchmark uses the IPC classification directly, and the other builds a TF-IDF representation, a classical text-mining technique that weights words by how distinctive they are. For the TF-IDF comparison, the authors reuse the same 100,000-abstract sample, tokenize the Chinese text with the Jieba tokenizer, cap the vocabulary at 50,000 features, and apply the same K-means clustering with 500 clusters so that the two representations are directly comparable. The results show that the embedding-based measure captures technological relatedness beyond the IPC benchmark and remains robust across alternative constructions of the network.</p>
<p>Robustness checks extend to the similarity function itself. Besides cosine similarity, the authors construct an alternative measure based on the Hellinger distance, a metric from information theory that compares probability distributions, applied to the firms&#8217; normalized technology-share vectors. The main conclusions are qualitatively unchanged under this alternative, indicating that the findings are not an artifact of one particular way of measuring similarity. The framework also survives stricter network thresholds: tightening the linking criteria reduces the average number of linked firms per firm from 437.4 to 140.8, 97.1 and 70.3 across the specifications tested, while preserving more than 99.9 percent of firm-month observations, showing that the signal does not depend on an unusually dense network.</p>
<p>The empirical payoff comes in the form of a financial application. Using a large sample of Chinese listed firms from March 2015 to December 2025, with financial data from Tushare, patent abstracts from QuantData and additional academic datasets from the China National Research Data Service, the researchers build a technological momentum measure from the similarity network. The idea builds on a well-known line of asset pricing research: information about economically linked firms diffuses slowly across markets, so news that has already moved the stocks of technological peers may predict the returns of firms that have not yet reacted. Portfolio sorts based on the embedding-based momentum measure, orthogonalized and sorted into quintiles each month-end, yield positive and economically meaningful abnormal returns across standard factor models, under both equal- and value-weighting, complementing Fama-MacBeth cross-sectional regression evidence reported in the main text.</p>
<p>The construction of the underlying variables follows established practice for the Chinese market, where trading restrictions distinguish A shares from B and H shares. Market capitalization is computed from total A shares using point-in-time share counts that are adjusted only after announcements, avoiding look-ahead bias. The book-to-market ratio uses trailing twelve-month equity over market capitalization, return on equity is based on trailing net profit scaled by average equity, and volatility, turnover and one-month reversal are computed from rolling twenty-day windows of daily log returns and trading data, resampled to month-end values. Industry classification follows the CITIC Level 1 standard, and the risk-free rate is the one-month Shanghai interbank offered rate. This attention to point-in-time data matters, because any predictive result in finance is only credible if the information used to form portfolios was actually available to investors at the time.</p>
<p>Beyond the specific momentum result, the broader contribution of the paper is methodological. It provides a transparent and scalable recipe for translating patent-level semantic representations into firm-level measures of technological relatedness, a problem that touches innovation analytics, competitive analysis, and any empirical setting where researchers need to know which organizations are working on similar technology. The framework is deliberately modular: the embedding model, the clustering granularity, the similarity function and the network thresholds are all separate design choices that can be examined and swapped, and the paper documents the sensitivity of the results to each one. As transformer-based language models continue to improve in specialized domains, with prior work such as SciBERT for scientific text and PatentSBERTa for patent distance and classification, the approach demonstrated here suggests that the text of innovation itself, read by machines at scale, can become a quantitative lens on the competitive structure of technology industries and the markets that price them.</p>
<p><strong>Subject of Research:</strong> Measuring firm-level technological similarity from patent texts using semantic embeddings and its application to stock return prediction</p>
<p><strong>Article Title:</strong> Measuring firm-level technological similarity from patent texts using semantic embeddings: a data-driven framework and financial market application</p>
<p><strong>Article References:</strong> Shi, C., Luo, R., Zhao, S., Geng, Q., &amp; Wu, D. (2026). Measuring firm-level technological similarity from patent texts using semantic embeddings: a data-driven framework and financial market application. <em>International Journal of Data Science and Analytics, 22</em>(1), Article 339. <a href="https://doi.org/10.1007/s41060-026-01297-1" rel="noopener noreferrer">https://doi.org/10.1007/s41060-026-01297-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s41060-026-01297-1" rel="noopener noreferrer">10.1007/s41060-026-01297-1</a></p>
<p><strong>Keywords:</strong> semantic embeddings, patent text analysis, technological similarity, firm networks, technological momentum, stock returns, innovation analytics, natural language processing, K-means clustering, asset pricing, Chinese listed firms, transformer models</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">262006</post-id>	</item>
	</channel>
</rss>
