<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>blockchain data analysis &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/blockchain-data-analysis/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Mon, 05 Oct 2026 01:08:28 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>blockchain data analysis &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Self-Taught AI Reads Blockchain Money Trails to Catch Crypto Laundering</title>
		<link>https://scienmag.com/self-taught-ai-reads-blockchain-money-trails-to-catch-crypto-laundering/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Mon, 05 Oct 2026 01:08:28 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI for financial crime detection]]></category>
		<category><![CDATA[anomaly detection]]></category>
		<category><![CDATA[anti-money laundering in cryptocurrency]]></category>
		<category><![CDATA[Bitcoin]]></category>
		<category><![CDATA[blockchain]]></category>
		<category><![CDATA[blockchain data analysis]]></category>
		<category><![CDATA[blockchain transaction analysis]]></category>
		<category><![CDATA[crypto transaction tracing]]></category>
		<category><![CDATA[cryptocurrency]]></category>
		<category><![CDATA[cryptocurrency money laundering detection]]></category>
		<category><![CDATA[decentralized finance security]]></category>
		<category><![CDATA[financial crime]]></category>
		<category><![CDATA[Graph Neural Networks]]></category>
		<category><![CDATA[graph-structured Transformer models]]></category>
		<category><![CDATA[illicit transaction identification]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning for crypto crime]]></category>
		<category><![CDATA[money laundering]]></category>
		<category><![CDATA[pseudo-labels]]></category>
		<category><![CDATA[self-supervised learning]]></category>
		<category><![CDATA[self-supervised learning in finance]]></category>
		<category><![CDATA[transaction graphs]]></category>
		<category><![CDATA[Transformer]]></category>
		<category><![CDATA[unlabeled blockchain datasets]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=236322</guid>

					<description><![CDATA[A new graph-structured Transformer model trained with self-supervised learning and community-consensus pseudo-labeling achieves record F1-scores for detecting cryptocurrency money laundering when almost no labeled illicit transactions are available.]]></description>
										<content:encoded><![CDATA[<p>Cryptocurrency has given the world a financial system that moves value across borders in seconds, but it has also given money launderers an environment where anonymity is built into the architecture. A new study published in Discover Artificial Intelligence tackles one of the hardest problems in financial crime detection: how to identify illicit transactions when almost none of them carry labels telling investigators what they are. The research, led by Yong Shang of Henan Judicial Police Vocational College in Zhengzhou, China, introduces a model called GT-SSL, a graph-structured Transformer trained through self-supervised learning, and reports striking results on two of the most widely used benchmark datasets in the field.</p>
<p>The core challenge is deceptively simple to state and brutally hard to solve. In real-world blockchain data, confirmed illicit transactions represent only a tiny fraction of the network. On the Elliptic dataset used in the study, roughly 200,000 Bitcoin transactions include just 2 percent labeled as illicit and 21 percent as licit, with the vast majority unlabeled. Traditional supervised machine learning starves in such conditions, and rule-based systems that once anchored anti-money laundering compliance struggle to keep pace with laundering strategies that mutate constantly. Handcrafted features and manual annotation are expensive, and by the time a rule is written, the criminals have moved on.</p>
<p>Graph neural networks emerged as a promising answer because blockchain transactions naturally form networks: money flows from one transaction to another through the unspent transaction output mechanism, creating directed chains, fan-out patterns, and circular loops that are the fingerprints of laundering. But conventional graph neural networks have their own weaknesses. They can suffer from over-smoothing, in which node representations become indistinguishable after repeated aggregation, and they model long-range dependencies poorly, which matters because laundering schemes often stretch across many hops of transfers. Meanwhile, existing self-supervised methods frequently fail to exploit the graph structure itself, blunting their advantage when labels are scarce.</p>
<p>GT-SSL attacks the problem in three stages. First, raw blockchain records, including transaction hashes, inputs, outputs, timestamps and amounts, are converted into a directed attributed graph in which nodes are transactions and edges represent fund transfers. The model then samples local neighborhoods using a biased restart random walk, generating fixed-length sequences of transactions that serve as input to a Transformer. The walk is deliberately engineered: transition probabilities weigh transaction amount similarity, temporal proximity, edge direction and a degree penalty that prevents massive hub nodes such as exchanges and mixing services from dominating every sampled context. A restart mechanism keeps each sequence anchored near its target transaction, preserving the local laundering path while retaining neighborhood diversity.</p>
<p>The second stage is where the architecture departs most sharply from a standard Transformer. Attention weights are not determined solely by feature similarity in the serialized sequence; they are also constrained by the topology of the underlying fund-flow graph. The study introduces a soft multi-hop structural bias: transactions one hop apart in the graph receive strong attention guidance, two-hop and three-hop neighbors receive attenuated guidance, and distant or unreachable nodes are penalized. This design choice is grounded in how laundering actually works. Layered schemes typically involve multi-level account transfers, fund splitting and cross-node aggregation, so a strict one-hop mask would blind the model to crucial multi-hop paths, while unconstrained global attention drowns it in irrelevant noise. The ablation experiments bear this out: without structural constraints, recall fell to 88.65 percent, whereas the multi-hop soft bias achieved the best results across F1-score, AUC and Matthews correlation coefficient.</p>
<p>Pre-training then proceeds through two complementary self-supervised tasks that share the same encoder. In masked feature reconstruction, 15 percent of node features are replaced with a mask token and the model must reconstruct them from context, forcing it to learn fine-grained transaction attributes. In graph contrastive learning, two augmented views of the graph are generated through edge dropping and feature perturbation, and the model learns to pull the representations of the same node together while pushing different nodes apart, using a temperature-scaled InfoNCE loss with the temperature set to 0.1. The dual-task design proved essential: removing masked reconstruction dropped the F1-score to 93.64 percent, removing contrastive learning dropped it to 93.04 percent, and removing both collapsed it to 81.82 percent, confirming that self-supervised pre-training is the single most important ingredient for learning under label scarcity.</p>
<p>The third stage addresses the pseudo-label problem, the Achilles heel of semi-supervised detection. The model first adopts only predictions whose confidence exceeds a high threshold, then applies a second-stage filter based on graph community structure. Using the Louvain algorithm, the transaction graph is partitioned into tightly connected communities, and a medium-confidence pseudo-label is accepted only if the surrounding community, assumed to be behaviorally homogeneous, provides sufficient consensus support. The thresholds were tuned on validation data and set at 0.90 for confidence and 0.70 for consensus. Across five rounds of self-training, first-stage pseudo-label accuracy stayed above 96.58 percent, second-stage accuracy above 94.62 percent, and cumulative error propagation reached only 3.63 percent, suggesting the mechanism genuinely suppresses the noise amplification that plagues naive self-training.</p>
<p>The headline numbers are impressive. On the Elliptic dataset, GT-SSL achieved an F1-score of 95.80 percent and an AUC of 97.62 percent; on AML-Bitcoin, a larger dataset of roughly 500,000 transactions with about 2,100 labeled laundering cases, it reached an F1-score of 93.78 percent and an AUC of 95.91 percent. It outperformed a broad field of baselines including GCN, GAT, Skip-GCN, EvolveGCN, Inspection-L, GCAF-AML and GNN-GRU, and also beat two post-2023 competitors, Elliptic++-GNN and BERT4ETH-AML, improving F1-score by 2.72 and 2.81 percent respectively while cutting false positives and false negatives. Against recent competitive baselines, the model reduced the average false positive rate to 2.58 percent and the average false negative rate to 7.09 percent, a meaningful margin in a domain where false alarms waste investigator time and missed detections let criminals escape.</p>
<p>Perhaps most striking is the model&#8217;s resilience when labels are nearly absent. With only 5 percent of training labels visible, GT-SSL still achieved 91.23 percent accuracy and 82.15 percent recall, and repeated runs with different random seeds showed stable results, with a paired t-test confirming the improvement over the best baseline was statistically significant at p below 0.01. The model also performed best across four specific laundering categories: ransomware, darknet markets, fraud and Ponzi schemes, reducing false positives for fraud and Ponzi cases to 3.12 and 4.25 percent respectively and cutting the false negative rate for the highly concealed Ponzi category to 11.36 percent. An error analysis showed the remaining failures concentrated in low-frequency Ponzi transactions and small-value multi-hop transfers, cases where risk signals are inherently weak and illicit behavior closely mimics normal activity.</p>
<p>The author is candid about the limits. GT-SSL is a static graph model: timestamps and time-step indices are encoded into node features, and the Elliptic experiments use chronological splits to test temporal generalization, but the graph itself is not dynamically updated, communities are computed once before self-training, and no online distribution-shift adaptation is performed. When evaluated on windows progressively farther from the training period, the F1-score declined from 93.50 to 91.21 percent, evidence of partial but not unlimited temporal robustness. The computational cost is also substantial, driven by structure-aware attention, dual-task pre-training and iterative pseudo-label screening. Future work, the study suggests, lies in temporal graph encoders, dynamic community detection and near-real-time incremental inference. Even so, the framework offers regulators something they rarely get: for each flagged transaction, the model can output the high-attention neighbors, the fund-flow sequence and the community consensus score, providing traceable evidence that could survive compliance review rather than a bare, unexplainable alert.</p>
<p><strong>Subject of Research:</strong> Self-supervised graph Transformer learning for detecting cryptocurrency money laundering in label-scarce blockchain transaction networks</p>
<p><strong>Article Title:</strong> Digital currency money laundering identification model based on graph structure and transformer self-supervised learning</p>
<p><strong>Article References:</strong> Shang, Y. (2026). Digital currency money laundering identification model based on graph structure and transformer self-supervised learning. <em>Discover Artificial Intelligence, 6</em>(1), Article 1288. <a href="https://doi.org/10.1007/s44163-026-02264-2" rel="noopener noreferrer">https://doi.org/10.1007/s44163-026-02264-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44163-026-02264-2" rel="noopener noreferrer">10.1007/s44163-026-02264-2</a></p>
<p><strong>Keywords:</strong> cryptocurrency, money laundering, blockchain, graph neural networks, Transformer, self-supervised learning, pseudo-labels, Bitcoin, financial crime, anomaly detection, machine learning, transaction graphs</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">236322</post-id>	</item>
	</channel>
</rss>
