<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Bi-LSTM &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/bi-lstm/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 09 Oct 2026 10:57:02 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>Bi-LSTM &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Attention-Based AI Reads News Word by Word to Catch Multilingual Fake Stories</title>
		<link>https://scienmag.com/attention-based-ai-reads-news-word-by-word-to-catch-multilingual-fake-stories/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Fri, 09 Oct 2026 10:57:02 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI accuracy in fake news classification]]></category>
		<category><![CDATA[AI-based multilingual news verification]]></category>
		<category><![CDATA[automated fake news detection in Indian languages]]></category>
		<category><![CDATA[Bi-LSTM]]></category>
		<category><![CDATA[challenges of misinformation in diverse languages]]></category>
		<category><![CDATA[COVID-19 misinformation in regional languages]]></category>
		<category><![CDATA[cross-lingual misinformation identification]]></category>
		<category><![CDATA[data augmentation]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[fake news detection]]></category>
		<category><![CDATA[fake news impact on public health]]></category>
		<category><![CDATA[FastText embeddings]]></category>
		<category><![CDATA[hierarchical attention network]]></category>
		<category><![CDATA[Hierarchical Attention Network for misinformation]]></category>
		<category><![CDATA[Hindi language]]></category>
		<category><![CDATA[low-resource languages]]></category>
		<category><![CDATA[Marathi language]]></category>
		<category><![CDATA[misinformation]]></category>
		<category><![CDATA[multilingual AI news analysis]]></category>
		<category><![CDATA[Multilingual fake news detection]]></category>
		<category><![CDATA[multilingual NLP]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[scalable solutions for multilingual misinformation]]></category>
		<category><![CDATA[social media misinformation in Hindi and Marathi]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=253369</guid>

					<description><![CDATA[Researchers have built a hierarchical attention network that detects fake news in English, Hindi, and Marathi, reaching 98.9 percent accuracy and introducing a new Marathi fake news dataset.]]></description>
										<content:encoded><![CDATA[<p>Fake news does not respect borders, and it certainly does not respect languages. While social media platforms have invested heavily in automated misinformation detection for English content, the vast majority of the world&#8217;s languages remain poorly served by these systems. A new study published in Discover Artificial Intelligence by Sushama Nandgaonkar and Sunil Mane of COEP Technological University in Pune tackles this gap head-on, presenting a Hierarchical Attention Network that detects false news across English, Hindi, and Marathi with remarkable accuracy, achieving 98.9 percent accuracy and a 98.8 percent F1-score on a combined multilingual test set.</p>
<p>The stakes are particularly high in India, which has the second-largest population of internet users worldwide. According to figures cited in the study, India had more than 375 million social media users in 2019, rising to 518 million in 2020, with projections suggesting the number could reach 1.5 billion by 2040. Affordable mobile data has driven this explosive growth, and platforms such as Facebook, Twitter, and WhatsApp now carry content in dozens of languages. Misinformation that circulates in regional languages can provoke local sentiments and influence health decisions, as demonstrated during the COVID-19 pandemic, when false claims about remedies and vaccines spread faster than official corrections could keep up.</p>
<p>The researchers distinguish carefully between related concepts that are often conflated. Disinformation refers to deliberately created and spread falsehoods, rumors are unverified pieces of information circulated with intent to mislead, and misinformation is false content shared without necessarily intending to deceive. The motivations behind such content range from spreading religious hatred to gaining political advantage or disseminating false health information. Because fact-checking websites themselves can carry biases, and because human verification cannot scale to billions of posts, automated detection systems have become an essential line of defense, and the study argues that these systems must work across the languages people actually use.</p>
<p>At the heart of the new work is a two-level attention architecture that mirrors how humans read documents. The model first processes each sentence word by word using a Bidirectional Long Short-Term Memory network, which reads text in both forward and backward directions to capture context from either side of every word. A custom attention layer then assigns dynamic weights to individual words, learning during training which terms matter most for distinguishing fake from genuine content. The weighted word representations are combined into sentence vectors, and a second Bi-LSTM layer with its own attention mechanism models relationships between sentences, producing a document-level representation that feeds into a final sigmoid classifier.</p>
<p>This hierarchical design is not merely an engineering convenience. Fake news articles often contain misleading lexical patterns, emotionally polarized expressions, and contextual inconsistencies distributed across multiple sentences rather than concentrated in any single phrase. By attending to both words and sentences, the model can capture local semantic cues and document-level structure simultaneously. The attention weights also improve interpretability, allowing researchers to see which parts of an article drove a classification decision, a significant advantage over opaque black-box approaches.</p>
<p>A crucial contribution of the study is linguistic rather than architectural. Because no publicly available fake news dataset existed for Marathi, the team built one from scratch, collecting 1,957 news articles from Marathi outlets including Loksatta, Lokmat, Maharashtra Times, and Zee News, of which 707 were labeled fake and 1,250 true. Labels were assigned based on source credibility, fact-checking reports, and consistency across multiple platforms, with manual review of every article. For Hindi, the researchers combined two existing datasets from Kaggle and GitHub, while English experiments used the widely adopted ISOT dataset. To address severe class imbalance in the low-resource languages, the team applied translation-based augmentation using the IndicTrans neural translation framework, generating additional training samples while keeping the test set untouched to prevent data leakage. The final augmented corpus contained 45,386 English, 10,481 Hindi, and 5,000 Marathi samples.</p>
<p>The choice of word embeddings proved important. FastText, developed by Facebook&#8217;s AI Research lab, represents each word as a bag of character n-grams, allowing it to capture morphological information and generate vectors for out-of-vocabulary words from their sub-word components. This property is especially valuable for morphologically rich languages like Hindi and Marathi, where word forms vary extensively. The researchers compared FastText against GloVe and random initialization under identical hyperparameters, finding that pretrained embeddings substantially improved performance, with FastText also converging fastest at 383.06 seconds of training time compared with 501.22 seconds for GloVe.</p>
<p>The experimental results were striking. The HAN model achieved 99.42 percent accuracy on English news, 99.04 percent on Hindi, and 80.90 percent on Marathi, outperforming traditional machine learning baselines including Logistic Regression, Linear Support Vector Machines, and Random Forest, as well as deep learning models such as CNN and standalone Bi-LSTM. Against multilingual transformers, the picture was more nuanced: mBERT reached 98.72 percent overall accuracy and DeBERTa-v3-base 98.66 percent, slightly below the proposed model&#8217;s 98.90 percent, but DeBERTa&#8217;s performance collapsed to 74.42 percent on Marathi, exposing the sensitivity of large pretrained transformers to data imbalance. Notably, the instruction-tuned LLaMA 3.1 8B model, evaluated through prompting rather than fine-tuning, managed only 67.98 percent accuracy in zero-shot settings and 84.76 percent with six examples in context, with 1,638 failed predictions out of 10,965 test instances.</p>
<p>Ablation experiments confirmed that both attention levels contribute meaningfully. A Bi-LSTM-only model achieved an F1-score of 97.43 percent, rising to 97.82 percent with word-level attention alone and 98.17 percent with sentence-level attention alone, while the complete hierarchical architecture reached 98.72 percent. Sentence-level attention proved more influential than word-level attention in isolation, suggesting that contextual dependencies across sentences play a decisive role in identifying deceptive content. Paired t-tests across multiple random seeds showed all improvements over baselines were statistically significant, with p-values below 0.01.</p>
<p>The study is candid about its limitations. Marathi&#8217;s error rate of 19.10 percent, compared with just 0.58 percent for English and 0.96 percent for Hindi, reflects the challenge of low-resource detection, and qualitative analysis showed that the worst failures involved fake articles written in a style so similar to legitimate news that the model assigned incorrect predictions with high confidence. The authors note that generalization to other Indic languages remains unexplored and that multimodal signals such as images and metadata are not yet incorporated. Future work will expand the dataset to additional Indian regional languages and integrate textual and visual information using transformer-based Indic language models. For now, the research demonstrates that carefully designed attention mechanisms, paired with sub-word embeddings and thoughtful data augmentation, can bring state-of-the-art misinformation detection to languages that have long been left behind.</p>
<p><strong>Subject of Research:</strong> Multilingual fake news detection using hierarchical attention networks for English, Hindi, and Marathi news articles</p>
<p><strong>Article Title:</strong> Hierarchical attention network for multilingual fake news detection</p>
<p><strong>Article References:</strong> Nandgaonkar, S., &amp; Mane, S. (2026). Hierarchical attention network for multilingual fake news detection. <em>Discover Artificial Intelligence, 6</em>(1), Article 1417. <a href="https://doi.org/10.1007/s44163-026-02425-3" rel="noopener noreferrer">https://doi.org/10.1007/s44163-026-02425-3</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44163-026-02425-3" rel="noopener noreferrer">10.1007/s44163-026-02425-3</a></p>
<p><strong>Keywords:</strong> fake news detection, hierarchical attention network, multilingual NLP, Bi-LSTM, FastText embeddings, Marathi language, Hindi language, misinformation, deep learning, low-resource languages, data augmentation, natural language processing</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">253369</post-id>	</item>
		<item>
		<title>Smarter Features, Not Bigger Models, Crack Earthquake Forecasting in Central Asia</title>
		<link>https://scienmag.com/smarter-features-not-bigger-models-crack-earthquake-forecasting-in-central-asia/</link>
		
		<dc:creator><![CDATA[Violet Maxwell]]></dc:creator>
		<pubDate>Mon, 21 Sep 2026 00:10:06 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI approaches to seismic activity]]></category>
		<category><![CDATA[Bi-LSTM]]></category>
		<category><![CDATA[CatBoost]]></category>
		<category><![CDATA[Central Asia]]></category>
		<category><![CDATA[Central Asia earthquake risk]]></category>
		<category><![CDATA[class imbalance]]></category>
		<category><![CDATA[class imbalance in seismic data]]></category>
		<category><![CDATA[earthquake forecasting]]></category>
		<category><![CDATA[earthquake forecasting frameworks]]></category>
		<category><![CDATA[Earthquake prediction]]></category>
		<category><![CDATA[earthquake prediction accuracy]]></category>
		<category><![CDATA[fault descriptors]]></category>
		<category><![CDATA[geophysical data analysis]]></category>
		<category><![CDATA[gradient boosting]]></category>
		<category><![CDATA[Kazakhstan]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in seismology]]></category>
		<category><![CDATA[neural network applications in earthquakes]]></category>
		<category><![CDATA[Omori decay]]></category>
		<category><![CDATA[PR-AUC]]></category>
		<category><![CDATA[predictive modeling for natural disasters]]></category>
		<category><![CDATA[seismic forecasting]]></category>
		<category><![CDATA[spatio-temporal prediction]]></category>
		<category><![CDATA[statistical evaluation of earthquake models]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=204452</guid>

					<description><![CDATA[A Kazakhstani research team shows that physically structured features and calibrated baselines, not model architecture, drive a tenfold improvement in macro-scale earthquake forecasting across Central Asia.]]></description>
										<content:encoded><![CDATA[<p>Earthquakes are among the most stubborn prediction problems in all of science, and a new study from Kazakhstan suggests that the path forward may lie less in exotic neural architectures and more in how the problem itself is framed. Writing in the Journal of Big Data, a team led by Marat Nurtas of the Ionosphere Institute and the International Information Technology University in Almaty reports a macro-scale earthquake forecasting framework for Central Asia that achieves a roughly tenfold improvement over a naive statistical baseline, reaching a Precision-Recall Area Under the Curve of approximately 0.451 and a Receiver Operating Characteristic Area Under the Curve of about 0.844. The striking twist is that six fundamentally different machine learning architectures, from gradient boosting to recurrent neural networks, all converged on nearly identical performance, pointing to a fundamental predictability ceiling rather than a modeling shortfall.</p>
<p>The research tackles a problem that has long plagued computational seismology: extreme class imbalance. When earthquake forecasting is cast as a grid-based classification task, the region under study is divided into spatial cells, and the model must predict whether a seismic event will occur in each cell during each time window. For Central Asia, the team discretized the territory into one-degree by one-degree grid cells and attempted to forecast earthquakes of magnitude 3.0 or greater on a weekly basis. Under this formulation, roughly 96 percent of all cell-week combinations contain no event at all, a phenomenon known as zero inflation. The true event prevalence sits at only about 4.5 percent, which means that a model doing nothing more than predicting &#8216;no earthquake&#8217; everywhere would still appear superficially accurate while being scientifically useless.</p>
<p>This imbalance has profound consequences for how forecasting models must be evaluated. Standard accuracy metrics become meaningless when negative cases dominate by more than twenty to one. The researchers therefore anchored their evaluation in the Precision-Recall Area Under the Curve, a metric that is far more sensitive to performance on the rare positive class. With a prevalence of 4.5 percent, the constant baseline for PR-AUC is 0.045, meaning any model must substantially exceed that value to demonstrate genuine predictive skill. The achieved score of 0.451 represents a tenfold improvement over this baseline, a substantial margin in a domain where even modest gains above chance are considered meaningful by the seismological community.</p>
<p>Central to the study is a carefully engineered feature space grounded in earthquake physics rather than raw statistical patterns. The framework integrates tectonic regime-conditioned normalization, which allows the model to account for the fact that different tectonic settings produce fundamentally different seismic behavior, so that features extracted from a thrust-fault environment are not treated as directly comparable to those from a strike-slip regime. It also incorporates Omori energy decay proxies, mathematical representations of the well-documented tendency of earthquake sequences to produce aftershocks whose frequency decays over time following a mainshock. These proxies give the models a physically interpretable signal about the temporal clustering of seismicity, encoding decades of seismological understanding directly into the input data.</p>
<p>Structural fault descriptors form a third pillar of the feature design. The geometry, orientation, and proximity of mapped fault systems are among the strongest known controls on where earthquakes occur, and by encoding these structural characteristics as model inputs, the framework ensures that the learning algorithms operate on geologically meaningful quantities rather than arbitrary grid statistics. The final and perhaps most consequential innovation is log-odds baseline initialization, a technique that encodes the historical cell-specific event rate directly into the learning objective. Instead of forcing each model to rediscover from scratch the simple fact that some grid cells are historically far more seismically active than others, the initialization embeds this prior knowledge into the model&#8217;s starting point, allowing learning effort to focus on deviations from the historical pattern.</p>
<p>To determine whether performance under such extreme imbalance is governed primarily by model architecture or by structured feature design, the researchers evaluated six heterogeneous architectures under a strict chronological split, ensuring that models were trained only on past data and tested on future periods, exactly as an operational forecasting system would be deployed. The architectures spanned a wide methodological range, including CatBoost and other gradient boosting methods, which excel at tabular data, and Bi-LSTM networks, a bidirectional long short-term memory architecture capable of capturing temporal dependencies in sequential data. Despite their radically different inductive biases and internal mechanics, the models converged on nearly identical PR-AUC and ROC-AUC values, a result the authors interpret as evidence that the information content of the feature space, not the capacity of the learner, is the binding constraint.</p>
<p>Equally notable is what the framework does not do. Many studies confronting severe class imbalance resort to synthetic resampling techniques, such as oversampling the rare event class or undersampling the dominant negative class, to artificially balance the training distribution. These methods can distort the learned probability calibration, producing models whose confidence scores no longer correspond to real-world event likelihoods. The Central Asia framework achieves its tenfold improvement entirely without synthetic resampling, preserving the integrity of the probability estimates. This matters enormously for practical applications, because emergency management authorities require calibrated forecasts whose stated probabilities can be trusted when weighing evacuation decisions, infrastructure inspections, and public warnings.</p>
<p>The authors argue that their findings point to the existence of a macro-scale predictability ceiling in seismic forecasting. If architecturally diverse models, given the same physically structured inputs, all plateau at the same performance level, the implication is that the remaining unpredictability reflects genuine stochasticity in the earthquake process at this spatial and temporal resolution, rather than a deficiency of current algorithms. This interpretation carries a sobering but valuable message for the field: further architectural innovation alone is unlikely to break through the ceiling, while improvements in physical understanding, richer observational data streams, and better-calibrated baselines may still push the boundary outward. It also cautions against the common practice of claiming architectural superiority from small performance differences that may fall within the noise of a shared predictability limit.</p>
<p>The work was funded by the Committee of Science of the Ministry of Science and Higher Education of the Republic of Kazakhstan under a grant for developing a multifunctional system of ground-space monitoring and early warning of natural and technogenic emergencies, underscoring its operational motivation. For a country situated in one of the most seismically active zones of Central Asia, where the collision of the Indian and Eurasian plates drives hazardous tectonics through the Tien Shan and surrounding mountain belts, reliable macro-scale forecasting is not an academic curiosity but a matter of public safety. By demonstrating that disciplined feature engineering, physically informed priors, and rigorous baseline calibration can deliver a tenfold gain in predictive skill without exotic machinery, the Almaty team has provided both a practical forecasting tool and a methodological lesson that resonates far beyond seismology: in data-starved, imbalance-dominated problems, how you frame the question often matters more than how elaborate your model is.</p>
<p><strong>Subject of Research:</strong> Machine learning earthquake forecasting under extreme class imbalance in Central Asia</p>
<p><strong>Article Title:</strong> Macro-scale earthquake forecasting under class imbalance in Central Asia</p>
<p><strong>Article References:</strong> Nurtas, M., Nurakynov, S., Sakabekov, A., Altaibek, A., Kumarkhanova, A., &amp; Merekeyev, A. (2026). Macro-scale earthquake forecasting under class imbalance in Central Asia. <em>Journal of Big Data</em>. <a href="https://doi.org/10.1186/s40537-026-01544-z" rel="noopener noreferrer">https://doi.org/10.1186/s40537-026-01544-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s40537-026-01544-z" rel="noopener noreferrer">10.1186/s40537-026-01544-z</a></p>
<p><strong>Keywords:</strong> earthquake forecasting, Central Asia, class imbalance, machine learning, CatBoost, gradient boosting, Bi-LSTM, PR-AUC, spatio-temporal prediction, Omori decay, fault descriptors, Kazakhstan</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">204452</post-id>	</item>
	</channel>
</rss>
