<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>streaming data &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/streaming-data/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 09 Oct 2026 17:22:03 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>streaming data &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New Adaptive Learning Method Keeps AI Models Sharp as Streaming Data Shifts</title>
		<link>https://scienmag.com/new-adaptive-learning-method-keeps-ai-models-sharp-as-streaming-data-shifts/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Fri, 09 Oct 2026 17:22:03 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adaptive learning]]></category>
		<category><![CDATA[adaptive machine learning models]]></category>
		<category><![CDATA[AI model robustness in evolving data]]></category>
		<category><![CDATA[Applied Intelligence]]></category>
		<category><![CDATA[class imbalance]]></category>
		<category><![CDATA[classification]]></category>
		<category><![CDATA[concept drift]]></category>
		<category><![CDATA[concept drift detection]]></category>
		<category><![CDATA[concept drift handling techniques]]></category>
		<category><![CDATA[CUSUM]]></category>
		<category><![CDATA[data streams]]></category>
		<category><![CDATA[drift detection]]></category>
		<category><![CDATA[dynamic learning algorithms]]></category>
		<category><![CDATA[fraud detection in streaming data]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in changing environments]]></category>
		<category><![CDATA[neural network self-reinvention]]></category>
		<category><![CDATA[neural networks]]></category>
		<category><![CDATA[online learning]]></category>
		<category><![CDATA[real-time data pattern shifts]]></category>
		<category><![CDATA[reliable analytical systems for streaming data]]></category>
		<category><![CDATA[sensor data processing]]></category>
		<category><![CDATA[streaming data]]></category>
		<category><![CDATA[streaming data analysis]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=255073</guid>

					<description><![CDATA[Researchers have introduced DACD, a dynamic adaptive learning method that detects concept drift in streaming data with a window-based CUSUM algorithm and retrains a neural network according to the type of drift detected.]]></description>
										<content:encoded><![CDATA[<p>Every second, the world&#8217;s data pipelines deliver torrents of information: sensor readings from factories, clickstreams from websites, transactions from banks, telemetry from vehicles. Machine learning models trained on yesterday&#8217;s data are expected to make sense of all of it today. The trouble is that the statistical patterns those models learned rarely stay still. User preferences shift, equipment ages, fraudsters change tactics, and the very definition of what a model is supposed to predict quietly morphs underneath it. Researchers call this phenomenon concept drift, and it remains one of the most stubborn obstacles to building reliable analytical systems for streaming data. A team of Chinese computer scientists now reports a new approach that tackles the problem head-on, combining a sensitive statistical detector with a neural network that knows how to reinvent itself when the ground shifts.</p>
<p>The method, described in the journal Applied Intelligence, is called DACD, short for Dynamic Adaptive Concept Drift learning. Its authors, Hui Qi, Chang Liu, Xiaobo Qi, Ying Shi, and Gaoxia Jiang, affiliated with Taiyuan Normal University, Shanxi University, and a Shanxi provincial key laboratory, set out to fix two chronic weaknesses of traditional drift-handling algorithms. First, many existing methods perform well only in narrow scenarios, faltering when classification tasks become diverse or when the data is dominated by binary features, the ones-and-zeros that pervade real-world streams. Second, many detectors either cry wolf at harmless fluctuations or sleep through genuine changes, and either mistake degrades the model that depends on them.</p>
<p>At the heart of DACD lies a window-based version of the Cumulative Sum algorithm, a classical statistical technique known as CUSUM that has been used for decades in industrial quality control. The idea behind CUSUM is elegant: rather than asking whether any single data point looks unusual, it accumulates evidence over time, summing deviations from a reference distribution until the accumulated score crosses a threshold. DACD applies this logic across sliding windows of historical data chunks. For each feature in the incoming data, the method computes standardized differences against the statistics of the recent past, accumulates positive and negative CUSUM statistics, and produces a drift score. When a feature&#8217;s score exceeds a detection threshold, that feature is flagged as drifted. Crucially, the method does not require every feature to scream at once: when at least one third of the features show drift scores above the threshold, the system declares that an overall concept drift has occurred. This multi-feature joint judgment makes the detector robust to noise that might trip up a single-feature test.</p>
<p>Detecting drift is only half the battle; the response matters just as much. DACD distinguishes between two fundamentally different kinds of change. Abrupt drift is a sudden, violent shift in the data distribution, concentrated within a short segment of the incoming chunk. Gradual drift is a slow, creeping transformation that unfolds over many samples. The method separates the two using a severity threshold: an event is classified as abrupt only when the intensity of the feature-level change is large and the change is localized within less than a tenth of the current chunk. This distinction is not academic hair-splitting. The two drift types demand opposite remedies, and DACD updates its neural network differently depending on which one it has just witnessed. For abrupt changes, the model pivots quickly toward the newest data; for gradual changes, it leans on carefully curated history to stay stable.</p>
<p>That curation happens through what the authors call an adaptive sample processing strategy. When drift strikes, DACD does not simply discard the past or blindly retrain on everything it has seen. Instead, it filters historical chunks, retaining a limited number of the most relevant ones, then augments the current training set through class-proportional sampling. The augmentation is guided by class priorities computed from the current chunk, so underrepresented categories get a boost and the model does not collapse into predicting only the majority class. All of this happens under temporal constraints, meaning the strategy is designed to keep the pipeline fast enough for genuine streaming deployment. The result, according to the paper, is a model that is more resilient because its training data is filtered and enriched rather than merely replaced.</p>
<p>The neural network itself is trained with a set of engineering choices that will be familiar to deep learning practitioners but are rarely assembled this carefully in a streaming context. DACD selects suitable loss functions and optimizes them with class weights, a technique that penalizes mistakes on rare classes more heavily and helps the model cope with imbalanced streams. Training proceeds in batches, and learning rate schedulers modulate how aggressively the network updates its weights over time, improving stability when the data distribution is in flux. The authors also draw on modern optimization ideas, citing work on decoupled weight decay and the convergence behavior of AdamW, the optimizer that has become a default in much of deep learning. These choices matter because a model that must adapt continuously is especially vulnerable to training instability: a poorly tuned learning rate can either erase useful knowledge in one violent step or crawl too slowly to track a fast-moving distribution.</p>
<p>Performance was evaluated on a battery of benchmark datasets, with simulated streams generated using MOA, the Massive Online Analysis framework that has become a standard testbed for stream classification research. Across these benchmarks, DACD demonstrated strong results in accuracy, convergence speed, and robustness, outperforming comparison methods in diverse classification scenarios. The evaluation followed rigorous statistical practice, using Friedman&#8217;s test and post-hoc analysis, the standard toolkit for comparing classifiers across multiple datasets, in the tradition established by Demšar&#8217;s influential work on statistical comparisons of classifiers. The authors also tested the method against real-world data, with the real-world dataset and core code available from the corresponding authors upon reasonable request.</p>
<p>One of the paper&#8217;s most valuable contributions is a thorough sensitivity analysis of the method&#8217;s parameters, an honesty about tuning that is often missing from algorithmic literature. The size of the data chunk, denoted S, governs the granularity of processing: chunks that are too small invite noise and excessive computation, while chunks that are too large delay the response to change. Experiments showed that a chunk size of 500 delivered the best overall balance, and the authors describe how this size can be dynamically adjusted, shrinking when abrupt drift is suspected to sharpen sensitivity and growing during gradual drift to accumulate enough evidence for a reliable judgment. The detection threshold and drift boundary were tuned across nine combinations, with the best average performance at a threshold of 4.5 and a boundary of 0.3. Intriguingly, the boundary then adapts itself: after an abrupt drift is detected, the next detection window uses a tighter boundary of 0.1, reverting to 0.3 for gradual conditions. A separate severity threshold of 3.0 cleanly separates abrupt from gradual events, chosen because smaller values over-classify slow changes as sudden ones while larger values delay recovery after genuine shocks.</p>
<p>The authors also provide a detailed complexity analysis, which matters enormously for anyone hoping to deploy the method outside a laboratory. The computational bottleneck is the neural network training itself, whose cost grows with the number of features, the hidden layer dimension, and the number of training epochs. The window CUSUM detector adds overhead that scales with window size and feature count, becoming significant in high-dimensional settings. Storage pressure comes from retaining historical chunks and augmented samples, but the authors note practical mitigations: limiting the number of retained chunks, reducing the augmentation ratio, and processing data in GPU-friendly batches. This kind of transparent accounting gives practitioners a realistic picture of what the method will cost to run at scale.</p>
<p>Why does this matter beyond the machine learning community? Concept drift is not a niche concern. Financial institutions monitoring transactions, hospitals tracking patient monitoring streams, energy grids balancing supply and demand, and online platforms personalizing content all face distributions that refuse to hold still. A model that silently degrades can cause real harm long before anyone notices, and retraining from scratch is often too slow or too expensive. Methods like DACD point toward a different paradigm: systems that continuously sense their own obsolescence and repair themselves in real time, distinguishing genuine change from statistical noise and responding with exactly the right amount of upheaval. As streaming data continues to grow in volume and importance, the ability to learn from a world that never stops changing may prove to be one of the defining capabilities of practical artificial intelligence. The work was supported by the National Natural Science Foundation of China and several Shanxi provincial research programs, a reminder that the infrastructure of adaptive intelligence is being built in laboratories around the world, one drift detector at a time.</p>
<p><strong>Subject of Research:</strong> A dynamic adaptive concept drift detection and learning method for classifying streaming data</p>
<p><strong>Article Title:</strong> DACD: a dynamic adaptive concept drift learning method for streaming data</p>
<p><strong>Article References:</strong> Qi, H., Liu, C., Qi, X., Shi, Y., &amp; Jiang, G. (2026). DACD: a dynamic adaptive concept drift learning method for streaming data. <em>Applied Intelligence, 56</em>(15), Article 483. <a href="https://doi.org/10.1007/s10489-026-07520-7" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07520-7</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07520-7" rel="noopener noreferrer">10.1007/s10489-026-07520-7</a></p>
<p><strong>Keywords:</strong> concept drift, streaming data, machine learning, CUSUM, neural networks, adaptive learning, data streams, classification, Applied Intelligence, drift detection, online learning, class imbalance</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">255073</post-id>	</item>
		<item>
		<title>Smart Thresholds, Not Retraining: A Lighter Way to Keep IoT Defenses Sharp</title>
		<link>https://scienmag.com/smart-thresholds-not-retraining-a-lighter-way-to-keep-iot-defenses-sharp/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 08 Oct 2026 01:41:20 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adaptive thresholding]]></category>
		<category><![CDATA[adaptive thresholding in IoT devices]]></category>
		<category><![CDATA[anomaly detection]]></category>
		<category><![CDATA[concept drift]]></category>
		<category><![CDATA[control engineering in cybersecurity]]></category>
		<category><![CDATA[cybersecurity]]></category>
		<category><![CDATA[decision boundary]]></category>
		<category><![CDATA[distribution shift]]></category>
		<category><![CDATA[dynamic decision boundaries for IoT defenses]]></category>
		<category><![CDATA[edge computing]]></category>
		<category><![CDATA[feedback control]]></category>
		<category><![CDATA[feedback loop-based model calibration]]></category>
		<category><![CDATA[handling distribution shift in streaming data]]></category>
		<category><![CDATA[improving IoT device security without retraining]]></category>
		<category><![CDATA[Internet of Things]]></category>
		<category><![CDATA[intrusion detection]]></category>
		<category><![CDATA[IoT security]]></category>
		<category><![CDATA[lightweight cybersecurity solutions for IoT]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning for network anomaly detection]]></category>
		<category><![CDATA[model retraining vs. threshold adaptation]]></category>
		<category><![CDATA[online statistical analysis for intrusion detection]]></category>
		<category><![CDATA[real-time network traffic analysis]]></category>
		<category><![CDATA[streaming data]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=246042</guid>

					<description><![CDATA[Researchers have developed a feedback-guided framework that adapts decision thresholds instead of retraining classifiers, achieving near-equal accuracy with lower false positives, greater stability, and reduced computational cost for streaming IoT intrusion detection.]]></description>
										<content:encoded><![CDATA[<p>Every smart thermostat, industrial sensor, and connected camera in the Internet of Things is a potential doorway for attackers, and the machine-learning models guarding those doorways face a stubborn problem: network traffic never stops changing. A detector trained last month may be quietly mis-calibrated this month as devices join the network, firmware updates shift traffic patterns, and attackers evolve their tactics. The conventional responses—leaving the model frozen or retraining it continuously—each carry a cost. A new study published in Discover Artificial Intelligence proposes a third path: keep the classifier fixed and let the decision boundary itself adapt, using a feedback loop borrowed in spirit from control engineering.</p>
<p>The research, led by Hung-Cuong Nguyen of Hung Vuong University and colleagues at Le Quy Don Technical University and Thai Nguyen University of Information and Communication Technology, addresses what the authors call distribution shift in streaming intrusion detection. Their framework operates on two distinct timescales. At the fast timescale, every incoming network sample is scored using online statistical estimates of feature means and variances, updated through exponential moving averages after each prediction. At the slower timescale, two families of interpretable threshold states—feature-specific abnormality thresholds and a global abnormality gate—are updated only once per non-overlapping batch of 200 samples, using batch-averaged predictive uncertainty and a bounded drift-intensity signal.</p>
<p>The mechanics are deliberately simple. For each of four monitored numerical features, the system computes a Z-score against pre-update statistics, ensuring that a sample never normalizes itself before being judged. A feature is flagged as abnormal when its standardized score exceeds that feature&#8217;s current threshold, and the proportion of flagged features forms an abnormality ratio. If that ratio crosses the global gate threshold, the sample is diverted away from the classifier entirely and treated as suspicious. Everything else flows through a lightweight Random Forest that was trained during an initial warm-up phase and never touched again during streaming operation.</p>
<p>What makes the thresholds move is a feedback law with two opposing forces. Predictive uncertainty, measured as the normalized entropy of the classifier&#8217;s output, pushes thresholds upward: when the model is confused, the gate becomes less sensitive, preventing ambiguous traffic from triggering excessive alarms. Drift intensity, estimated by comparing a recent window of 100 abnormality ratios against a non-overlapping reference window of 300 through a Hoeffding-type tolerance, pushes thresholds downward: when the traffic distribution genuinely departs from its reference state, sensitivity increases. Both signals are clipped to the interval from zero to one, and the resulting threshold updates are themselves clipped to fixed bounds, so the maximum unconstrained one-batch change is 0.025 for a feature threshold and 0.006 for the global gate.</p>
<p>The authors are careful about what this boundedness does and does not prove. The mathematical analysis establishes that threshold states and their one-step changes remain finite, but it explicitly does not establish convergence to an optimal boundary, closed-loop stability in a formal control-theoretic sense, or robustness against an adversary who deliberately manipulates the feedback sequence. In fact, the paper identifies a genuine security limitation: sustained traffic that repeatedly produces high uncertainty without a strong distribution-change signal can drive the detector toward a less sensitive operating point until clipping intervenes. The team treats this desensitization risk as an honest caveat rather than glossing over it.</p>
<p>The experimental evaluation is unusually thorough for this class of work. The framework was tested on three public IoT intrusion-detection benchmarks—RT-IoT2022, IoTID20, and CICIoT2023—with 46,000 streaming observations per dataset processed as 230 batches, repeated across five random seeds for a total of fifteen paired runs per comparison. Against Adaptive Random Forest, the strongest baseline representing continuous model-level adaptation, the proposed method scored an overall F1 of 0.9387 versus 0.9407. That small accuracy gap of roughly 0.21 percentage points was statistically detectable, and the authors decline to claim equivalence. What the fixed-classifier approach bought instead was a lower false-positive rate of 0.0312 versus 0.0349, a 24.5 percent reduction in temporal F1 variance, and a total processing cost of 2.79 milliseconds per sample versus 4.15—a 32.7 percent reduction overall and a 64.7 percent reduction in update time.</p>
<p>Perhaps the most striking result comes from the leave-one-attack-family-out evaluation, where entire attack categories—grouped DDoS for RT-IoT2022, Mirai for IoTID20, and Spoofing for CICIoT2023—were withheld from training entirely. Here the adaptive abnormality gate achieved an F1 of 0.800 with a false-positive rate of 0.031, outperforming fixed statistical baselines such as rolling Z-score rules, median absolute deviation, and interquartile-range detectors. This matters because conventional closed-set classifiers can only assign labels they saw during training, whereas new IoT attack families emerge constantly. The gate reframes the problem: rather than forcing unknown traffic into predefined categories, it simply asks whether a sample deviates from learned statistical patterns, with sensitivity that adjusts as the stream evolves.</p>
<p>A controlled-stream analysis added a stress-test dimension, imposing a predefined distribution shift at batch 120 and tracking how each method responded. The proposed method&#8217;s threshold trajectories revealed smooth, bounded evolution driven by the uncertainty component of the feedback loop, and a criterion-based recovery measure quantified how many batches were needed for F1 to return to at least 95 percent of its pre-shift level. An ablation study disentangled the contributions of the adaptive gate, the adaptive feature thresholds, and the drift signal, showing that the full configuration achieved the best combination of F1, false-positive control, and temporal stability, while no single component dominated every metric. A sensitivity analysis over a grid of feedback coefficients confirmed that the nominal settings represent a stability-oriented operating point rather than a universally optimal choice.</p>
<p>The authors frame their contribution not as a replacement for model adaptation but as a distinct operating point in an accuracy–stability–efficiency trade-off. For resource-constrained edge devices, where GPU acceleration is unavailable and every millisecond counts, a detector that sacrifices a sliver of peak accuracy in exchange for fewer false alarms, smoother behavior over time, and dramatically lower update cost may be exactly the right bargain. The explicit threshold states also offer a form of operational transparency: an operator can inspect how gate sensitivity evolves and, when the drift detector remains silent, know that observed threshold movement is attributable to uncertainty alone. The team acknowledges the limits of this transparency—it is not causal feature explanation—and points toward future work on feature-specific feedback controllers, hybrid schemes that trigger selective model updating under sustained change, and explicit adversarial evaluation. For now, the study offers a compelling demonstration that sometimes the smartest way to adapt a defender is not to rebuild it, but to nudge the line it draws.</p>
<p><strong>Subject of Research:</strong> Feedback-guided decision-boundary adaptation for stable streaming intrusion detection in Internet of Things networks under distribution shift</p>
<p><strong>Article Title:</strong> Feedback guided decision boundary adaptation for stable IoT intrusion detection under distribution shift</p>
<p><strong>Article References:</strong> Nguyen, H.-C., Ta, T. M., Dao, N.-T., Nguyen, Q.-H., &amp; Phung, T.-N. (2026). Feedback guided decision boundary adaptation for stable IoT intrusion detection under distribution shift. <em>Discover Artificial Intelligence, 6</em>(1), Article 1394. <a href="https://doi.org/10.1007/s44163-026-02370-1" rel="noopener noreferrer">https://doi.org/10.1007/s44163-026-02370-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44163-026-02370-1" rel="noopener noreferrer">10.1007/s44163-026-02370-1</a></p>
<p><strong>Keywords:</strong> Internet of Things, intrusion detection, distribution shift, concept drift, adaptive thresholding, decision boundary, feedback control, machine learning, streaming data, anomaly detection, cybersecurity, edge computing</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">246042</post-id>	</item>
		<item>
		<title>New AI Framework Slashes Labeling Costs in Shifting, Imbalanced Data Streams</title>
		<link>https://scienmag.com/new-ai-framework-slashes-labeling-costs-in-shifting-imbalanced-data-streams/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Mon, 05 Oct 2026 08:16:41 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[active learning]]></category>
		<category><![CDATA[adaptive algorithms]]></category>
		<category><![CDATA[adaptive machine learning for dynamic environments]]></category>
		<category><![CDATA[addressing concept drift in AI models]]></category>
		<category><![CDATA[applications of MLIDSC in industry]]></category>
		<category><![CDATA[Applied Intelligence]]></category>
		<category><![CDATA[class imbalance]]></category>
		<category><![CDATA[class imbalance handling in streaming data]]></category>
		<category><![CDATA[concept drift]]></category>
		<category><![CDATA[concept drift detection and adaptation]]></category>
		<category><![CDATA[cost-effective data annotation for real-time ML]]></category>
		<category><![CDATA[data streams]]></category>
		<category><![CDATA[imbalanced data classification in AI systems]]></category>
		<category><![CDATA[labeling strategy]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning data stream labeling]]></category>
		<category><![CDATA[MLIDSC]]></category>
		<category><![CDATA[multiclass classification]]></category>
		<category><![CDATA[online learning]]></category>
		<category><![CDATA[real-time machine learning framework]]></category>
		<category><![CDATA[self-adaptive labeling strategies]]></category>
		<category><![CDATA[streaming data]]></category>
		<category><![CDATA[streaming data anomaly detection]]></category>
		<category><![CDATA[techniques for reducing labeling costs in data streams]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=237324</guid>

					<description><![CDATA[Researchers have developed MLIDSC, a self-adaptive online active learning framework that maintains high accuracy on multiclass imbalanced data streams while sharply reducing expert labeling costs under concept drift.]]></description>
										<content:encoded><![CDATA[<p>Every second, an unrelenting torrent of data flows through the world&#8217;s networks: sensor readings from industrial machinery, packets of internet traffic, transactions from financial systems, and streams of images from cameras watching roads, factories, and hospitals. For machine learning systems tasked with making sense of this flood in real time, two stubborn problems have long conspired to undermine accuracy. The first is concept drift, the phenomenon whereby the statistical patterns underlying the data quietly shift as the world changes. The second is class imbalance, in which some categories of data are overwhelmingly common while others, often the most important ones, appear only rarely. A new framework called MLIDSC, published in Applied Intelligence by researchers at Khulna University of Engineering &amp; Technology and the University of Barishal in Bangladesh, tackles both problems at once with a self-adaptive labeling strategy that promises high accuracy at a fraction of the usual cost.</p>
<p>The labeling problem sits at the heart of why streaming machine learning is so difficult. In classical supervised learning, algorithms are trained on datasets in which every example already carries a label, a process that is expensive and time-consuming but at least finite. In a data stream, new examples arrive continuously and without end, and if every one of them had to be sent to a human expert for annotation, the cost would quickly become prohibitive. Active learning offers a way out: instead of labeling everything, the system itself decides which incoming examples are worth the expense of an expert query, and which can be safely ignored or inferred. The art lies in choosing well. Query too often and the budget evaporates; query too rarely and the model drifts out of date, silently degrading as the world moves on.</p>
<p>Concept drift sharpens this dilemma considerably. When the underlying relationship between features and classes changes, a model trained on yesterday&#8217;s data may be actively misleading today. Drift can arrive suddenly, as when a network intrusion technique appears out of nowhere, or gradually, as customer behavior evolves season by season. Either way, the moments immediately after a drift begins are precisely when fresh labels are most valuable, and precisely when a passive system is least likely to seek them. Existing active learning methods for streams have developed various heuristics for detecting uncertainty and triggering queries, but many of them depend on parameters that must be tuned in advance, assumptions about the data that the designers of MLIDSC argue are rarely justified in real deployments where prior knowledge is scarce.</p>
<p>The second challenge, class imbalance, compounds the difficulty in a way that is particularly insidious in multiclass settings. In a binary problem with a rare class, resampling techniques and cost-sensitive learning have well-understood analogues. But when a stream contains many classes with wildly different frequencies, and when the identity of the rare classes can itself change over time, the problem becomes far more tangled. Rare classes are often the ones that matter most: the fraudulent transaction among millions of legitimate ones, the failing component signal buried in routine telemetry, the emerging disease pattern in a stream of ordinary diagnoses. A labeling strategy that samples uniformly across the stream will spend most of its budget on the abundant classes, leaving the rare ones underrepresented and the model blind to exactly the events it was built to catch.</p>
<p>MLIDSC, whose name abbreviates Multiclass Imbalanced Data Stream with Concept Drift, addresses these intertwined challenges with what the authors describe as a self-adaptive online labeling strategy. The key idea is that the decision to query an expert should be made based on the current context of the stream, rather than on fixed thresholds or pre-set parameters. The framework introduces an automated labeling mechanism that eliminates the need for prior parameter assumptions, allowing the system to calibrate its own behavior against the data as it actually arrives. This matters in practice because streaming deployments are often set up once and left to run for months or years, with little opportunity for the kind of manual retuning that offline machine learning pipelines take for granted.</p>
<p>The most distinctive technical contribution of the framework is a novel weighting scheme designed to prioritize informative minority-class data in nonstationary environments. The scheme combines two signals: the imbalance ratio, which captures how underrepresented a given class is at a given moment, and the importance of individual instances at that point in time. By fusing these two measures, MLIDSC directs its limited labeling budget toward the data points that are simultaneously rare and critical, rather than spreading attention evenly or relying on static class priors. In a stream where the balance of classes shifts as drift proceeds, this dynamic weighting allows the system to notice when a formerly abundant class is fading and a formerly rare one is rising, and to reallocate expert attention accordingly.</p>
<p>Another important design choice distinguishes MLIDSC from much of the prior literature: it processes data online, example by example, rather than in chunks. Many existing methods for drifting and imbalanced streams operate on fixed-size batches, accumulating a buffer of examples before making labeling and model-update decisions. Chunk-based processing simplifies some algorithmic questions, but the authors argue it is impractical for scenarios that demand continuous processing, where waiting for a buffer to fill introduces latency and can delay the detection of abrupt changes. By working in a truly online fashion, MLIDSC can react to each new example as it arrives, which is essential in applications such as network security or industrial monitoring where a delayed response is often as bad as no response at all.</p>
<p>The experimental evaluation behind the framework is notably comprehensive. The researchers tested MLIDSC on both real and synthetic data streams, with synthetic datasets generated using the Scikit-multiflow Python package and real datasets drawn from the Massive Online Analysis repository and the UCI machine learning repository, including the Statlog project data. Crucially, the experiments varied both the degree of concept drift and the imbalance ratio, allowing the team to probe how the method behaves across the full spectrum of difficulty that streaming environments present. The results reported in the paper show that MLIDSC achieves high accuracy while substantially reducing labeling costs compared to existing methods, a combination that is the central promise of active learning and the standard against which such frameworks are judged.</p>
<p>The implications reach well beyond the machine learning research community. Any organization that deploys predictive models on live data faces the twin pressures that MLIDSC targets: expert annotation is expensive, and the world refuses to hold still. Network traffic classification, one of the motivating applications in this line of research, is a vivid example. The mix of applications and protocols traversing a network changes constantly, and traffic classes are wildly imbalanced, with a handful of dominant services dwarfing the rare but security-critical flows that analysts most need to identify. A framework that keeps a classifier accurate in such an environment while asking human experts to label only a small, well-chosen fraction of the stream translates directly into operational savings and faster detection of emerging threats.</p>
<p>The work also fits into a broader and rapidly evolving research conversation. The references anchoring the study trace a decade of progress in the field, from early active learning methods for drifting streams developed by Žliobaitė and colleagues, through ensemble approaches that pair drift detection with resampling strategies, to recent work on multiclass imbalance by Liu, Li, Han and others. The authors of MLIDSC build on their own prior contributions as well, including hybrid labeling strategies and drift detection techniques based on outlier computation. What sets the new framework apart in this crowded landscape is its insistence on self-adaptation: rather than asking practitioners to guess at parameters before deployment, it derives its labeling decisions from the stream itself. As machine learning systems are increasingly entrusted with continuous, high-stakes decisions in environments that never stop changing, that kind of autonomy may prove to be not just a convenience but a necessity, and MLIDSC offers a concrete, experimentally validated template for how to achieve it.</p>
<p><strong>Subject of Research:</strong> Online active learning for multiclass imbalanced data streams with concept drift</p>
<p><strong>Article Title:</strong> MLIDSC: A self-adaptive online active learning framework for multiclass imbalanced data stream with concept drift</p>
<p><strong>Article References:</strong> Halder, B., Hasan, K. M. A., &amp; Ahmed, M. M. (2026). MLIDSC: A self-adaptive online active learning framework for multiclass imbalanced data stream with concept drift. <em>Applied Intelligence, 56</em>(15), Article 431. <a href="https://doi.org/10.1007/s10489-026-07473-x" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07473-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07473-x" rel="noopener noreferrer">10.1007/s10489-026-07473-x</a></p>
<p><strong>Keywords:</strong> machine learning, active learning, concept drift, data streams, class imbalance, online learning, labeling strategy, multiclass classification, adaptive algorithms, streaming data, Applied Intelligence, MLIDSC</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">237324</post-id>	</item>
		<item>
		<title>KL Divergence and Tolerance Time Sharpen Detection of Shifting Data Streams</title>
		<link>https://scienmag.com/kl-divergence-and-tolerance-time-sharpen-detection-of-shifting-data-streams/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sat, 03 Oct 2026 20:00:19 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adaptive drift detection models]]></category>
		<category><![CDATA[ADWIN]]></category>
		<category><![CDATA[beta distribution]]></category>
		<category><![CDATA[concept drift]]></category>
		<category><![CDATA[concept drift detection]]></category>
		<category><![CDATA[data streams]]></category>
		<category><![CDATA[drift detection]]></category>
		<category><![CDATA[evolving fraud detection]]></category>
		<category><![CDATA[information theory]]></category>
		<category><![CDATA[information theory in machine learning]]></category>
		<category><![CDATA[KL divergence]]></category>
		<category><![CDATA[KL divergence in streaming data]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[real-time concept drift monitoring]]></category>
		<category><![CDATA[SABeDM]]></category>
		<category><![CDATA[SABeDM beta distribution]]></category>
		<category><![CDATA[sensor network degradation]]></category>
		<category><![CDATA[SRP]]></category>
		<category><![CDATA[statistical change detection]]></category>
		<category><![CDATA[streaming classifier performance]]></category>
		<category><![CDATA[streaming data]]></category>
		<category><![CDATA[tolerance time]]></category>
		<category><![CDATA[tolerance time in data streams]]></category>
		<category><![CDATA[user preference shifts]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=231726</guid>

					<description><![CDATA[Researchers have enhanced the SABeDM concept drift detector with KL divergence and a tolerance time mechanism, achieving consistently higher accuracy and F1-scores than leading baselines across six streaming benchmarks.]]></description>
										<content:encoded><![CDATA[<p>Machine learning models deployed in the real world rarely enjoy the luxury of a stable environment. Fraud patterns evolve, sensor networks degrade, user preferences shift with the news cycle, and the statistical relationships a model learned yesterday can quietly dissolve overnight. Researchers call this phenomenon concept drift, and detecting it quickly and reliably is one of the central unsolved headaches of streaming data science. A new study published in Knowledge and Information Systems by Wenjun Bian, Angbera Ature, Chan Huah Yong, Shamsuddeen Rabiu and colleagues proposes a fresh answer: an enhanced drift detection framework that fuses a classic tool from information theory, Kullback–Leibler divergence, with a temporal safety valve the authors call tolerance time, all built on top of their earlier Sliding Adaptive Beta Distribution Model, known as SABeDM.</p>
<p>The core idea behind the original SABeDM was to track the behavior of a streaming classifier using a beta distribution, a flexible probabilistic description of error rates that lives naturally between zero and one. As data arrives, the model slides a window across the stream and adapts its distributional estimate of how well the learner is performing. When the picture of performance changes sharply, that is a signal that the underlying data distribution may have changed too. The approach already showed promise in earlier work, but like all window-based detectors it faced a familiar dilemma: windows that are too sensitive drown the system in false alarms caused by ordinary noise, while windows that are too forgiving let genuine drifts slip by unnoticed until accuracy has already collapsed.</p>
<p>The new framework attacks that dilemma from two directions at once. The first is information-theoretic. Rather than relying only on summary statistics of the error stream, the enhanced model, denoted SABeDM with a superscript kl, computes the Kullback–Leibler divergence between the probability distributions estimated from consecutive data windows. KL divergence, a foundational quantity in information theory, measures how many extra bits of information are lost, on average, when one distribution is used to approximate another. In practical terms, it gives the detector a single, principled number that quantifies how different the recent past looks from the present. Small divergences correspond to business as usual; large ones flag a genuine redistribution of the data. Because the measure compares full distributions rather than just means or variances, it can pick up both abrupt ruptures and subtle, gradual shifts that simpler statistics tend to smooth over.</p>
<p>The second innovation is temporal. The authors introduce tolerance time, a grace period that allows the detector to absorb minor delays and short-lived fluctuations without immediately declaring a drift. In a live data stream, brief disturbances are ubiquitous: a batch of noisy measurements, a temporary outage, a transient spike in traffic. A detector that reacts to every one of these burns computational resources and, worse, triggers unnecessary model retraining, which can itself degrade performance. Tolerance time introduces deliberate patience into the system, requiring that evidence of a distributional shift persist for a defined interval before an alarm is raised. The combination is elegant in its symmetry: KL divergence supplies sensitivity to real change, while tolerance time supplies immunity to phantom change.</p>
<p>To find out whether this dual mechanism actually delivers, the team evaluated the framework on six benchmark datasets spanning the standard menagerie of drift scenarios: SEA_a, SEA_g, MIXD, HYP, PHI and WET. These benchmarks are deliberately engineered to stress different failure modes, from sudden abrupt concept changes to gradual and incremental drifts, and the WET dataset adds a real-world, high-dimensional test case where clean theoretical behavior often falls apart. The proposed detector was pitted against a roster of state-of-the-art baselines that includes Streaming Random Patches (SRP), ADWIN, DDM and EDDM, methods that have anchored the drift detection literature for years and, in the case of ADWIN and DDM, remain default choices in many production pipelines.</p>
<p>The results were consistent and, in several cases, striking. On the SEA_a dataset, the enhanced SABeDM achieved an accuracy of 83.50 percent and an F1-score of 82.75 percent, compared with 73.68 percent and 73.87 percent respectively for SRP, the strongest competitor in that comparison. On the more complex PHI dataset, the framework reached 96.10 percent accuracy and a 96.15 percent F1-score, outperforming every baseline tested. Gains extended across precision and recall as well, indicating that the improvements were not an artifact of trading one error type for another but reflected a genuinely better-calibrated detector. The authors also report minimal detection latency, meaning the system tends to notice drift soon after it begins, which is critical because every example processed between the onset of drift and its detection is an example a deployed model may classify wrongly.</p>
<p>Perhaps the most persuasive evidence came from the real-world WET dataset, where the gap between theory-friendly benchmarks and messy practice usually narrows dramatically. There, the enhanced framework lifted the F1-score to 69.00 percent from 60.19 percent for SRP, a gain of nearly nine percentage points in a setting where such margins are rare. The authors attribute this robustness to the interplay of the two new components: KL divergence detects the fine-grained distributional texture of real data, while tolerance time prevents the noise inherent in real streams from overwhelming the detector. The framework also held up in high-dimensional settings, a known weak point for many statistical detectors whose assumptions strain as dimensionality grows.</p>
<p>The significance of the work extends beyond one benchmark table. Concept drift detection sits at the foundation of trustworthy machine learning in dynamic environments, from fraud monitoring and network intrusion detection to predictive maintenance and adaptive recommendation. When drift goes undetected, models fail silently, and the cost is measured in bad decisions rather than error bars. When detection is too twitchy, systems churn through retraining cycles and lose the stability that operators depend on. By grounding the detection decision in a rigorous information-theoretic quantity and pairing it with an explicit model of temporal tolerance, the new framework offers a principled middle path, one that other researchers can extend to related problems such as detecting recurrent concepts, distinguishing real from virtual drift, and handling drift in image and other complex data streams, areas the authors and a growing survey literature identify as open challenges.</p>
<p>There are, of course, caveats worth keeping in mind. The reported evaluations rest on six datasets, and the authors note that no new datasets were generated or analyzed beyond those used in the study, so independent replication on additional industrial streams will be the natural next test. The choice of tolerance time parameters will also matter in practice, since the right grace period likely depends on the application&#8217;s tolerance for delayed detection versus false alarms. Still, the direction is clear and the evidence is strong. As data streams grow faster and models are asked to learn continuously in environments that refuse to stand still, detectors that can both sense the faintest whisper of change and ignore the routine noise of the stream will only become more valuable. This study suggests that an old idea from information theory, applied inside an adaptive probabilistic window with a well-timed pause, may be exactly the combination the field has been waiting for.</p>
<p><strong>Subject of Research:</strong> Concept drift detection in data streams using KL divergence and tolerance time within sliding adaptive beta distribution models</p>
<p><strong>Article Title:</strong> Information-theoretic drift detection using KL divergence and tolerance time in Sliding Adaptive Beta Distribution Models</p>
<p><strong>Article References:</strong> Bian, W., Ature, A., Yong, C. H., &amp; Rabiu, S. (2026). Information-theoretic drift detection using KL divergence and tolerance time in Sliding Adaptive Beta Distribution Models. <em>Knowledge and Information Systems, 68</em>(1), Article 273. <a href="https://doi.org/10.1007/s10115-026-02893-0" rel="noopener noreferrer">https://doi.org/10.1007/s10115-026-02893-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10115-026-02893-0" rel="noopener noreferrer">10.1007/s10115-026-02893-0</a></p>
<p><strong>Keywords:</strong> concept drift, KL divergence, SABeDM, data streams, machine learning, beta distribution, tolerance time, drift detection, streaming data, information theory, ADWIN, SRP</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">231726</post-id>	</item>
		<item>
		<title>Adaptive Splitting Trees Boost Data Stream Learning Under Concept Drift</title>
		<link>https://scienmag.com/adaptive-splitting-trees-boost-data-stream-learning-under-concept-drift/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 19:10:18 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adaptive decision trees]]></category>
		<category><![CDATA[adaptive splitting trees]]></category>
		<category><![CDATA[change detection]]></category>
		<category><![CDATA[classification]]></category>
		<category><![CDATA[concept drift]]></category>
		<category><![CDATA[concept drift detection]]></category>
		<category><![CDATA[concept drift handling]]></category>
		<category><![CDATA[data stream classification]]></category>
		<category><![CDATA[data stream mining]]></category>
		<category><![CDATA[decision trees]]></category>
		<category><![CDATA[dynamic model adaptation]]></category>
		<category><![CDATA[ensemble learning]]></category>
		<category><![CDATA[evolving data streams]]></category>
		<category><![CDATA[HAST]]></category>
		<category><![CDATA[Hoeffding Tree]]></category>
		<category><![CDATA[Hoeffding trees]]></category>
		<category><![CDATA[incremental machine learning]]></category>
		<category><![CDATA[LAST]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[online learning algorithms]]></category>
		<category><![CDATA[online machine learning]]></category>
		<category><![CDATA[real-time data mining]]></category>
		<category><![CDATA[streaming data]]></category>
		<category><![CDATA[streaming data analysis]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=201568</guid>

					<description><![CDATA[Researchers have developed Hoeffding Adaptive Splitting Trees that combine periodic and change-driven splitting to improve ensemble classification of data streams under concept drift.]]></description>
										<content:encoded><![CDATA[<p>Every second, the world&#8217;s sensors, financial networks, and online platforms emit torrents of data that never stop arriving. Unlike traditional machine learning, where algorithms study a fixed dataset, streaming systems must learn on the fly, processing each example once and discarding it. A new study published in Data Mining and Knowledge Discovery tackles one of the hardest problems in this setting: how to keep classification models accurate when the very rules governing the data shift underneath them. Researchers Daniel Nowak Assis, Jean Paul Barddal, and Fabrício Enembreck from the Pontifícia Universidade Católica do Paraná have introduced a family of decision trees called Hoeffding Adaptive Splitting Trees, or HASTs, that promise to make the workhorse algorithms of stream mining both more adaptable and more diverse.</p>
<p>The core challenge is known as concept drift. In a classification problem, a drift occurs when the joint probability distribution linking features and labels changes over time, formally expressed as the distribution at time t differing from the distribution at a later time. Drifts can be abrupt, gradual, incremental, or even recurring, where an old pattern resurfaces after a period of absence. A model trained on yesterday&#8217;s behavior can silently degrade, and streaming systems must detect and react to these shifts almost instantly, all while respecting strict constraints on memory, processing speed, and the ability to produce predictions at any moment.</p>
<p>For two decades, the standard tool for building decision trees on streams has been the Hoeffding Tree, introduced by Domingos and Hulten in 2000. Rather than revisiting stored data, it accumulates statistics at each leaf node and periodically attempts a split, using the Hoeffding bound, a probability inequality that guarantees with confidence level delta how close a sample mean is to the true expected value. If the difference between the best and second-best splitting attributes exceeds a calculated threshold, the tree splits. The approach is elegant and memory efficient, but recent research has exposed a weakness: split attempts happen at fixed intervals regardless of what the data is doing, so the tree keeps searching for splits even during long stretches of stability, wasting computation, and may miss the precise moments when accuracy actually deteriorates.</p>
<p>The same research group previously proposed the Local Adaptive Streaming Tree, or LAST, which flips this logic. Instead of splitting on a schedule, LAST attaches change detectors to leaf nodes that continuously monitor either the error rate or the class-distribution purity. When a detector flags a change, the leaf splits, provided a minimal impurity condition is met. As a standalone classifier, LAST outperformed Hoeffding Trees. But when the authors tried to use it as the base learner inside ensembles, the state-of-the-art approach for stream classification, problems emerged. Because incremental trees all start from a single root, early predictions come from nearly identical majority-class or Naive Bayes models, so the change detectors across ensemble members receive highly similar inputs. Detectors then trigger at closely aligned times, producing correlated trees and undermining the diversity that ensembles depend on for accuracy.</p>
<p>There was a second, subtler flaw. Ensemble methods such as online bagging assign each incoming instance to base learners through Poisson sampling, a random process that simulates drawing samples with replacement. When the only source of variation among detectors is this random weighting, split decisions become governed by sampling noise rather than genuine changes in performance or distribution. Combined with LAST&#8217;s very permissive split condition, this can produce poor, suboptimal splits that a monolithic tree would never make.</p>
<p>The new Hoeffding Adaptive Splitting Trees resolve this tension by combining both splitting philosophies. The first variant, HLAST, keeps the periodic Hoeffding-bound split attempts of a classic Hoeffding Tree, which naturally fosters diversity because different ensemble members receive different numbers of instance copies and therefore split at different times, while also retaining change detectors that can trigger an immediate split when performance or purity degrades. The second variant, EFLAST, builds on the Extremely Fast Decision Tree, or EFDT, which compares the best attribute against not splitting at all and includes a mechanism to re-evaluate and replace earlier splits as better options emerge. EFLAST layers the same adaptive, detector-driven splitting on top of this eager framework. Both models use the HDDM_A drift detector, which prior ablation studies identified as the most efficient and accurate option.</p>
<p>To test the idea, the researchers implemented the trees in the Massive Online Analysis framework and plugged them into five leading ensemble algorithms: Leveraging Bagging, Adaptive Random Forests, Streaming Random Patches, the Adaptive Regularized Ensemble, and the Adaptive Random Tree Ensemble, each running one hundred base learners. The evaluation covered thirteen real-world datasets, including electricity pricing, airline delays, weather data, and insect occurrence records, plus twenty-four synthetic streams generated by classic benchmarks such as AGRAWAL, SEA, LED, RBF, and HYPER, which simulate abrupt, gradual, incremental, and recurring drifts. Performance was measured with the prequential test-then-train protocol, and differences were validated with Friedman tests and Wilcoxon post-hoc comparisons.</p>
<p>The results were striking on real-world data. HLAST and its distribution-monitoring variant HLAST_D beat the standard Hoeffding Tree in 75 percent and 63 percent of cases respectively, while the original adaptive trees won only around 40 percent of the time, confirming that pairing the adaptive mechanism with periodic Hoeffding splits is what unlocks the gains. The improvements were largest on multi-class problems such as Outdoor, Rialto, Poker, LADPU, and CoverType, reaching up to sixteen percentage points of F1-Score improvement, because leaves in such problems stay impure longer and adaptive splitting has more opportunity to act. Crucially, the biggest wins appeared on the hardest streams, those lacking temporal autocorrelation, showing the gains reflect genuine concept learning rather than exploitation of easy, repetitive data. On simple synthetic and binary problems, where concepts are learned quickly, the new trees left results essentially unchanged.</p>
<p>The study also revealed that the pairing of tree and ensemble matters. ARTE combined with HLAST_D delivered the strongest and most consistent results across the benchmark, outperforming ARTE with a standard Hoeffding Tree on nearly every real-world dataset, while remaining cheaper in CPU time and peak memory than Streaming Random Patches with Hoeffding Trees. HLAST_D, which monitors class-distribution purity rather than error rate, proved the right choice for SRP, because the error-driven HLAST can let trees built on weak random feature subsets keep growing and bias accuracy-weighted voting. The drift analysis added further nuance: on synthetic streams all ensembles recovered quickly after each drift, but ARTE struggled on sharp-boundary concepts like AGRAWAL, SRP and Leveraging Bagging faltered on feature-dependent SEA concepts, and Leveraging Bagging with plain Hoeffding Trees degraded sharply as incremental RBF drift accelerated.</p>
<p>The work, funded by CAPES and conducted at PUCPR with a collaboration at Sorbonne Université, is fully open access, with source code and raw results publicly available for reproducibility. The authors point toward several future directions, including extending the approach to regression, applying pre-pruning techniques, and designing even more efficient ensembles that vary sampling intensity based on whether instances are misclassified. For a field where models must learn forever from data that never stops changing, Hoeffding Adaptive Splitting Trees offer a compelling recipe: split when it matters, diversify by design, and adapt at the first sign of change.</p>
<p><strong>Subject of Research:</strong> Adaptive decision tree splitting for data stream classification under concept drift with ensemble learning</p>
<p><strong>Article Title:</strong> Hoeffding adaptive splitting trees for data stream classification with concept drift and ensemble learning</p>
<p><strong>Article References:</strong> Nowak Assis, D., Barddal, J. P., &amp; Enembreck, F. (2026). Hoeffding adaptive splitting trees for data stream classification with concept drift and ensemble learning. <em>Data Mining and Knowledge Discovery, 40</em>(6), Article 91. <a href="https://doi.org/10.1007/s10618-026-01255-2" rel="noopener noreferrer">https://doi.org/10.1007/s10618-026-01255-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10618-026-01255-2" rel="noopener noreferrer">10.1007/s10618-026-01255-2</a></p>
<p><strong>Keywords:</strong> data stream mining, concept drift, Hoeffding Tree, ensemble learning, decision trees, online machine learning, change detection, HAST, LAST, classification, streaming data, machine learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">201568</post-id>	</item>
	</channel>
</rss>
