<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>classification algorithms &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/classification-algorithms/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 13:23:20 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>classification algorithms &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AutoML Ensemble Predicts Medical Student Performance with Near-Perfect Accuracy</title>
		<link>https://scienmag.com/automl-ensemble-predicts-medical-student-performance-with-near-perfect-accuracy/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 13:23:20 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[AI-driven academic performance prediction]]></category>
		<category><![CDATA[Auto-Weka framework for educational datasets]]></category>
		<category><![CDATA[automated hyper-parameter optimization]]></category>
		<category><![CDATA[AutoML]]></category>
		<category><![CDATA[AutoML in medical education]]></category>
		<category><![CDATA[bagging]]></category>
		<category><![CDATA[classification algorithms]]></category>
		<category><![CDATA[Damietta University]]></category>
		<category><![CDATA[dropout prediction]]></category>
		<category><![CDATA[early warning systems for student attrition]]></category>
		<category><![CDATA[educational data mining]]></category>
		<category><![CDATA[ensemble machine learning models]]></category>
		<category><![CDATA[ensemble methods]]></category>
		<category><![CDATA[high-accuracy student risk classification]]></category>
		<category><![CDATA[learning analytics]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning model selection automation]]></category>
		<category><![CDATA[Medical Education]]></category>
		<category><![CDATA[medical student performance prediction]]></category>
		<category><![CDATA[predictive analytics in healthcare]]></category>
		<category><![CDATA[preventing medical student failure]]></category>
		<category><![CDATA[Random Forest]]></category>
		<category><![CDATA[student performance prediction]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=222970</guid>

					<description><![CDATA[Researchers at Damietta University used automated machine learning to predict medical student performance, finding that ensemble methods like Bagging and Random Forest achieved up to 100 percent accuracy on five years of academic records.]]></description>
										<content:encoded><![CDATA[<p>Every year, universities lose students not because they lack talent, but because nobody spotted the warning signs in time. In medical education, where the stakes are unusually high and the cost of attrition is measured in both money and lost clinicians, the ability to predict which students will struggle before they fail has become a pressing institutional priority. A new study from researchers at Damietta University in Egypt, working with collaborators at Taibah University and Qassim University in Saudi Arabia, shows that automated machine learning, or AutoML, can identify at-risk medical students with startling precision, and that ensemble methods in particular can approach flawless classification on real academic records.</p>
<p>The research, published in the Journal of New Approaches in Educational Research, tackles a problem that has long frustrated educational data scientists: with so many machine learning models available, each with its own tuning requirements, how does an institution without a dedicated data science team find the one that works best for its data? The authors&#8217; answer was to let an algorithm do the searching. Using the Auto-Weka framework, which automates both model selection and hyper-parameter optimization, the team allowed a search procedure to iterate through a long list of predictive strategies and their associated settings until it converged on the configuration that delivered the highest classification accuracy.</p>
<p>The search landed on an ensemble model, a strategy that combines the outputs of several base classifiers rather than relying on any single one. The ensemble evaluated in the study drew on five constituent techniques: artificial neural networks, K-nearest neighbors, Naive Bayes, support vector machines, and logistic regression. Each of these brings a distinct mathematical personality to the task. Neural networks learn layered, nonlinear representations of the data through weighted connections between artificial neurons. K-nearest neighbors classifies a new student by looking at the most similar labeled cases in the training set, using distance functions such as the Euclidean metric. Naive Bayes applies Bayesian probability under the simplifying assumption that features are conditionally independent, which makes it fast to train. Support vector machines construct an optimal separating hyperplane between classes, using kernel functions to handle data that cannot be split linearly. Logistic regression maps a linear combination of inputs through a sigmoid function to produce a probability between zero and one.</p>
<p>In the neural network component used here, the architecture consisted of an input layer representing the categorized data features, two hidden layers containing twelve and seven neurons respectively, and a single output neuron representing the binary outcome. The sigmoid activation function was chosen because it modulates values smoothly between zero and one, making it well suited to probability-style outputs. For the support vector machine, the researchers employed a polynomial kernel of a specified degree, which implicitly casts the data into a higher-dimensional space where, according to Cover&#8217;s theorem, a hyperplane is more likely to separate the two classes cleanly. These technical choices, normally the province of experienced practitioners, were arrived at automatically through the AutoML search rather than by manual trial and error.</p>
<p>The dataset behind the study came from the academic records of students enrolled in a university course at Damietta University between 2016 and 2021. The raw collection contained 480 records, but after preprocessing to remove outliers, missing values, and inconsistencies, 461 usable instances remained. The data was split into a training set of 329 instances, roughly seventy percent, and a test set of 132 instances, roughly thirty percent. Five features described each record: the academic year, the midterm score, the writing exam score, the final degree, and the overall grade. The grades were distributed across five categories from A to F, with D grades, representing a pass, the most common at nearly thirty-eight percent of students, followed by F grades, representing failure, at nearly twenty-nine percent.</p>
<p>The failure statistics buried in those records are striking. In 2016, 70.21 percent of students in the course failed, and in 2017 the rate climbed to a peak of 71.21 percent. The picture improved dramatically in later years, with failure rates of 19.66 percent in 2018, 9.78 percent in 2019, 16.13 percent in 2020, and 17.07 percent in 2021, but the early years illustrate exactly why institutional decision makers want early-warning tools. Only twelve students across the entire five-year span achieved the top A grade, a mere 2.6 percent of the sample, underscoring how demanding the course was and how much room there is for targeted intervention.</p>
<p>When the researchers benchmarked seven classification methods on the test set, the ensemble approaches dominated. Bagging, which trains multiple models on resampled subsets of the data and aggregates their votes, classified all 132 test instances correctly, achieving one hundred percent accuracy along with perfect precision, recall, F-measure, and kappa statistics. Random Forest, a related ensemble built from decision trees, misclassified just one instance, yielding 99.26 percent accuracy and scores of 0.99 across the other metrics. Naive Bayes followed at 95.68 percent accuracy, then the artificial neural network at 91.6 percent, logistic regression at 90.13 percent, K-nearest neighbors at 89.82 percent, and the support vector machine at 79.3 percent. The kappa coefficients, which correct for agreement that could occur by chance, told the same story, ranging from 0.74 for the support vector machine to a perfect 1.0 for Bagging.</p>
<p>The comparison with earlier literature is instructive. Previous studies had reported Random Forest accuracy of 72.4 percent and Naive Bayes accuracy of 88.3 percent on comparable tasks, while K-nearest neighbors had reached 92.6 percent in one survey. The substantially higher figures in the current work suggest that the combination of careful preprocessing, the AutoML-driven selection of hyper-parameters, and the intrinsic strength of ensemble methods can push performance well beyond what individual classifiers typically achieve. The authors attribute the ensemble advantage to the interdependencies among features: because the predictors are correlated, combining multiple base learners that each capture different aspects of those relationships produces a more robust and more accurate overall model than any single technique can manage alone.</p>
<p>The practical implications extend beyond the leaderboard. The researchers frame the predictive model as a decision-support tool for medical sector colleges, one that could inform amendments to admission systems and student selection methods using statistics and grades accumulated over the preceding five years. Identifying weak students early, particularly in the first year when dropout risk peaks, would allow institutions to intervene without lowering educational standards. The study is candid about its limitations, however: the dataset covered a single course at one institution, contained only five features, and relied on records from students who had already begun their studies rather than applicants. The authors note that most published work using methods like Naive Bayes similarly depends on data from enrolled students, which limits how early in the pipeline predictions can be made.</p>
<p>Future work, the team writes, will expand both the number of features and the number of instances in the dataset to enable deeper analysis of educational data, with the goal of better distinguishing struggling students from thriving ones and ultimately reducing failure rates. They also plan to layer optimization techniques such as differential evolution and genetic algorithms onto the predictive framework, potentially squeezing out further gains. For now, the study stands as a compelling demonstration that automated machine learning can compress what used to be a specialist&#8217;s weeks of model tuning into an algorithmic search, and that when it comes to forecasting the academic fate of medical students, the wisdom of many models combined decisively outperforms the judgment of any one.</p>
<p><strong>Subject of Research:</strong> Automated machine learning for predicting academic performance of medical students</p>
<p><strong>Article Title:</strong> Predicting student performance academic using Automated Machine Learning (AutoML): in medical academic institutions</p>
<p><strong>Article References:</strong> Abougalala, R. A., Alharbi, N., Amasha, M. A., Areed, M. F., Alkhalaf, S., &amp; Khairy, D. (2025). Predicting student performance academic using Automated Machine Learning (AutoML): in medical academic institutions. <em>Journal of New Approaches in Educational Research, 14</em>(1), Article 19. <a href="https://doi.org/10.1007/s44322-025-00038-9" rel="noopener noreferrer">https://doi.org/10.1007/s44322-025-00038-9</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44322-025-00038-9" rel="noopener noreferrer">10.1007/s44322-025-00038-9</a></p>
<p><strong>Keywords:</strong> AutoML, machine learning, student performance prediction, medical education, ensemble methods, Random Forest, Bagging, educational data mining, Damietta University, classification algorithms, dropout prediction, learning analytics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">222970</post-id>	</item>
	</channel>
</rss>
