<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>role of evolutionary algorithms in AI model configuration &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/role-of-evolutionary-algorithms-in-ai-model-configuration/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 01:18:05 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>role of evolutionary algorithms in AI model configuration &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Genetic Algorithms Meet LoRA: New Study Tests Whether Smarter Search Really Beats Simple Tuning</title>
		<link>https://scienmag.com/genetic-algorithms-meet-lora-new-study-tests-whether-smarter-search-really-beats-simple-tuning/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 01:18:05 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[automated hyperparameter search in AI]]></category>
		<category><![CDATA[comparison of LoRA and manual tuning strategies]]></category>
		<category><![CDATA[cost-benefit analysis of model fine-tuning methods]]></category>
		<category><![CDATA[cost-effective fine-tuning of large language models]]></category>
		<category><![CDATA[DistilBERT]]></category>
		<category><![CDATA[efficiency of low-rank updates in neural networks]]></category>
		<category><![CDATA[evaluation]]></category>
		<category><![CDATA[evaluation of search algorithms in AI model adaptation]]></category>
		<category><![CDATA[genetic algorithm]]></category>
		<category><![CDATA[Genetic algorithms for model tuning]]></category>
		<category><![CDATA[hyperparameter optimization]]></category>
		<category><![CDATA[impact of automated search on model performance]]></category>
		<category><![CDATA[inference latency]]></category>
		<category><![CDATA[LoRa]]></category>
		<category><![CDATA[Low-Rank Adaptation (LoRA) in transformer models]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[model sparsity]]></category>
		<category><![CDATA[optimization techniques for pretrained transformers]]></category>
		<category><![CDATA[parameter-efficient fine-tuning]]></category>
		<category><![CDATA[Pareto optimization]]></category>
		<category><![CDATA[role of evolutionary algorithms in AI model configuration]]></category>
		<category><![CDATA[selective unfreezing]]></category>
		<category><![CDATA[text classification]]></category>
		<category><![CDATA[transfer learning]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=220682</guid>

					<description><![CDATA[A matched-budget study of LoRA fine-tuning finds that genetic algorithms and other search optimizers offer no consistent advantage over well-tuned baselines, and that apparent gains often come from extra model capacity rather than smarter search.]]></description>
										<content:encoded><![CDATA[<p>Fine-tuning large language models has become one of the most expensive rituals in modern artificial intelligence. Every time a pretrained transformer must be adapted to a new task, engineers face a choice: retrain the entire network, or find a shortcut that preserves most of the pretrained knowledge while updating only a sliver of the parameters. A new study published in Machine Learning with Applications by Seda Bayat Toksoz and colleagues takes a hard, unusually honest look at one of the most popular shortcuts — Low-Rank Adaptation, or LoRA — and asks a question the field has largely avoided: when automated search algorithms pick the LoRA configuration, are they actually better than a well-tuned baseline, or are they just getting a bigger slice of the evaluation budget?</p>
<p>The appeal of LoRA is easy to explain. Instead of updating a full weight matrix inside a transformer, LoRA freezes the pretrained weights and learns a small low-rank residual update, expressed as the product of two much thinner matrices. For a hidden size of 768, as in DistilBERT, a rank-4 update on a single attention projection adds only 6,144 trainable coefficients. Once training is complete, the residual can be merged back into the original projection, leaving the deployed model essentially identical in shape and speed to the original. But beneath that elegance lies a thicket of decisions: which layers to adapt, which attention projections to target, what rank to use, how to set the scaling factor, dropout, and batch size. Each choice shifts the balance between predictive power and efficiency, and no universal recipe exists.</p>
<p>Recent research has increasingly treated these decisions as an optimization problem in their own right. Methods such as AdaLoRA reallocate rank during training, AutoPEFT searches over entire parameter-efficient configurations, AutoLoRA estimates layer-wise ranks through meta-learning, and BIPEFT separates structural selection from rank-budget allocation. What these studies share is a tendency to compare their search strategies against fixed, sometimes under-tuned baselines. The new work attacks that weakness directly by imposing a matched-budget protocol: every method — a genetic algorithm, the Slime Mould Algorithm, Optuna&#8217;s Bayesian optimization, a Pareto-based genetic algorithm, tuned Vanilla LoRA, and tuned full fine-tuning — received exactly 25 validation-only candidate evaluations per task. After selection, each winning configuration was retrained independently with five random seeds, and held-out test data were touched only after every configuration decision was locked in.</p>
<p>The experimental scope was deliberately heterogeneous. DistilBERT served as the main encoder across three very different text classification tasks: the Emotion benchmark, an imbalanced six-class affective dataset with a 9.4-to-1 ratio between its largest and smallest classes; AG News, a large balanced four-class topic benchmark; and Financial PhraseBank, a small domain-specific three-class sentiment dataset. A secondary check used RoBERTa-Base on the GLUE SST-2 task. This diversity matters because a search strategy that shines on one regime may quietly fail on another, and the authors wanted to know whether any optimizer&#8217;s advantage survived contact with different data conditions.</p>
<p>The headline finding is sobering for fans of clever search. On AG News, where full fine-tuning reached 94.4 percent accuracy, GA-guided selective unfreezing came closest at 94.3 percent, but the tuned Vanilla LoRA baseline itself hit 94.1 percent — and the three direct searched-LoRA variants (Optuna, GA, and SMA) landed at 93.9, 93.7, and 93.4 percent respectively. On Financial PhraseBank, where run-to-run variability was three to five times larger, GA+SU-LoRA posted the highest numerical mean at 86.2 percent versus 86.0 percent for full fine-tuning, a gap the authors explicitly decline to interpret as superiority. Across tasks, the ordering of optimizers shifted, and none consistently beat the tuned references. Configuration search, the authors conclude, is best understood as a mechanism for locating task-specific operating points, not as a way to crown a universally dominant algorithm.</p>
<p>Perhaps the most revealing result came from the component ablations on the Emotion dataset. The study&#8217;s GA+SU-LoRA method combines a genetic-algorithm-selected layer and projection support with selective unfreezing, which makes the corresponding pretrained attention matrices trainable. It achieved 93.6 percent accuracy. But a control condition — selective unfreezing applied without any GA-selected support — reached 93.8 percent. The conclusion is uncomfortable but important: the gain came primarily from exposing additional backbone capacity, not from the intelligence of the search. Similarly, a sparse variant that masked roughly half of the LoRA coefficients scored 92.2 percent, essentially indistinguishable from a sparsification control without GA support and only 0.4 points below the tuned Vanilla LoRA baseline, with an approximate Welch 95 percent confidence interval spanning −0.84 to +0.04 points.</p>
<p>The sparsity result also carries a deployment warning. The masked adapters achieved 51.6 percent mean coefficient sparsity across the LoRA factors, but because the zeros were unstructured, the dense tensor shapes remained unchanged and conventional GPU kernels executed the same matrix operations as before. After merging the adapter updates into the base projections, latency stayed within roughly 1 to 3 percent of the full fine-tuning reference, and peak inference memory was essentially unchanged. Compactness on paper, in other words, does not automatically translate into speed in production. The authors argue that future work should distinguish carefully between dense trainable allocation, effective active coefficients, structural model size, and realized deployment cost — four quantities that are routinely conflated in the parameter-efficient fine-tuning literature.</p>
<p>The study also scrutinized the objective function used to guide search. The original scalar fitness combined validation macro-F1 with a compactness penalty weighted by a coefficient lambda. A sensitivity analysis showed this penalty was treacherous: at lambda equal to 0.05, validation accuracy dropped to the 89.8 to 90.8 percent range as the search shrank the layer support to roughly four layers, while removing the penalty entirely recovered 1.3 to 2.0 percentage points. A Pareto-based genetic algorithm that treated accuracy and parameter count as separate objectives exposed the trade-off directly, producing a knee region at 91.9 to 92.5 percent validation accuracy with four to five active layers — a more transparent and robust way to balance performance against compactness than baking the trade-off into a single coefficient whose meaning shifts with dataset scale.</p>
<p>Resource measurements rounded out the deployment picture. The genetic search itself consumed about 2.5 GPU-hours on a single RTX 4070, a modest cost by modern standards. Final retraining runs took between 8 and 15 minutes depending on the method, with peak training memory of 7.4 gigabytes for full fine-tuning, 5.8 gigabytes for the selective-unfreezing variant, and 3.2 gigabytes for the sparse variant. Notably, selective unfreezing raised the trainable ratio to about 11.56 percent of the model — far above conventional LoRA&#8217;s roughly 1 percent — which means it should never be described as adapter-only tuning, even though it delivers accuracy close to full fine-tuning at a fraction of the memory.</p>
<p>The broader lesson of this study may outlast any of its specific numbers. In a field racing to automate every design decision, the authors demonstrate that credible comparisons require matched tuning effort, explicit component-isolating controls, repeated final training across seeds, and deployment-aware reporting. When those safeguards are in place, the mystique of sophisticated search algorithms fades: what remains is a set of practical tools for finding task-specific trade-offs, with the real performance gains coming from controlled access to pretrained capacity rather than from the search itself. For practitioners choosing between optimizers, the message is clear — spend the effort tuning your baseline first, because a well-tuned simple method is a harder target than most papers admit.</p>
<p><strong>Subject of Research:</strong> Parameter-efficient fine-tuning of encoder language models using automated LoRA configuration search and sparse low-rank adaptation</p>
<p><strong>Article Title:</strong> Efficient encoder fine-tuning through configuration search and sparse low-rank adaptation</p>
<p><strong>Article References:</strong> Toksoz, S. B., Isik, G., Turkmen, H., Pacal, I., &amp; Keles, A. (2026). Efficient encoder fine-tuning through configuration search and sparse low-rank adaptation. <em>Machine Learning with Applications, 26</em>, Article 101024. <a href="https://doi.org/10.1016/j.mlwa.2026.101024" rel="noopener noreferrer">https://doi.org/10.1016/j.mlwa.2026.101024</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.mlwa.2026.101024" rel="noopener noreferrer">10.1016/j.mlwa.2026.101024</a></p>
<p><strong>Keywords:</strong> LoRA, parameter-efficient fine-tuning, DistilBERT, genetic algorithm, hyperparameter optimization, model sparsity, selective unfreezing, Pareto optimization, text classification, transfer learning, inference latency, machine learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">220682</post-id>	</item>
	</channel>
</rss>
