<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>inference latency &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/inference-latency/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 01:18:05 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>inference latency &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Genetic Algorithms Meet LoRA: New Study Tests Whether Smarter Search Really Beats Simple Tuning</title>
		<link>https://scienmag.com/genetic-algorithms-meet-lora-new-study-tests-whether-smarter-search-really-beats-simple-tuning/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 01:18:05 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[automated hyperparameter search in AI]]></category>
		<category><![CDATA[comparison of LoRA and manual tuning strategies]]></category>
		<category><![CDATA[cost-benefit analysis of model fine-tuning methods]]></category>
		<category><![CDATA[cost-effective fine-tuning of large language models]]></category>
		<category><![CDATA[DistilBERT]]></category>
		<category><![CDATA[efficiency of low-rank updates in neural networks]]></category>
		<category><![CDATA[evaluation]]></category>
		<category><![CDATA[evaluation of search algorithms in AI model adaptation]]></category>
		<category><![CDATA[genetic algorithm]]></category>
		<category><![CDATA[Genetic algorithms for model tuning]]></category>
		<category><![CDATA[hyperparameter optimization]]></category>
		<category><![CDATA[impact of automated search on model performance]]></category>
		<category><![CDATA[inference latency]]></category>
		<category><![CDATA[LoRa]]></category>
		<category><![CDATA[Low-Rank Adaptation (LoRA) in transformer models]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[model sparsity]]></category>
		<category><![CDATA[optimization techniques for pretrained transformers]]></category>
		<category><![CDATA[parameter-efficient fine-tuning]]></category>
		<category><![CDATA[Pareto optimization]]></category>
		<category><![CDATA[role of evolutionary algorithms in AI model configuration]]></category>
		<category><![CDATA[selective unfreezing]]></category>
		<category><![CDATA[text classification]]></category>
		<category><![CDATA[transfer learning]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=220682</guid>

					<description><![CDATA[A matched-budget study of LoRA fine-tuning finds that genetic algorithms and other search optimizers offer no consistent advantage over well-tuned baselines, and that apparent gains often come from extra model capacity rather than smarter search.]]></description>
										<content:encoded><![CDATA[<p>Fine-tuning large language models has become one of the most expensive rituals in modern artificial intelligence. Every time a pretrained transformer must be adapted to a new task, engineers face a choice: retrain the entire network, or find a shortcut that preserves most of the pretrained knowledge while updating only a sliver of the parameters. A new study published in Machine Learning with Applications by Seda Bayat Toksoz and colleagues takes a hard, unusually honest look at one of the most popular shortcuts — Low-Rank Adaptation, or LoRA — and asks a question the field has largely avoided: when automated search algorithms pick the LoRA configuration, are they actually better than a well-tuned baseline, or are they just getting a bigger slice of the evaluation budget?</p>
<p>The appeal of LoRA is easy to explain. Instead of updating a full weight matrix inside a transformer, LoRA freezes the pretrained weights and learns a small low-rank residual update, expressed as the product of two much thinner matrices. For a hidden size of 768, as in DistilBERT, a rank-4 update on a single attention projection adds only 6,144 trainable coefficients. Once training is complete, the residual can be merged back into the original projection, leaving the deployed model essentially identical in shape and speed to the original. But beneath that elegance lies a thicket of decisions: which layers to adapt, which attention projections to target, what rank to use, how to set the scaling factor, dropout, and batch size. Each choice shifts the balance between predictive power and efficiency, and no universal recipe exists.</p>
<p>Recent research has increasingly treated these decisions as an optimization problem in their own right. Methods such as AdaLoRA reallocate rank during training, AutoPEFT searches over entire parameter-efficient configurations, AutoLoRA estimates layer-wise ranks through meta-learning, and BIPEFT separates structural selection from rank-budget allocation. What these studies share is a tendency to compare their search strategies against fixed, sometimes under-tuned baselines. The new work attacks that weakness directly by imposing a matched-budget protocol: every method — a genetic algorithm, the Slime Mould Algorithm, Optuna&#8217;s Bayesian optimization, a Pareto-based genetic algorithm, tuned Vanilla LoRA, and tuned full fine-tuning — received exactly 25 validation-only candidate evaluations per task. After selection, each winning configuration was retrained independently with five random seeds, and held-out test data were touched only after every configuration decision was locked in.</p>
<p>The experimental scope was deliberately heterogeneous. DistilBERT served as the main encoder across three very different text classification tasks: the Emotion benchmark, an imbalanced six-class affective dataset with a 9.4-to-1 ratio between its largest and smallest classes; AG News, a large balanced four-class topic benchmark; and Financial PhraseBank, a small domain-specific three-class sentiment dataset. A secondary check used RoBERTa-Base on the GLUE SST-2 task. This diversity matters because a search strategy that shines on one regime may quietly fail on another, and the authors wanted to know whether any optimizer&#8217;s advantage survived contact with different data conditions.</p>
<p>The headline finding is sobering for fans of clever search. On AG News, where full fine-tuning reached 94.4 percent accuracy, GA-guided selective unfreezing came closest at 94.3 percent, but the tuned Vanilla LoRA baseline itself hit 94.1 percent — and the three direct searched-LoRA variants (Optuna, GA, and SMA) landed at 93.9, 93.7, and 93.4 percent respectively. On Financial PhraseBank, where run-to-run variability was three to five times larger, GA+SU-LoRA posted the highest numerical mean at 86.2 percent versus 86.0 percent for full fine-tuning, a gap the authors explicitly decline to interpret as superiority. Across tasks, the ordering of optimizers shifted, and none consistently beat the tuned references. Configuration search, the authors conclude, is best understood as a mechanism for locating task-specific operating points, not as a way to crown a universally dominant algorithm.</p>
<p>Perhaps the most revealing result came from the component ablations on the Emotion dataset. The study&#8217;s GA+SU-LoRA method combines a genetic-algorithm-selected layer and projection support with selective unfreezing, which makes the corresponding pretrained attention matrices trainable. It achieved 93.6 percent accuracy. But a control condition — selective unfreezing applied without any GA-selected support — reached 93.8 percent. The conclusion is uncomfortable but important: the gain came primarily from exposing additional backbone capacity, not from the intelligence of the search. Similarly, a sparse variant that masked roughly half of the LoRA coefficients scored 92.2 percent, essentially indistinguishable from a sparsification control without GA support and only 0.4 points below the tuned Vanilla LoRA baseline, with an approximate Welch 95 percent confidence interval spanning −0.84 to +0.04 points.</p>
<p>The sparsity result also carries a deployment warning. The masked adapters achieved 51.6 percent mean coefficient sparsity across the LoRA factors, but because the zeros were unstructured, the dense tensor shapes remained unchanged and conventional GPU kernels executed the same matrix operations as before. After merging the adapter updates into the base projections, latency stayed within roughly 1 to 3 percent of the full fine-tuning reference, and peak inference memory was essentially unchanged. Compactness on paper, in other words, does not automatically translate into speed in production. The authors argue that future work should distinguish carefully between dense trainable allocation, effective active coefficients, structural model size, and realized deployment cost — four quantities that are routinely conflated in the parameter-efficient fine-tuning literature.</p>
<p>The study also scrutinized the objective function used to guide search. The original scalar fitness combined validation macro-F1 with a compactness penalty weighted by a coefficient lambda. A sensitivity analysis showed this penalty was treacherous: at lambda equal to 0.05, validation accuracy dropped to the 89.8 to 90.8 percent range as the search shrank the layer support to roughly four layers, while removing the penalty entirely recovered 1.3 to 2.0 percentage points. A Pareto-based genetic algorithm that treated accuracy and parameter count as separate objectives exposed the trade-off directly, producing a knee region at 91.9 to 92.5 percent validation accuracy with four to five active layers — a more transparent and robust way to balance performance against compactness than baking the trade-off into a single coefficient whose meaning shifts with dataset scale.</p>
<p>Resource measurements rounded out the deployment picture. The genetic search itself consumed about 2.5 GPU-hours on a single RTX 4070, a modest cost by modern standards. Final retraining runs took between 8 and 15 minutes depending on the method, with peak training memory of 7.4 gigabytes for full fine-tuning, 5.8 gigabytes for the selective-unfreezing variant, and 3.2 gigabytes for the sparse variant. Notably, selective unfreezing raised the trainable ratio to about 11.56 percent of the model — far above conventional LoRA&#8217;s roughly 1 percent — which means it should never be described as adapter-only tuning, even though it delivers accuracy close to full fine-tuning at a fraction of the memory.</p>
<p>The broader lesson of this study may outlast any of its specific numbers. In a field racing to automate every design decision, the authors demonstrate that credible comparisons require matched tuning effort, explicit component-isolating controls, repeated final training across seeds, and deployment-aware reporting. When those safeguards are in place, the mystique of sophisticated search algorithms fades: what remains is a set of practical tools for finding task-specific trade-offs, with the real performance gains coming from controlled access to pretrained capacity rather than from the search itself. For practitioners choosing between optimizers, the message is clear — spend the effort tuning your baseline first, because a well-tuned simple method is a harder target than most papers admit.</p>
<p><strong>Subject of Research:</strong> Parameter-efficient fine-tuning of encoder language models using automated LoRA configuration search and sparse low-rank adaptation</p>
<p><strong>Article Title:</strong> Efficient encoder fine-tuning through configuration search and sparse low-rank adaptation</p>
<p><strong>Article References:</strong> Toksoz, S. B., Isik, G., Turkmen, H., Pacal, I., &amp; Keles, A. (2026). Efficient encoder fine-tuning through configuration search and sparse low-rank adaptation. <em>Machine Learning with Applications, 26</em>, Article 101024. <a href="https://doi.org/10.1016/j.mlwa.2026.101024" rel="noopener noreferrer">https://doi.org/10.1016/j.mlwa.2026.101024</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.mlwa.2026.101024" rel="noopener noreferrer">10.1016/j.mlwa.2026.101024</a></p>
<p><strong>Keywords:</strong> LoRA, parameter-efficient fine-tuning, DistilBERT, genetic algorithm, hyperparameter optimization, model sparsity, selective unfreezing, Pareto optimization, text classification, transfer learning, inference latency, machine learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">220682</post-id>	</item>
		<item>
		<title>Prompt Robustness and Fine-Tuning Tested in Open-Vocabulary Object Detection Showdown</title>
		<link>https://scienmag.com/prompt-robustness-and-fine-tuning-tested-in-open-vocabulary-object-detection-showdown/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Mon, 21 Sep 2026 01:58:51 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[domain shift]]></category>
		<category><![CDATA[fine-tuning]]></category>
		<category><![CDATA[inference latency]]></category>
		<category><![CDATA[mAP]]></category>
		<category><![CDATA[object recognition]]></category>
		<category><![CDATA[open-vocabulary detection]]></category>
		<category><![CDATA[prompt sensitivity]]></category>
		<category><![CDATA[vision-language models]]></category>
		<category><![CDATA[YOLO-World]]></category>
		<category><![CDATA[YOLOE]]></category>
		<category><![CDATA[zero-shot object detection]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=204992</guid>

					<description><![CDATA[A new comparative study finds that YOLO-World v2 offers the best balance of accuracy, speed, and prompt robustness among open-vocabulary object detectors, while YOLOE loses its open-vocabulary behavior after fine-tuning.]]></description>
										<content:encoded><![CDATA[<p>Open-vocabulary object detection has quietly become one of the most consequential ideas in modern computer vision. Instead of being locked to a fixed list of categories learned during training, these models can recognize objects described in plain language—a text prompt such as &#8220;traffic sign&#8221; or &#8220;kitchen appliance&#8221; is enough to make them find and localize instances they were never explicitly trained on. The promise is enormous: robots that understand novel instructions, surveillance systems that adapt to new threats, and annotation pipelines that label images without human effort. Yet a new systematic study from Selcuk University in Konya, Türkiye, suggests that the field&#8217;s enthusiasm for headline accuracy numbers has obscured a more complicated reality, one in which the choice of words, the cost of inference, and the fate of unseen classes after fine-tuning can matter as much as raw performance.</p>
<p>The research, published in Multimedia Tools and Applications by Melisa Alara Ozuberk and Ilkay Cinar, delivers one of the first head-to-head evaluations of three leading real-time open-vocabulary detectors: YOLO-World, its successor YOLO-World v2, and YOLOE. Rather than benchmarking accuracy alone, the authors designed their experiments around three scenarios that mirror how these systems are actually deployed: zero-shot inference on entirely new datasets, sensitivity to variations in the textual prompts that steer detection, and fine-tuning on domain-specific data followed by tests of whether open-vocabulary generalization survives. The evaluation spans three datasets with deliberately different characteristics—the classic VOC2012 segmentation subset, the HomeObjects-3K indoor detection dataset, and the demanding KITTI autonomous driving benchmark.</p>
<p>The zero-shot results reveal a striking dependence on domain. The highest performance was achieved on HomeObjects-3K, where YOLO-World v2 reached 0.443 mAP@0.5:0.95, a metric that rewards both accurate localization and correct classification across a range of overlap thresholds. KITTI, by contrast, produced the weakest results across all three models, a consequence of domain shift: the driving imagery, with its unusual viewpoints, small distant objects, and harsh lighting conditions, differs substantially from the data distributions these models encountered during pre-training. On the VOC dataset, YOLOE claimed the highest zero-shot accuracy at 0.310 mAP@0.5:0.95, outperforming both YOLO-World variants on that benchmark, although this advantage came with a caveat that emerged clearly in the timing analysis.</p>
<p>That caveat is speed. YOLO-World v2 proved to offer the best overall balance between accuracy and throughput, sustaining between 20 and 30 frames per second—comfortably real-time for many applications. YOLOE, despite its stronger zero-shot accuracy on VOC, exhibited lower inference speed in some configurations, a trade-off that could prove decisive in latency-sensitive settings such as autonomous navigation or live video analytics. The YOLO-World family also benefited from an embedding cache mechanism, which pre-computes text embeddings for the prompt vocabulary and reuses them across frames. Latency analysis with increasing prompt counts showed that this design keeps inference efficient even as the number of textual categories grows, an architectural advantage that becomes more valuable the richer the vocabulary deployed in production.</p>
<p>Perhaps the most practically important finding concerns prompt robustness—the question of how much detection quality degrades when the words fed to the model change. The authors constructed four categories of prompts for each dataset: base prompts, attribute prompts that add descriptive modifiers, longer descriptive prompts, and noisy prompts containing degraded or perturbed language. The results showed measurable performance drops under noisy prompt conditions for all models, confirming that open-vocabulary detectors are not immune to the fragility of language interfaces that has been documented across the broader vision-language literature. However, YOLO-World v2 maintained better stability across these variations than its competitors, suggesting that its training recipe or text-encoding pathway confers a degree of resilience that practitioners should weigh when deploying systems in the hands of non-expert users who cannot be relied upon to craft optimal prompts.</p>
<p>Fine-tuning delivered the expected gains but also exposed an uncomfortable truth about what adaptation costs. After fine-tuning on each dataset, mAP scores improved for all three models, demonstrating that standard transfer learning techniques remain effective when open-vocabulary detectors are specialized to a target domain. But the fine-tuned models showed zero performance on some unseen categories—classes that were never part of the fine-tuning data. This is precisely the failure mode that open-vocabulary detection is supposed to prevent, and its appearance after adaptation indicates that the boundary between open and closed vocabulary is thinner than the field often assumes.</p>
<p>The divergence between the model families was especially pronounced here. The YOLO-World family retained some of its open-vocabulary generalization ability after fine-tuning, continuing to respond to textual prompts for categories outside the training set. YOLOE, under the fine-tuning protocol adopted in the study, exhibited closed-set-like behavior: its predictions remained insensitive to the evaluated prompt variations, effectively behaving as if the text interface had been switched off and the model had reverted to a conventional fixed-category detector. For teams choosing between these architectures, the implication is significant—fine-tuning YOLOE may buy accuracy on known classes at the price of the very flexibility that motivated choosing an open-vocabulary model in the first place.</p>
<p>The study&#8217;s methodology reflects a growing recognition that evaluation practices in this field have been too narrow. The authors note that existing research has focused mainly on accuracy metrics while prompt robustness, unseen class generalization, and computational costs are rarely assessed together. By combining confusion-matrix-based analysis, standard detection metrics such as mAP at multiple intersection-over-union thresholds, and latency profiling under varying prompt counts, the work offers a template for more honest benchmarking. The datasets themselves are all publicly available—VOC2012 from the PASCAL repository, KITTI from the KITTI Vision Benchmark Suite, and HomeObjects-3K through its original repository—making the evaluation pipeline reproducible by other groups.</p>
<p>The broader context makes these findings timely. Open-vocabulary detection builds on a lineage that runs from the original YOLO real-time detector through open-set recognition and open-world detection to caption-supervised methods and the CLIP-style vision-language models that supply the text-image alignment these detectors depend on. YOLO-World, introduced in 2024, brought this capability to real-time speeds, and YOLOE pushed the concept further with its &#8220;see anything&#8221; design. Applications documented in the literature now span automatic image annotation, number plate recognition, wildlife monitoring, medical imaging, underwater fish counting, robotic navigation, and anomaly detection in surveillance—domains where the ability to name new categories without retraining is transformative.</p>
<p>For practitioners, the study&#8217;s bottom line is that model selection should be a multi-dimensional decision. Accuracy, inference speed, prompt robustness, and unseen class generalization form a set of trade-offs that no single model dominates. YOLO-World v2 emerges as the most balanced option, combining competitive accuracy, real-time throughput, an efficient embedding cache, and the strongest prompt stability. YOLOE offers the best zero-shot accuracy in some settings but pays in speed and, critically, appears to surrender its open-vocabulary character when fine-tuned. As these systems move from research demos into safety-relevant deployments—self-driving perception, medical triage, industrial inspection—the lesson of this comparative study is that the questions worth asking about a detector extend well beyond its leaderboard score, reaching into how it behaves when the words change, the domain shifts, and the training data runs out.</p>
<p><strong>Subject of Research:</strong> Comparative evaluation of prompt robustness, fine-tuning, and generalization in open-vocabulary object detection models</p>
<p><strong>Article Title:</strong> Prompt robustness, fine-tuning, and Generalization in open-vocabulary object detection: a comparative study of YOLO-World, YOLO-World v2 and YOLOE</p>
<p><strong>Article References:</strong> Ozuberk, M. A., &amp; Cinar, I. (2026). Prompt robustness, fine-tuning, and Generalization in open-vocabulary object detection: a comparative study of YOLO-World, YOLO-World v2 and YOLOE. <em>Multimedia Tools and Applications, 85</em>(10), Article 767. <a href="https://doi.org/10.1007/s11042-026-21928-w" rel="noopener noreferrer">https://doi.org/10.1007/s11042-026-21928-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11042-026-21928-w" rel="noopener noreferrer">10.1007/s11042-026-21928-w</a></p>
<p><strong>Keywords:</strong> open-vocabulary detection, YOLO-World, YOLOE, zero-shot object detection, prompt sensitivity, fine-tuning, computer vision, mAP, inference latency, domain shift, vision-language models, object recognition</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">204992</post-id>	</item>
	</channel>
</rss>
