<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>LoRa &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/lora/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 22:42:32 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>LoRa &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Language Models Learn to Hunt Hackers in Wireless Sensor Networks</title>
		<link>https://scienmag.com/language-models-learn-to-hunt-hackers-in-wireless-sensor-networks/</link>
		
		<dc:creator><![CDATA[Hailey Crawford]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 22:42:32 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI-based network monitoring]]></category>
		<category><![CDATA[AI-powered hacking prevention]]></category>
		<category><![CDATA[anomaly detection in wireless sensors]]></category>
		<category><![CDATA[Blackhole attack]]></category>
		<category><![CDATA[ChatTracer]]></category>
		<category><![CDATA[ChatTracer network defense]]></category>
		<category><![CDATA[cyber security]]></category>
		<category><![CDATA[DeepSeek-R1]]></category>
		<category><![CDATA[Flooding attack]]></category>
		<category><![CDATA[intrusion detection]]></category>
		<category><![CDATA[intrusion detection with language models]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models for cybersecurity]]></category>
		<category><![CDATA[LoRa]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in IoT security]]></category>
		<category><![CDATA[open wireless architecture vulnerabilities]]></category>
		<category><![CDATA[Selective Forwarding]]></category>
		<category><![CDATA[sensor network threat mitigation]]></category>
		<category><![CDATA[sensor network vulnerability]]></category>
		<category><![CDATA[translating network traffic to natural language]]></category>
		<category><![CDATA[Wireless sensor network security]]></category>
		<category><![CDATA[wireless sensor networks]]></category>
		<category><![CDATA[WSN-BFSF dataset]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=224066</guid>

					<description><![CDATA[Researchers have built ChatTracer, a lightweight intrusion detection framework that fine-tunes a compact large language model to classify wireless sensor network attacks with up to 99 percent accuracy.]]></description>
										<content:encoded><![CDATA[<p>Wireless sensor networks have quietly become the nervous system of the modern world. They monitor soil moisture across vast agricultural fields, stream patient vitals in hospitals, coordinate traffic flows in smart cities, and track conditions in environments too dangerous or remote for humans to occupy. Yet the very feature that makes these networks so useful—their open, wireless architecture—also makes them extraordinarily vulnerable. Any attacker with modest equipment can intercept, inject, or drop packets in transit, and the tiny, battery-powered sensor nodes that populate these networks rarely have the computational muscle to defend themselves. A new study published in Cluster Computing by Ayah Khaldi, Amani Krieshan, and Firas Al Balas of Jordan University of Science and Technology proposes an unexpected ally in this struggle: a compact large language model, repurposed as a network security guard.</p>
<p>The framework, called ChatTracer, represents a striking departure from how intrusion detection has traditionally been approached. Conventional machine learning systems for network defense treat traffic data as rows of numbers, feeding statistical features such as packet counts, energy levels, and routing behavior into classifiers trained to spot anomalies. ChatTracer instead does something conceptually radical: it converts structured network traffic features into language-like prompts, essentially translating the behavior of a sensor network into text that a language model can read and reason about semantically. In this framing, a suspicious pattern of dropped packets becomes a sentence the model can interpret, and a flood of spurious traffic becomes a phrase with recognizable malicious intent.</p>
<p>The choice of underlying model reflects a careful engineering trade-off between capability and practicality. Rather than deploying a massive, data-center-scale language model, the researchers built ChatTracer on DeepSeek-R1 Distill 1.5B, a distilled model small enough to be considered lightweight by modern standards. Distillation compresses the knowledge of a larger model into a smaller architecture, preserving much of its reasoning ability while dramatically reducing memory and compute requirements—an essential consideration for security tools that may eventually need to run close to the networks they protect. To adapt this general-purpose model to the specialized task of attack classification, the team applied Low-Rank Adaptation, or LoRA, a fine-tuning technique that freezes the original model weights and trains only small, low-rank update matrices injected into the network&#8217;s layers.</p>
<p>LoRA has become one of the most influential techniques in efficient model customization precisely because it sidesteps the enormous cost of full fine-tuning. Instead of updating billions of parameters, LoRA trains a tiny fraction of them, which means the adaptation process demands far less data, far less time, and far less hardware. For a security application like ChatTracer, this matters in two ways. First, it makes the framework reproducible and deployable on modest infrastructure, including edge devices that sit near sensor deployments. Second, it means the model can be retrained quickly as new attack patterns emerge, a critical property in a domain where adversaries constantly evolve their tactics.</p>
<p>To evaluate their system rigorously, the researchers turned to WSN-BFSF, a publicly available benchmark dataset specifically designed for wireless sensor network security research. The dataset captures traffic under three of the most damaging attack classes that plague these networks. Blackhole attacks lure traffic toward a malicious node that then silently discards everything it receives, creating an information sinkhole that can blind an entire region of the network. Flooding attacks overwhelm nodes and links with a torrent of spurious packets, exhausting the limited energy and bandwidth of sensor hardware. Selective forwarding attacks are subtler and often harder to catch: the malicious node forwards most packets normally but selectively drops the ones that matter most, degrading the network while maintaining an appearance of health.</p>
<p>The results reported in the study are remarkable. Under optimized training settings, ChatTracer achieved classification accuracy of up to 99 percent in detecting Blackhole, Flooding, and Selective Forwarding attacks. That figure places the language-model-based approach at the very top of what conventional machine learning methods have achieved on similar tasks, and it does so with a model that is orders of magnitude smaller than the frontier systems dominating headlines. The semantic analysis enabled by the language model appears to give it an edge in recognizing the behavioral signatures of attacks, particularly the nuanced patterns that distinguish selective forwarding from ordinary packet loss caused by unreliable wireless channels.</p>
<p>What makes this work especially compelling is its lineage. The name ChatTracer is borrowed from earlier systems that applied large language models to entirely different tracking problems, including real-time Bluetooth device tracking and multimodal visual tracking. The underlying insight of those projects—that language models can serve as powerful reasoning engines over structured, non-linguistic data when that data is rendered as prompts—turns out to generalize beautifully to network security. A parallel research thread, exemplified by systems like TrafficLLM, has been exploring generic traffic representations for network analysis, and ChatTracer extends this idea into the specific and under-served domain of wireless sensor networks, where resource constraints rule out heavyweight solutions.</p>
<p>The implications extend well beyond the laboratory. As smart cities, precision agriculture, and remote healthcare monitoring expand, the number of deployed sensor nodes is climbing into the billions, and each one is a potential entry point for attackers. A compromised sensor network in a hospital could falsify patient data; one in an agricultural system could sabotage irrigation; one in an industrial setting could mask dangerous conditions. Traditional intrusion detection systems, often designed for resource-rich servers guarding enterprise networks, translate poorly to this environment. A lightweight framework that can achieve near-perfect detection accuracy while remaining small enough for edge deployment addresses precisely the gap that has left sensor deployments exposed.</p>
<p>There are, of course, caveats and open questions. The 99 percent accuracy figure comes from a specific benchmark dataset with known attack types, and real-world deployments will present noisier data, novel attacks, and distribution shifts that challenge any trained model. The study evaluated three attack classes; wireless sensor networks face a broader menagerie of threats, including Sybil attacks, sinkhole attacks, and spoofing, each with distinct signatures. Whether the prompt-based semantic approach generalizes to these other attack families, and how the system performs under adversarial pressure from attackers who know they are being watched by a language model, remain subjects for future work. The authors received no external funding for the research and declare no competing interests, and the public availability of the WSN-BFSF dataset means other teams can immediately test and extend the approach.</p>
<p>Still, the study marks a genuine milestone in the convergence of two fields that seemed, until recently, to have little to say to each other. Language models were built to read and write; sensor networks were built to sense and transmit. By teaching a small, efficient model to read a network the way it reads text, ChatTracer demonstrates that the semantic reasoning power of large language models can be compressed, adapted, and pointed at one of the most stubborn security problems of the connected age. If the approach scales from benchmark to battlefield, the tiny sensors scattered across farms, cities, and hospitals may soon carry with them a guardian that understands their traffic not as numbers, but as a story—one in which attacks, however subtle, betray themselves in the telling.</p>
<p><strong>Subject of Research:</strong> Using fine-tuned large language models for intrusion detection in wireless sensor networks</p>
<p><strong>Article Title:</strong> Enhancing cyber security in wireless sensor networks using ChatTracer in Large Language Models (LLMs)</p>
<p><strong>Article References:</strong> Khaldi, A., Krieshan, A., &amp; Balas, F. A. (2026). Enhancing cyber security in wireless sensor networks using ChatTracer in Large Language Models (LLMs). <em>Cluster Computing, 29</em>(14), Article 811. <a href="https://doi.org/10.1007/s10586-026-06622-8" rel="noopener noreferrer">https://doi.org/10.1007/s10586-026-06622-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10586-026-06622-8" rel="noopener noreferrer">10.1007/s10586-026-06622-8</a></p>
<p><strong>Keywords:</strong> wireless sensor networks, large language models, ChatTracer, intrusion detection, cyber security, DeepSeek-R1, LoRA, WSN-BFSF dataset, Blackhole attack, Flooding attack, Selective Forwarding, machine learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">224066</post-id>	</item>
		<item>
		<title>AI Image Editing Gets Surgical: New GGIP2P System Pins Down Exactly What to Change</title>
		<link>https://scienmag.com/ai-image-editing-gets-surgical-new-ggip2p-system-pins-down-exactly-what-to-change/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 02:41:03 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[addressing instruction ambiguity in image editing]]></category>
		<category><![CDATA[advancements in AI-based photo editing]]></category>
		<category><![CDATA[AI image editing]]></category>
		<category><![CDATA[AI system for precise object modifications]]></category>
		<category><![CDATA[BERT]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[diffusion architecture in image manipulation]]></category>
		<category><![CDATA[diffusion models]]></category>
		<category><![CDATA[Guided-Grounded-InstructPix2Pix (GGIP2P)]]></category>
		<category><![CDATA[handling pronouns and distractor objects in image modification]]></category>
		<category><![CDATA[instruction-based image editing]]></category>
		<category><![CDATA[InstructPix2Pix]]></category>
		<category><![CDATA[interactive AI image editing systems]]></category>
		<category><![CDATA[LoRa]]></category>
		<category><![CDATA[mask-guided generation]]></category>
		<category><![CDATA[named entity recognition]]></category>
		<category><![CDATA[object grounding]]></category>
		<category><![CDATA[precise spatial guidance in AI image editing]]></category>
		<category><![CDATA[pronoun resolution]]></category>
		<category><![CDATA[size prediction]]></category>
		<category><![CDATA[spatial reasoning]]></category>
		<category><![CDATA[spatially guided image synthesis]]></category>
		<category><![CDATA[target localization in image editing]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=220966</guid>

					<description><![CDATA[Researchers at the University of Kashan have developed GGIP2P, a modular pipeline that combines BERT-based entity recognition, pronoun resolution, spatial reasoning, and size prediction to give instruction-based image editors precise control over where edits occur.]]></description>
										<content:encoded><![CDATA[<p>Instruction-based image editing has promised a future in which anyone can transform a photograph simply by typing what they want changed. Tell a model to make the sky stormy or turn a dog into a fox, and the software should comply. In practice, however, even the most advanced systems stumble the moment an instruction becomes genuinely conversational. Ask an editor to change the bird on the right, and it may alter the wrong bird, both birds, or something else entirely. A new study published in Multimedia Tools and Applications by Zahra Esmaily and Hossein Ebrahimpour-Komleh of the University of Kashan tackles this failure head-on with a system called Guided-Grounded-InstructPix2Pix, or GGIP2P, which gives instruction-following editors a far more precise sense of where, exactly, the edit should happen.</p>
<p>The core problem the researchers identified is target localization. State-of-the-art editing models built on diffusion architectures can generate stunning imagery, but they treat instructions as loose global guidance rather than precise spatial commands. When instructions contain pronouns, distractor objects, or intricate spatial relations, the models frequently misidentify what should be edited. Existing approaches also lack any mechanism for spatially guiding the generation of new objects that are absent from the original image, meaning a request to place a butterfly atop a candle is handled with little control over where the butterfly actually lands. GGIP2P addresses both weaknesses with a modular pipeline that performs multi-step grounding and instruction disambiguation before any pixels are touched.</p>
<p>The first stage of the pipeline reformulates target identification as a Named Entity Recognition task, a technique borrowed from natural language processing where systems are trained to spot and classify mentions of entities within text. Rather than relying on generic noun-phrase extraction, the team fine-tuned a BERT language model using LoRA, or Low-Rank Adaptation, a parameter-efficient fine-tuning method that trains small sets of low-rank weight matrices instead of updating the full network. This approach keeps the adapted model remarkably lightweight, at roughly 10 megabytes, while giving it the ability to recognize which tokens in an instruction actually refer to the editing target, even when the reference is indirect or wrapped in spatial modifiers.</p>
<p>Once candidate targets are identified linguistically, GGIP2P applies a series of reasoning filters to resolve ambiguity. A pronoun resolution mechanism rewrites instructions containing words like them or it into explicit references, eliminating referential ambiguity before grounding begins. A plurality-aware bounding-box filter handles cases where an instruction mentions a category with multiple instances in the scene, deciding whether the instruction refers to one object or several. A spatial reasoning unit then interprets both absolute directional cues, such as the upper fan or the right bird, and relative directional cues, such as the horse on the left side of the gray horse. Together these modules convert a free-form sentence into an unambiguous spatial specification that can be mapped onto the image.</p>
<p>The second major innovation is a guided object generation component for edits that introduce things not present in the original scene. Because there is no existing object to ground, the system must decide where the new object should appear and how large it should be. GGIP2P employs a size-prediction model that estimates the relative dimensions of the absent object, enabling precise mask-guided placement within the diffusion editing process. The authors report that their size-prediction module is also highly lightweight, around 10 megabytes, and the ablation studies show that the placement masks it produces are accurate enough to support coherent object synthesis while preserving the surrounding background.</p>
<p>Quantitative results back up the design. In an ablation study measuring localization accuracy with Intersection-over-Union, replacing the NER-based target detector with a simple noun phrase extractor caused performance to fall from 0.71 to 0.42, demonstrating that token-level target identification is a critical driver of correct grounding. Removing the spatial reasoning module, which unifies plurality-based bounding-box disambiguation and directional cue reasoning, degraded IoU to 0.44. The researchers also isolated the effect of pronoun handling on the 63 pronoun-based instructions in their test set: enabling the pronoun-replacement module lifted IoU from 0.40 to 0.64, indicating that most grounding failures in pronoun instructions stem from unresolved referential ambiguity rather than any weakness in the underlying vision system.</p>
<p>Efficiency benchmarks add an important practical dimension. Tested under identical conditions on a standardized NVIDIA T4 GPU in a Google Colab environment, GGIP2P achieved an end-to-end latency of 20.23 seconds per image while consuming 9.27 gigabytes of VRAM, outperforming the earlier Grounded-InstructPix2Pix baseline, which required 21.85 seconds and 10.10 gigabytes. The performance dividend comes directly from the specialized BERT plus LoRA NER module, which filters target entities entirely within the textual domain and thereby avoids the redundant, iterative cross-modal CLIP scoring that slows down competing pipelines. By contrast, InstructEdit, despite a slightly lower VRAM footprint of 8.98 gigabytes, incurred a prohibitive latency of 70.48 seconds per image due to its inversion-based optimization process. The base InstructPix2Pix model remains fastest simply because it lacks grounding and masking altogether, but it pays for that speed with far poorer targeting fidelity.</p>
<p>The evaluation methodology itself reflects careful engineering. To measure the semantic correctness of generated content, the team created a Masked-Targeted CLIP Score, for which they manually wrote simplified target descriptions for each complex instruction in their test set, stripping away contextual, spatial, and action-related words to isolate the description of the final desired object. They also supplied ground-truth textual phrases to a competing method that optionally uses BLIP and GPT-3 for phrase generation, ensuring that comparisons focused on editing performance rather than language-generation reliability. Qualitative comparisons across challenge categories, including plural pronoun resolution, absolute and relative spatial reasoning, spatially guided object generation, and targetless directional edits such as adding a sunset to the top of an image, showed GGIP2P consistently editing the intended region while leaving the rest of the scene untouched.</p>
<p>The study is candid about limitations. An error analysis of the size-prediction model, broken down by object scale, revealed that small objects under 50 pixels, which make up roughly a third of the evaluation set, exhibit a relative percentage error of 44.8 percent, a regime where minor absolute pixel errors translate into disproportionately large relative errors. When predicted masks become too small, the diffusion model often fails to generate complete or coherent objects, producing truncated or visually implausible outputs. The team&#8217;s mitigation is a clamping strategy that enforces a minimum mask size of 50 pixels, which they show significantly improves generation quality and perceptual consistency. This kind of honest error accounting, paired with open availability of the source code, datasets, and pretrained models in a public GitHub repository, strengthens confidence in the reported gains.</p>
<p>Taken together, the results suggest that the future of conversational image editing lies less in ever-larger generative backbones and more in the intelligent scaffolding wrapped around them. GGIP2P demonstrates that lightweight language-side reasoning, efficient fine-tuning, and explicit spatial modeling can deliver targeting precision that brute-force generation cannot match, while adding negligible computational overhead. For applications ranging from photo retouching to design workflows and content creation, the ability to say exactly what should change and have the system understand both the words and the geometry behind them marks a meaningful step toward editing tools that behave the way users intuitively expect. The work was partially supported by the Research Council of the University of Kashan, and the authors report no competing interests.</p>
<p><strong>Subject of Research:</strong> Instruction-based image editing with precise target grounding and spatially guided object generation</p>
<p><strong>Article Title:</strong> Guided-Grounded-InstructPix2Pix (GGIP2P): Precise targeting and spatially guided control for instruction-based image editing</p>
<p><strong>Article References:</strong> Esmaily, Z., &amp; Ebrahimpour-Komleh, H. (2026). Guided-Grounded-InstructPix2Pix (GGIP2P): Precise targeting and spatially guided control for instruction-based image editing. <em>Multimedia Tools and Applications, 85</em>(10), Article 784. <a href="https://doi.org/10.1007/s11042-026-21941-z" rel="noopener noreferrer">https://doi.org/10.1007/s11042-026-21941-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11042-026-21941-z" rel="noopener noreferrer">10.1007/s11042-026-21941-z</a></p>
<p><strong>Keywords:</strong> instruction-based image editing, diffusion models, object grounding, BERT, LoRA, named entity recognition, spatial reasoning, pronoun resolution, InstructPix2Pix, mask-guided generation, computer vision, size prediction</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">220966</post-id>	</item>
		<item>
		<title>Genetic Algorithms Meet LoRA: New Study Tests Whether Smarter Search Really Beats Simple Tuning</title>
		<link>https://scienmag.com/genetic-algorithms-meet-lora-new-study-tests-whether-smarter-search-really-beats-simple-tuning/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 01:18:05 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[automated hyperparameter search in AI]]></category>
		<category><![CDATA[comparison of LoRA and manual tuning strategies]]></category>
		<category><![CDATA[cost-benefit analysis of model fine-tuning methods]]></category>
		<category><![CDATA[cost-effective fine-tuning of large language models]]></category>
		<category><![CDATA[DistilBERT]]></category>
		<category><![CDATA[efficiency of low-rank updates in neural networks]]></category>
		<category><![CDATA[evaluation]]></category>
		<category><![CDATA[evaluation of search algorithms in AI model adaptation]]></category>
		<category><![CDATA[genetic algorithm]]></category>
		<category><![CDATA[Genetic algorithms for model tuning]]></category>
		<category><![CDATA[hyperparameter optimization]]></category>
		<category><![CDATA[impact of automated search on model performance]]></category>
		<category><![CDATA[inference latency]]></category>
		<category><![CDATA[LoRa]]></category>
		<category><![CDATA[Low-Rank Adaptation (LoRA) in transformer models]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[model sparsity]]></category>
		<category><![CDATA[optimization techniques for pretrained transformers]]></category>
		<category><![CDATA[parameter-efficient fine-tuning]]></category>
		<category><![CDATA[Pareto optimization]]></category>
		<category><![CDATA[role of evolutionary algorithms in AI model configuration]]></category>
		<category><![CDATA[selective unfreezing]]></category>
		<category><![CDATA[text classification]]></category>
		<category><![CDATA[transfer learning]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=220682</guid>

					<description><![CDATA[A matched-budget study of LoRA fine-tuning finds that genetic algorithms and other search optimizers offer no consistent advantage over well-tuned baselines, and that apparent gains often come from extra model capacity rather than smarter search.]]></description>
										<content:encoded><![CDATA[<p>Fine-tuning large language models has become one of the most expensive rituals in modern artificial intelligence. Every time a pretrained transformer must be adapted to a new task, engineers face a choice: retrain the entire network, or find a shortcut that preserves most of the pretrained knowledge while updating only a sliver of the parameters. A new study published in Machine Learning with Applications by Seda Bayat Toksoz and colleagues takes a hard, unusually honest look at one of the most popular shortcuts — Low-Rank Adaptation, or LoRA — and asks a question the field has largely avoided: when automated search algorithms pick the LoRA configuration, are they actually better than a well-tuned baseline, or are they just getting a bigger slice of the evaluation budget?</p>
<p>The appeal of LoRA is easy to explain. Instead of updating a full weight matrix inside a transformer, LoRA freezes the pretrained weights and learns a small low-rank residual update, expressed as the product of two much thinner matrices. For a hidden size of 768, as in DistilBERT, a rank-4 update on a single attention projection adds only 6,144 trainable coefficients. Once training is complete, the residual can be merged back into the original projection, leaving the deployed model essentially identical in shape and speed to the original. But beneath that elegance lies a thicket of decisions: which layers to adapt, which attention projections to target, what rank to use, how to set the scaling factor, dropout, and batch size. Each choice shifts the balance between predictive power and efficiency, and no universal recipe exists.</p>
<p>Recent research has increasingly treated these decisions as an optimization problem in their own right. Methods such as AdaLoRA reallocate rank during training, AutoPEFT searches over entire parameter-efficient configurations, AutoLoRA estimates layer-wise ranks through meta-learning, and BIPEFT separates structural selection from rank-budget allocation. What these studies share is a tendency to compare their search strategies against fixed, sometimes under-tuned baselines. The new work attacks that weakness directly by imposing a matched-budget protocol: every method — a genetic algorithm, the Slime Mould Algorithm, Optuna&#8217;s Bayesian optimization, a Pareto-based genetic algorithm, tuned Vanilla LoRA, and tuned full fine-tuning — received exactly 25 validation-only candidate evaluations per task. After selection, each winning configuration was retrained independently with five random seeds, and held-out test data were touched only after every configuration decision was locked in.</p>
<p>The experimental scope was deliberately heterogeneous. DistilBERT served as the main encoder across three very different text classification tasks: the Emotion benchmark, an imbalanced six-class affective dataset with a 9.4-to-1 ratio between its largest and smallest classes; AG News, a large balanced four-class topic benchmark; and Financial PhraseBank, a small domain-specific three-class sentiment dataset. A secondary check used RoBERTa-Base on the GLUE SST-2 task. This diversity matters because a search strategy that shines on one regime may quietly fail on another, and the authors wanted to know whether any optimizer&#8217;s advantage survived contact with different data conditions.</p>
<p>The headline finding is sobering for fans of clever search. On AG News, where full fine-tuning reached 94.4 percent accuracy, GA-guided selective unfreezing came closest at 94.3 percent, but the tuned Vanilla LoRA baseline itself hit 94.1 percent — and the three direct searched-LoRA variants (Optuna, GA, and SMA) landed at 93.9, 93.7, and 93.4 percent respectively. On Financial PhraseBank, where run-to-run variability was three to five times larger, GA+SU-LoRA posted the highest numerical mean at 86.2 percent versus 86.0 percent for full fine-tuning, a gap the authors explicitly decline to interpret as superiority. Across tasks, the ordering of optimizers shifted, and none consistently beat the tuned references. Configuration search, the authors conclude, is best understood as a mechanism for locating task-specific operating points, not as a way to crown a universally dominant algorithm.</p>
<p>Perhaps the most revealing result came from the component ablations on the Emotion dataset. The study&#8217;s GA+SU-LoRA method combines a genetic-algorithm-selected layer and projection support with selective unfreezing, which makes the corresponding pretrained attention matrices trainable. It achieved 93.6 percent accuracy. But a control condition — selective unfreezing applied without any GA-selected support — reached 93.8 percent. The conclusion is uncomfortable but important: the gain came primarily from exposing additional backbone capacity, not from the intelligence of the search. Similarly, a sparse variant that masked roughly half of the LoRA coefficients scored 92.2 percent, essentially indistinguishable from a sparsification control without GA support and only 0.4 points below the tuned Vanilla LoRA baseline, with an approximate Welch 95 percent confidence interval spanning −0.84 to +0.04 points.</p>
<p>The sparsity result also carries a deployment warning. The masked adapters achieved 51.6 percent mean coefficient sparsity across the LoRA factors, but because the zeros were unstructured, the dense tensor shapes remained unchanged and conventional GPU kernels executed the same matrix operations as before. After merging the adapter updates into the base projections, latency stayed within roughly 1 to 3 percent of the full fine-tuning reference, and peak inference memory was essentially unchanged. Compactness on paper, in other words, does not automatically translate into speed in production. The authors argue that future work should distinguish carefully between dense trainable allocation, effective active coefficients, structural model size, and realized deployment cost — four quantities that are routinely conflated in the parameter-efficient fine-tuning literature.</p>
<p>The study also scrutinized the objective function used to guide search. The original scalar fitness combined validation macro-F1 with a compactness penalty weighted by a coefficient lambda. A sensitivity analysis showed this penalty was treacherous: at lambda equal to 0.05, validation accuracy dropped to the 89.8 to 90.8 percent range as the search shrank the layer support to roughly four layers, while removing the penalty entirely recovered 1.3 to 2.0 percentage points. A Pareto-based genetic algorithm that treated accuracy and parameter count as separate objectives exposed the trade-off directly, producing a knee region at 91.9 to 92.5 percent validation accuracy with four to five active layers — a more transparent and robust way to balance performance against compactness than baking the trade-off into a single coefficient whose meaning shifts with dataset scale.</p>
<p>Resource measurements rounded out the deployment picture. The genetic search itself consumed about 2.5 GPU-hours on a single RTX 4070, a modest cost by modern standards. Final retraining runs took between 8 and 15 minutes depending on the method, with peak training memory of 7.4 gigabytes for full fine-tuning, 5.8 gigabytes for the selective-unfreezing variant, and 3.2 gigabytes for the sparse variant. Notably, selective unfreezing raised the trainable ratio to about 11.56 percent of the model — far above conventional LoRA&#8217;s roughly 1 percent — which means it should never be described as adapter-only tuning, even though it delivers accuracy close to full fine-tuning at a fraction of the memory.</p>
<p>The broader lesson of this study may outlast any of its specific numbers. In a field racing to automate every design decision, the authors demonstrate that credible comparisons require matched tuning effort, explicit component-isolating controls, repeated final training across seeds, and deployment-aware reporting. When those safeguards are in place, the mystique of sophisticated search algorithms fades: what remains is a set of practical tools for finding task-specific trade-offs, with the real performance gains coming from controlled access to pretrained capacity rather than from the search itself. For practitioners choosing between optimizers, the message is clear — spend the effort tuning your baseline first, because a well-tuned simple method is a harder target than most papers admit.</p>
<p><strong>Subject of Research:</strong> Parameter-efficient fine-tuning of encoder language models using automated LoRA configuration search and sparse low-rank adaptation</p>
<p><strong>Article Title:</strong> Efficient encoder fine-tuning through configuration search and sparse low-rank adaptation</p>
<p><strong>Article References:</strong> Toksoz, S. B., Isik, G., Turkmen, H., Pacal, I., &amp; Keles, A. (2026). Efficient encoder fine-tuning through configuration search and sparse low-rank adaptation. <em>Machine Learning with Applications, 26</em>, Article 101024. <a href="https://doi.org/10.1016/j.mlwa.2026.101024" rel="noopener noreferrer">https://doi.org/10.1016/j.mlwa.2026.101024</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.mlwa.2026.101024" rel="noopener noreferrer">10.1016/j.mlwa.2026.101024</a></p>
<p><strong>Keywords:</strong> LoRA, parameter-efficient fine-tuning, DistilBERT, genetic algorithm, hyperparameter optimization, model sparsity, selective unfreezing, Pareto optimization, text classification, transfer learning, inference latency, machine learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">220682</post-id>	</item>
		<item>
		<title>New Zeroth-Order Method Tames Gradient-Free Fine-Tuning of Large Language Models</title>
		<link>https://scienmag.com/new-zeroth-order-method-tames-gradient-free-fine-tuning-of-large-language-models/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 20:14:59 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[application of zeroth-order methods to AI model adaptation]]></category>
		<category><![CDATA[black-box model optimization techniques]]></category>
		<category><![CDATA[black-box optimization]]></category>
		<category><![CDATA[convergence guarantees]]></category>
		<category><![CDATA[cost reduction in zeroth-order optimization]]></category>
		<category><![CDATA[efficient parameter update strategies for large]]></category>
		<category><![CDATA[gradient-free fine-tuning of neural networks]]></category>
		<category><![CDATA[gradient-free methods]]></category>
		<category><![CDATA[gradient-free model training methods]]></category>
		<category><![CDATA[gradient-free optimization in restricted hardware environments]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models without backpropagation]]></category>
		<category><![CDATA[latent subspace]]></category>
		<category><![CDATA[LoRa]]></category>
		<category><![CDATA[low-rank projection]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[new zeroth-order method for large models]]></category>
		<category><![CDATA[optimization]]></category>
		<category><![CDATA[overcoming gradient computation limitations in deep learning]]></category>
		<category><![CDATA[parameter-efficient fine-tuning]]></category>
		<category><![CDATA[random subspace]]></category>
		<category><![CDATA[resource-efficient large language model fine-tuning]]></category>
		<category><![CDATA[zeroth-order optimization]]></category>
		<category><![CDATA[zeroth-order optimization for large language model fine-tuning]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=218826</guid>

					<description><![CDATA[Researchers have introduced PLS-ZO, a zeroth-order optimization framework that confines gradient-free perturbations to a low-dimensional latent subspace, dramatically reducing the variance that has limited backpropagation-free fine-tuning of large language models.]]></description>
										<content:encoded><![CDATA[<p>Fine-tuning a large language model usually demands something deceptively simple: gradients. Every step of conventional training relies on backpropagation, the algorithmic machinery that propagates error signals backward through billions of parameters to tell each weight how it should change. But there are increasingly common scenarios in which that machinery is unavailable. A model may be accessed only through an application programming interface that returns outputs but hides its internals. It may be so large, or stored so aggressively quantized, that keeping activations in memory for a backward pass is prohibitively expensive. It may sit on hardware where automatic differentiation simply is not supported. In all of these black-box and resource-constrained settings, researchers have turned to zeroth-order optimization, a family of methods that updates parameters using nothing more than the values the model returns when queried. A new study published in the International Journal of Machine Learning and Cybernetics by Ximin Zhang and Shuhua Yuan of the Henan Engineering Research Center of Fault-Tolerant Server in Zhengzhou, China, argues that the standard versions of these methods carry a hidden cost that grows catastrophically with model size, and proposes an elegant way to pay far less.</p>
<p>The core idea of zeroth-order optimization is old and intuitively appealing. Rather than computing a gradient analytically, one estimates it by perturbing the parameters slightly in random directions and observing how the loss changes. If nudging a weight upward makes the model worse and nudging it downward makes it better, the estimator knows, roughly, which way to push. Two function evaluations per step, in the so-called two-point estimator popularized in modern language model work such as the MeZO algorithm, are enough to produce a usable descent direction. The catch is statistical. A random direction in a space of millions or billions of dimensions is almost certainly nearly orthogonal to the true gradient, so the difference in loss between two nearby points along that direction is dominated by noise. The variance of the resulting gradient estimate scales with the dimension of the perturbation space, and in parameter-efficient fine-tuning, where the trainable set may still contain millions of parameters, that variance can be severe enough to make training crawl or stall.</p>
<p>Parameter-efficient fine-tuning itself was supposed to solve the scale problem. Techniques such as adapters, prefix-tuning, prompt tuning, BitFit and the now-ubiquitous low-rank adaptation known as LoRA freeze the bulk of a pretrained model and train only small inserted modules or subsets of weights. This slashes memory for optimizer states and makes adaptation feasible on modest hardware when gradients are available. But as Zhang and Yuan point out, even a reduced trainable set frequently reaches into the millions of parameters, and a zeroth-order estimator perturbing all of them simultaneously inherits variance proportional to that count. The reduction that makes PEFT attractive under backpropagation does relatively little for the statistical burden of gradient-free methods. Previous lines of work have attacked the problem from several directions, including sparse perturbations, transferable static sparsity, low-rank structures, curvature-aware estimates, preconditioned and accelerated variants, and random subspace methods that restrict perturbations to a randomly chosen subset of coordinates. The new paper builds on this momentum but changes where the compression happens.</p>
<p>The proposed framework, called Projected Latent Subspace Zeroth-Order Optimization, or PLS-ZO, restricts perturbations not to a subset of coordinates but to a deliberately constructed low-dimensional latent subspace, and it does so at the level of individual weight matrices rather than the flattened parameter vector. The authors construct hierarchical low-rank random projections that map each trainable matrix into a controllable latent space of much lower dimension. Random projection is a classical tool from linear algebra and randomized numerical computation: a high-dimensional vector multiplied by a random matrix retains, with high probability, the geometric information that matters, compressed into far fewer coordinates. By applying this idea hierarchically across the matrices of the model, PLS-ZO creates a latent coordinate system in which the optimization actually takes place. Perturbations are drawn in this latent space, and the difference of loss values between perturbed and unperturbed forward passes yields an estimated gradient with variance governed by the latent dimension rather than by the original parameter count.</p>
<p>A crucial and somewhat distinctive design choice follows from this. Many subspace methods treat the projection as a way to generate cheaper perturbations but still run the optimizer in the original high-dimensional coordinates. PLS-ZO instead performs the main optimizer operations, including momentum accumulation, preconditioning and the bookkeeping of optimizer states, directly in the low-dimensional latent coordinates. Momentum in particular benefits: an exponential moving average of past gradient estimates is a powerful variance reducer, but maintaining and applying it over millions of coordinates is exactly the kind of overhead zeroth-order methods are supposed to avoid. Working in latent space keeps those state tensors small and makes preconditioning, which rescales update directions to account for uneven curvature, computationally practical. The result is an optimizer whose memory footprint and per-step arithmetic scale with the chosen latent dimension, a hyperparameter the practitioner controls, rather than with the full size of the trainable parameter set.</p>
<p>One subtlety of any random-subspace approach is that the subspace must eventually change. A single fixed random projection risks missing important directions of variation in the loss landscape for the entire run, so PLS-ZO periodically refreshes its projections, drawing new random bases in which optimization continues. Naively, each refresh would force the optimizer to discard its accumulated momentum and preconditioner statistics, throwing away the very variance reduction that makes the method work, or would require expensive re-projection of large state tensors. The authors introduce what they call a lazy subspace refresh strategy, which defers and amortizes the cost of regenerating projections, together with a principled state migration rule that projects the latent optimizer states from the old basis into the new one. In effect, the optimizer&#8217;s memory of its trajectory survives the change of coordinates, so long-horizon training remains stable instead of repeatedly restarting from a cold state.</p>
<p>Beyond the algorithmic engineering, the paper contributes theory. The authors prove that the bias and variance of the two-point zeroth-order estimator under their scheme are governed by the latent subspace dimension rather than the original parameter dimension, formalizing the intuition that compressing the perturbation space compresses the noise. They further establish a nonconvex convergence guarantee for the setting in which the random subspaces are periodically refreshed, showing that the algorithm converges to a stationary point at a rate consistent with classical zeroth-order analysis despite the moving coordinate system. Nonconvex guarantees of this kind matter because the loss surfaces of fine-tuned neural networks are emphatically nonconvex, and a method whose theoretical behavior degrades under basis changes would be difficult to trust in practice. The migration rule and refresh schedule are precisely the components that make the analysis go through.</p>
<p>On the empirical side, the authors report extensive experiments demonstrating that PLS-ZO consistently achieves highly effective fine-tuning performance under the parameter-efficient setting. The evaluation draws on widely used natural language understanding benchmarks of the kind standard in this literature, including datasets such as SST-2, BoolQ, WiC, WinoGrande and MultiRC that have anchored comparisons among gradient-free fine-tuning methods, applied to robustly pretrained encoder models in the RoBERTa family. The paper&#8217;s data availability statement notes that no new datasets were generated or analyzed, meaning the contribution lies in optimization methodology rather than data collection. The experiments position PLS-ZO against the backdrop of a rapidly crowding field: recent work has introduced low-rank zeroth-order fine-tuning, temporally low-rank perturbations, gradient-aligned projections, minimum-variance two-point estimators, Hessian-informed optimizers, quantized zeroth-order training, and learned optimizers, all racing to make forward-pass-only adaptation competitive with full backpropagation.</p>
<p>The significance of this line of research extends well beyond academic benchmarks. As large models are deployed at the edge, embedded in privacy-sensitive pipelines, or offered as closed services, the assumption that a practitioner can differentiate through the model is increasingly fragile. Zeroth-order fine-tuning requires only the ability to run the model forward and read off a loss, which makes it compatible with API-based adaptation, on-device training under tight memory budgets, and models whose weights are frozen by policy as much as by architecture. Methods like PLS-ZO attack the central statistical obstacle, dimension-induced variance, with a control knob, the latent dimension, that lets practitioners trade a small amount of expressive freedom for large gains in estimate quality and compute. If the theoretical promise holds up under wider independent testing, the practical upshot is that adapting a frontier-scale model could become feasible for organizations with neither the hardware nor the internal access that today&#8217;s fine-tuning workflows assume, a shift that could meaningfully democratize who gets to specialize the most powerful models in existence.</p>
<p><strong>Subject of Research:</strong> Zeroth-order optimization for parameter-efficient fine-tuning of large language models</p>
<p><strong>Article Title:</strong> PLS-ZO: projected latent subspace zeroth-order optimization for parameter-efficient fine-tuning</p>
<p><strong>Article References:</strong> Zhang, X., &amp; Yuan, S. (2026). PLS-ZO: projected latent subspace zeroth-order optimization for parameter-efficient fine-tuning. <em>International Journal of Machine Learning and Cybernetics, 17</em>(10), Article 488. <a href="https://doi.org/10.1007/s13042-026-03317-9" rel="noopener noreferrer">https://doi.org/10.1007/s13042-026-03317-9</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s13042-026-03317-9" rel="noopener noreferrer">10.1007/s13042-026-03317-9</a></p>
<p><strong>Keywords:</strong> zeroth-order optimization, parameter-efficient fine-tuning, large language models, random subspace, low-rank projection, black-box optimization, gradient-free methods, LoRA, latent subspace, convergence guarantees, machine learning, optimization</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">218826</post-id>	</item>
		<item>
		<title>FuseDepth Combines Frozen AI Visual Priors to Unlock True Metric Depth From a Single Photo</title>
		<link>https://scienmag.com/fusedepth-combines-frozen-ai-visual-priors-to-unlock-true-metric-depth-from-a-single-photo/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sat, 26 Sep 2026 00:59:19 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI-based scene understanding]]></category>
		<category><![CDATA[combining relative and absolute depth models]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[computer vision depth sensing]]></category>
		<category><![CDATA[Depth Anything V2]]></category>
		<category><![CDATA[depth discontinuities]]></category>
		<category><![CDATA[domain generalization in depth estimation]]></category>
		<category><![CDATA[foundation models]]></category>
		<category><![CDATA[fusion of foundation models]]></category>
		<category><![CDATA[image matting]]></category>
		<category><![CDATA[improved depth boundary accuracy]]></category>
		<category><![CDATA[LoRa]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[metric depth]]></category>
		<category><![CDATA[metric depth from single images]]></category>
		<category><![CDATA[monocular depth estimation]]></category>
		<category><![CDATA[off-the-shelf AI models]]></category>
		<category><![CDATA[open-access depth estimation frameworks]]></category>
		<category><![CDATA[prior fusion]]></category>
		<category><![CDATA[scene scale calibration]]></category>
		<category><![CDATA[surface normals]]></category>
		<category><![CDATA[zero-shot depth prediction]]></category>
		<category><![CDATA[zero-shot learning]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=215787</guid>

					<description><![CDATA[A new framework called FuseDepth fuses frozen detection, matting and surface-normal models with an adapted Depth Anything backbone to achieve sharper, truly metric depth estimation on unseen scenes without any test-time calibration.]]></description>
										<content:encoded><![CDATA[<p>Estimating how far away every pixel in a photograph truly lies, in meters rather than merely in relative order, has long been one of computer vision&#8217;s most stubborn challenges. A new open-access study published in Machine Learning with Applications introduces FuseDepth, a framework that pushes zero-shot monocular depth estimation closer to practical deployment by fusing several frozen, off-the-shelf AI models into a single pipeline. Rather than training yet another giant network, the approach borrows the strengths of existing foundation models and composes them in a structured way, achieving sharper boundaries and more reliable absolute scale on scenes the system has never seen.</p>
<p>The central problem the paper identifies is a coupled failure mode. Relative-depth models such as Depth Anything v2 and Marigold produce visually convincing depth maps that generalize well across domains, but their outputs are not metrically calibrated, so converting them to true distances typically requires per-image scale and shift fitting against ground-truth data, which is unavailable in real deployments. Conversely, metric-aware systems such as ZoeDepth, Metric3Dv2, DepthPro, ZeroDepth and UniDepth anchor absolute scale better, yet can blur object boundaries and produce locally inconsistent geometry when the scene distribution shifts. FuseDepth attacks both weaknesses simultaneously instead of trading one for the other.</p>
<p>The architecture works in two stages. In the first, a frozen Depth Anything v2 backbone is adapted with lightweight low-rank adapters, known as LoRA modules, and paired with a boundary-regularized local metric-scale field. Instead of applying a single global affine calibration, a small convolutional head predicts spatially varying scale and shift values, biased toward low-frequency variation and piecewise smoothness. Crucially, the regularization is relaxed near estimated occlusion boundaries, allowing metric scale to change legitimately where objects end, while penalizing arbitrary pixelwise calibration everywhere else. No per-image or per-dataset scale fitting occurs at inference.</p>
<p>The second stage injects explicit high-fidelity visual priors. A pretrained YOLOv8 detector proposes object regions, and each proposal is refined with BiRefNet, an image-matting model that recovers far sharper alpha boundaries than conventional segmentation masks. The gradients of these alpha mattes are aggregated into a scene-wide contour map, and overlapping or adjacent surface hypotheses are linked into a deterministic surface-interaction graph. From this graph, the method derives a soft jump set marking locations where depth may legitimately discontinuously change and where smoothness assumptions should be suspended.</p>
<p>A complementary geometric prior comes from Omnidata, a frozen monocular surface-normal estimator. Surface normals describe the local orientation of planes and curved surfaces, information that is orthogonal to boundary evidence. FuseDepth enforces a projective compatibility between predicted depth gradients and the normal field, computed under the camera&#8217;s intrinsic matrix, but this constraint is deliberately suppressed wherever the graph-derived jump set indicates an occlusion boundary. The result is a piecewise regularization strategy: smoothness and normal consistency hold within surfaces, while boundary-preserving correction takes over at discontinuities.</p>
<p>A compact residual network of roughly 0.2 million parameters then refines the coarse metric depth. Rather than simply concatenating inputs, the refiner emits three residual proposals per pixel, one each for contour, normal and coarse-depth corrections, together with softmax-normalized arbitration weights that mix them locally. These learned gates allow the model to favor the contour branch near detected boundaries and lean on normals or the coarse depth in smooth or uncertain regions. The trainable components are learned only through the final depth losses, without any explicit reliability labels.</p>
<p>Evaluation follows a deliberately strict protocol. FuseDepth is trained on roughly 50,000 images from NYUv2, ScanNet and KITTI, and tested zero-shot on three unseen benchmarks spanning panoramic HDR scenes with LiDAR depth, long-range driving footage and diverse indoor and outdoor RGB-D imagery. Relative-depth baselines receive only one affine mapping fitted once on the training mixture and then frozen. Across all targets, FuseDepth posts the lowest errors, reaching an absolute relative error of 0.105 on SYNS, 0.171 on DDAD and 0.182 and 0.276 on the indoor and outdoor DIODE splits, while also recording the lowest per-image log-scale bias, a direct measure of absolute-scale stability.</p>
<p>Mechanism-isolation experiments support the design choices. Removing the contour prior, the normal prior, or the graph structure each degrades accuracy, and replacing the learned gate with plain concatenation or fixed equal weighting is consistently worse. Deliberately corrupted external priors, such as alpha masks taken from the wrong image or normals from a different scene, cause only graceful degradation because the gate shifts its mass away from the corrupted branch. Statistical analysis shows the improvements over the strongest baseline, UniDepth, are significant at the 95 percent level on SYNS and DIODE, while the DDAD advantage is small but reproduced across three independent training seeds and all geographic subsets.</p>
<p>The authors are candid about limitations. The pipeline depends on the quality and cost of its external priors, runs at 15.4 frames per second with 7.4 gigabytes of peak GPU memory, and cannot fully certify that the frozen upstream models never saw the evaluation datasets during their own pretraining. Low-quality contour priors naturally occur on roughly 6 to 12 percent of target images. Even so, the study demonstrates that thoughtfully composing existing foundation models can beat monolithic systems on strict metric transfer, suggesting a future where visual AI advances by orchestrating specialized experts rather than simply scaling up single networks.</p>
<p><strong>Subject of Research:</strong> Zero-shot metric monocular depth estimation via fusion of frozen foundation visual priors</p>
<p><strong>Article Title:</strong> FuseDepth: Zero-shot metric depth with semantic–geometric fusion of foundation visual priors</p>
<p><strong>Article References:</strong> Javidnia, H. (2026). FuseDepth: Zero-shot metric depth with semantic–geometric fusion of foundation visual priors. <em>Machine Learning with Applications, 26</em>, Article 101010. <a href="https://doi.org/10.1016/j.mlwa.2026.101010" rel="noopener noreferrer">https://doi.org/10.1016/j.mlwa.2026.101010</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.mlwa.2026.101010" rel="noopener noreferrer">10.1016/j.mlwa.2026.101010</a></p>
<p><strong>Keywords:</strong> monocular depth estimation, zero-shot learning, metric depth, computer vision, foundation models, Depth Anything v2, image matting, surface normals, LoRA, machine learning, depth discontinuities, prior fusion</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">215787</post-id>	</item>
		<item>
		<title>DuoDiT: A Dual-Stream Trick That Fine-Tunes Giant Image Generators by Training Just 2.6% of Their Weights</title>
		<link>https://scienmag.com/duodit-a-dual-stream-trick-that-fine-tunes-giant-image-generators-by-training-just-2-6-of-their-weights/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Fri, 25 Sep 2026 22:41:34 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[cost reduction in generative image models]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[diffusion transformer model adaptation]]></category>
		<category><![CDATA[diffusion transformers]]></category>
		<category><![CDATA[dual-stream architecture]]></category>
		<category><![CDATA[dual-stream architecture for AI models]]></category>
		<category><![CDATA[DuoDiT]]></category>
		<category><![CDATA[efficient training of large AI models]]></category>
		<category><![CDATA[FID]]></category>
		<category><![CDATA[generative AI]]></category>
		<category><![CDATA[high-quality generative imagery cost optimization]]></category>
		<category><![CDATA[image generation]]></category>
		<category><![CDATA[image generation with minimal parameter updates]]></category>
		<category><![CDATA[ImageNet]]></category>
		<category><![CDATA[large diffusion transformer model fine-tuning]]></category>
		<category><![CDATA[LoRa]]></category>
		<category><![CDATA[noise removal in diffusion models]]></category>
		<category><![CDATA[parameter-efficient fine-tuning]]></category>
		<category><![CDATA[parameter-efficient image generation]]></category>
		<category><![CDATA[reducing computational costs in AI image generation]]></category>
		<category><![CDATA[scalable attention mechanisms in image synthesis]]></category>
		<category><![CDATA[transformer-based image synthesis techniques]]></category>
		<category><![CDATA[Vision Transformers]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=215056</guid>

					<description><![CDATA[Researchers at Tarbiat Modares University have developed DuoDiT, a dual-stream architecture that adapts large diffusion transformers for image generation by training only 2.6 percent of the model's parameters while achieving state-of-the-art perceptual quality.]]></description>
										<content:encoded><![CDATA[<p>A team of computer scientists at Tarbiat Modares University in Tehran has unveiled a new architecture that promises to make one of the most expensive operations in modern artificial intelligence dramatically cheaper. In a study published in Multimedia Tools and Applications, Mostafa Shahbazi Dil, Mohammad Mahmoudabadi, and Mansoor Rezghi introduce DuoDiT, a parameter-efficient dual-stream design that adapts large diffusion transformer models for image generation while leaving the overwhelming majority of the network&#8217;s parameters untouched. The work arrives at a moment when the appetite for high-quality generative imagery is exploding, but the computational cost of customizing the underlying models remains a stubborn bottleneck for research labs and companies alike.</p>
<p>To understand why DuoDiT matters, it helps to recall how contemporary image generators actually work. Diffusion models learn to create images by iteratively removing noise from a random starting point, gradually sculpting a clean picture out of static. The most powerful recent versions of these systems replace the older convolutional backbones with transformers, the same architecture that revolutionized language modeling. These Diffusion Transformers, or DiTs, chop a latent representation of the image into small patches, treat each patch as a token, and process the sequence with scalable attention mechanisms. The result is state-of-the-art fidelity, but also enormous models: the DiT-XL/2 backbone that DuoDiT builds upon contains roughly 690 million parameters.</p>
<p>The problem the Iranian researchers set out to solve concerns adaptation rather than initial training. Training a diffusion transformer from scratch costs millions of dollars of compute, so practitioners increasingly rely on parameter-efficient fine-tuning, or PEFT, techniques that freeze the pre-trained backbone and update only a small subset of weights. Methods such as LoRA and DiffFit have shown that adding low-rank updates or tuning selected layers can customize a model at a fraction of the usual cost. Yet, as the DuoDiT authors point out, existing PEFT approaches for diffusion transformers share a structural limitation: they operate at the same compressed token resolution as the frozen backbone, which constrains their ability to recover fine texture and boundary details without updating a large number of parameters.</p>
<p>DuoDiT&#8217;s central idea is elegantly simple. Instead of modifying the frozen backbone at its own coarse resolution, the architecture adds a second, trainable stream that looks at the same noisy latent input through a much finer lens. This auxiliary stream uses a smaller patch size, producing a denser grid of tokens that capture localized visual information the coarse backbone tokens tend to blur over. The design echoes the spirit of side-tuning and multi-scale vision transformers, but applies it to the denoising process of a diffusion model, where detail recovery during generation is precisely what separates adequate images from striking ones.</p>
<p>The technical machinery of the second stream is where the paper&#8217;s engineering shows through. Because the fine-grained stream produces many more tokens than the frozen DiT, naively fusing the two representations would be computationally prohibitive. DuoDiT instead employs local CLS-token aggregation: fine-grained tokens are grouped, and a learnable summary token distills each group into a single aligned detail token. These aggregated detail tokens are then fused back into the frozen DiT representation just before the output head. In effect, the backbone continues to supply its powerful global denoising prior, the statistical knowledge of how images should look that it acquired during pre-training, while the second stream injects content-aware local refinements exactly where they are needed.</p>
<p>The numbers reported in the study are striking. On ImageNet-1K class-conditional generation at 256 by 256 resolution, the DuoDiT-200 configuration updates only 17.95 million of the backbone&#8217;s 690.39 million parameters, a mere 2.6 percent of the total. Despite this tiny trainable footprint, the model achieves a Fréchet Inception Distance, or FID, of 2.219, a Kernel Inception Distance of 0.0014, and an LPIPS-Mean of 0.7162. FID measures how statistically similar generated images are to real ones, with lower being better, while KID and LPIPS probe distributional fidelity and perceptual quality from complementary angles. Among all the methods evaluated in the paper, DuoDiT obtained the lowest KID and LPIPS-Mean scores while maintaining a competitive FID, suggesting that its gains are concentrated exactly where the dual-stream design predicts: in perceptual detail and distributional precision.</p>
<p>Automated metrics only tell part of the story, and the authors supplemented them with a human evaluation. In a preference study comparing DuoDiT against LightningDiT, an established efficient diffusion transformer baseline, human raters chose DuoDiT&#8217;s outputs in 53.26 percent of non-tie responses. That margin may appear modest, but in the world of image generation benchmarks, where differences between top systems are often imperceptible, a consistent human preference achieved with a fraction of the trainable parameters is a meaningful result. The study&#8217;s ethics declarations note that participation was voluntary and fully anonymized, with informed consent collected from all subjects.</p>
<p>The broader significance of DuoDiT lies in what it implies for the economics of generative AI. As diffusion transformers scale up to power photorealistic text-to-image systems, video generation, and multimodal assistants, the ability to adapt them cheaply becomes a strategic advantage. A technique that preserves a frozen backbone&#8217;s learned knowledge while adding a lightweight, task-specific refinement path means that multiple specializations can share a single expensive pre-trained model. It also lowers the barrier for academic groups and smaller organizations that cannot afford full fine-tuning runs, potentially diversifying who gets to push the frontier of image synthesis. The DuoDiT authors have released their data and implementation openly on GitHub, a move that should accelerate adoption and follow-up research.</p>
<p>The paper also situates itself within a rapidly evolving lineage. Diffusion models first overtook generative adversarial networks in image quality in 2021, latent diffusion made high-resolution synthesis practical, and vision transformers supplied the scalable backbone that modern systems like PixArt and video diffusion models now depend on. Parameter-efficient fine-tuning, meanwhile, migrated from natural language processing into vision and diffusion settings through methods like LoRA, SVDiff, and DiffFit. DuoDiT&#8217;s contribution to this lineage is architectural rather than merely algorithmic: rather than squeezing more efficiency out of the same token resolution, it changes the resolution at which adaptation happens, arguing that fine-grained detail recovery demands fine-grained tokens.</p>
<p>There are, of course, caveats worth keeping in mind. The reported results are for class-conditional generation on ImageNet-1K, a well-controlled benchmark that differs from the open-ended text-to-image setting of commercial systems, and the auxiliary stream still adds inference-time computation even if its trainable parameter count is small. The authors report no external funding and declare no competing interests, and the work was published in the journal&#8217;s special collection on rich media with generative AI. Whether the dual-stream principle transfers to text-conditioned and video-scale diffusion transformers remains an open question that the community will now be eager to test. For now, DuoDiT offers a compelling demonstration that in generative AI, as in so much of science, knowing precisely where to add a little capacity can matter far more than adding a lot of it everywhere.</p>
<p><strong>Subject of Research:</strong> Parameter-efficient fine-tuning of diffusion transformers for image generation using a dual-stream architecture</p>
<p><strong>Article Title:</strong> DuoDiT: a parameter-efficient dual-stream architecture for image generation in diffusion transformers</p>
<p><strong>Article References:</strong> Shahbazi Dil, M., Mahmoudabadi, M., &amp; Rezghi, M. (2026). DuoDiT: a parameter-efficient dual-stream architecture for image generation in diffusion transformers. <em>Multimedia Tools and Applications, 85</em>(10), Article 780. <a href="https://doi.org/10.1007/s11042-026-21926-y" rel="noopener noreferrer">https://doi.org/10.1007/s11042-026-21926-y</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11042-026-21926-y" rel="noopener noreferrer">10.1007/s11042-026-21926-y</a></p>
<p><strong>Keywords:</strong> diffusion transformers, image generation, parameter-efficient fine-tuning, DuoDiT, vision transformers, generative AI, deep learning, computer vision, LoRA, ImageNet, FID, dual-stream architecture</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">215056</post-id>	</item>
		<item>
		<title>Tiny Adapters, Big Results: LoRA Matches Full Fine-Tuning Across Sentiment Tasks</title>
		<link>https://scienmag.com/tiny-adapters-big-results-lora-matches-full-fine-tuning-across-sentiment-tasks/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Fri, 25 Sep 2026 01:51:19 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AdaLoRA]]></category>
		<category><![CDATA[aspect-based sentiment analysis]]></category>
		<category><![CDATA[BERT]]></category>
		<category><![CDATA[BERT sentiment classification]]></category>
		<category><![CDATA[DeBERTa]]></category>
		<category><![CDATA[emotion detection]]></category>
		<category><![CDATA[energy-efficient NLP]]></category>
		<category><![CDATA[impact of LoRA on NLP tasks]]></category>
		<category><![CDATA[large pretrained language models]]></category>
		<category><![CDATA[LoRa]]></category>
		<category><![CDATA[Low-Rank Adaptation (LoRA)]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[model adaptation techniques]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[parameter-efficient fine-tuning]]></category>
		<category><![CDATA[resource-efficient model training]]></category>
		<category><![CDATA[RoBERTa]]></category>
		<category><![CDATA[sentiment analysis]]></category>
		<category><![CDATA[sentiment classification framework]]></category>
		<category><![CDATA[systematic evaluation of fine-tuning methods]]></category>
		<category><![CDATA[transformer-based models]]></category>
		<category><![CDATA[transformers]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=214019</guid>

					<description><![CDATA[A new systematic study shows that low-rank adaptation can match or approach full fine-tuning of transformer models across four sentiment analysis paradigms while updating up to 99.8 percent fewer parameters.]]></description>
										<content:encoded><![CDATA[<p>Sentiment analysis has quietly become one of the most commercially and scientifically important applications of modern natural language processing. Every product review, tweet, and customer support ticket is a potential signal about how people feel, and companies and researchers alike have poured resources into transformer-based models that can decode those signals. But there is a catch: fully fine-tuning a large pretrained language model such as BERT, RoBERTa, or DeBERTa means updating hundreds of millions of parameters, which demands serious GPU memory, hours of training time, and considerable energy. A new study published in the journal Machine Learning argues that most of that effort may be unnecessary. The work, led by Md. Easin Arafat and Muhammad Usman Akmal of Eötvös Loránd University in Budapest together with colleagues, introduces SentiMatrix, a systematic evaluation framework showing that parameter-efficient fine-tuning can deliver competitive sentiment classification while updating less than two percent of a model&#8217;s parameters.</p>
<p>The core technique under scrutiny is Low-Rank Adaptation, or LoRA, a method first proposed in 2022 that has since become the workhorse of efficient model adaptation. Instead of adjusting every weight in a pretrained network, LoRA freezes the original weights entirely and injects small pairs of low-rank matrices into the transformer&#8217;s attention layers. Mathematically, the update to a weight matrix is constrained to the product of two much smaller matrices, so the number of trainable parameters shrinks from the full matrix size to a fraction determined by a rank hyperparameter. In the SentiMatrix experiments, this translated into reductions of trainable parameters by up to 99.8 percent compared with full fine-tuning, alongside lower GPU memory consumption and shorter training times in most evaluated settings.</p>
<p>What distinguishes SentiMatrix from earlier work is its breadth. Previous studies typically tested efficient adaptation on a single sentiment task or dataset, making it impossible to know whether the conclusions generalized. The Hungarian team instead evaluated four distinct sentiment analysis paradigms: intent-based classification, which captures a speaker&#8217;s underlying attitude in binary and three-class settings; aspect-based sentiment analysis, which assigns polarity to specific entities mentioned in a review; fine-grained classification on five-star rating scales; and emotion detection across six discrete affective categories such as joy, sadness, and fear. Seven benchmark datasets were used, including SST-2 and IMDb for binary polarity, the Twitter US Airline Sentiment corpus, the SemEval-2014 laptop and restaurant benchmarks, Yelp and Amazon e-commerce reviews, and the CARER emotion dataset, all under a consistent three-stage protocol.</p>
<p>That protocol compared three conditions. First, task-specific pretrained baselines from HuggingFace were run in inference mode, serving as informed upper-bound references rather than cold-start baselines. Second, full fine-tuning of the base architectures established the performance ceiling, at the cost of updating every parameter and consuming between roughly 2.0 and 4.2 gigabytes of peak GPU memory with training times reaching over 1,800 seconds on the longest tasks. Third, LoRA adaptation was applied to the same architectures, with Adaptive LoRA, or AdaLoRA, benchmarked as a direct comparator. AdaLoRA differs from standard LoRA by dynamically allocating its parameter budget across weight matrices using importance scores derived from singular value decomposition, pruning the least informative components as training proceeds.</p>
<p>The headline numbers are striking. LoRA reached 93.28 percent accuracy on SST-2 with RoBERTa, actually beating full fine-tuning by 2.37 percentage points while cutting trainable parameters from about 124.65 million to 1.20 million and reducing training time by roughly half. On the combined laptop and restaurant aspect-based benchmark with DeBERTa-v3, LoRA achieved 80.74 percent accuracy against 88.04 percent for full fine-tuning, while updating only 0.33 million parameters instead of 184.76 million. On the five-class e-commerce task, LoRA&#8217;s 67.63 percent came remarkably close to full fine-tuning&#8217;s 68.47 percent, and on six-class emotion detection with DistilBERT, LoRA scored 93.22 percent, essentially matching the full model&#8217;s 93.30 percent.</p>
<p>The results were not uniformly favorable, and the authors are candid about where parameter efficiency falls short. On the domain-shifted Twitter three-class task, full fine-tuning reached 93.64 percent while LoRA managed only 72.84 percent, the largest gap in the study. On the Yelp five-star dataset, AdaLoRA outperformed both LoRA and full fine-tuning, achieving 59.35 percent accuracy. The researchers offer an intuitive explanation for these patterns: LoRA&#8217;s fixed low-rank constraint acts as an implicit regularizer, forcing the model to capture only the most essential directional changes in weight space rather than memorizing task-specific noise. This helps most when the training data is noisy or label boundaries are ambiguous, as in social media text and fine-grained star ratings, but becomes a limitation when substantial domain adaptation is required.</p>
<p>One of the most practically interesting findings concerns label granularity. When the researchers collapsed five sentiment classes into three, merging one- and two-star reviews as negative, three stars as neutral, and four and five stars as positive, accuracy jumped dramatically, with LoRA gaining 26.98 percentage points on Yelp and 17.33 points on e-commerce data. The authors caution that this gap reflects a combination of reduced label ambiguity and a change in backbone architecture between the two settings, so it should not be attributed to granularity alone. Still, the result carries a clear message for practitioners: distinguishing adjacent ordinal categories, such as a three-star versus four-star review, may demand modeling subtleties that lie largely outside the vocabulary-level signal available to contextualized encoders.</p>
<p>The study also went beyond hard-decision accuracy by reporting calibration-oriented metrics. A confidence score measured the mean softmax probability assigned to the correct class, while a similarity score computed the cosine similarity between the predicted probability distribution and the one-hot ground truth. Across tasks, LoRA&#8217;s confidence and similarity scores were competitive with or superior to full fine-tuning, indicating that parameter-efficient adaptation preserves the quality of probability distributions, not merely the correctness of discrete predictions. This matters for real deployments, where downstream systems often rely on graded confidence rather than raw class labels.</p>
<p>The authors are equally transparent about limitations. The evaluation covers only English-language datasets and encoder-based architectures, so the findings may not extend to decoder-only generative large language models, whose adaptation dynamics can differ substantially. The LoRA and AdaLoRA comparisons used closely matched but not strictly identical parameter budgets, and the data splits for fine-grained datasets were not stratified by rating class, which may affect minority categories. Efficiency metrics reflect training time only; inference-time costs depend on whether adapters are merged into the base model before serving. Future work, the team says, will extend the framework to generative models such as LLaMA and Mistral, explore multilingual and low-resource settings, and run strictly controlled rank-sensitivity experiments.</p>
<p>Even with those caveats, SentiMatrix arrives at a moment when the economics of AI adaptation are under intense scrutiny. Training and retraining large models carries financial and environmental costs that scale with parameter counts, and the demonstration that a fraction of a percent of trainable parameters can nearly match full fine-tuning across four sentiment paradigms is a meaningful data point. For organizations that must update sentiment models frequently, across multiple domains and label schemes, the study suggests that LoRA-style adaptation is a practical default, reserving full fine-tuning for cases of severe domain shift. The source code and datasets are publicly available on GitHub, inviting the community to replicate and extend a benchmark that may well become a reference point for efficient adaptation of language models.</p>
<p><strong>Subject of Research:</strong> Parameter-efficient fine-tuning of encoder-based transformer models for multidimensional sentiment analysis</p>
<p><strong>Article Title:</strong> SentiMatrix: Parameter-Efficient Fine-Tuning of Encoder-Based Transformers for Multidimensional Sentiment Analysis</p>
<p><strong>Article References:</strong> Arafat, M. E., Akmal, M. U., Abosinnee, A. S., &amp; Orosz, T. (2026). SentiMatrix: Parameter-Efficient Fine-Tuning of Encoder-Based Transformers for Multidimensional Sentiment Analysis. <em>Machine Learning, 115</em>(10), Article 229. <a href="https://doi.org/10.1007/s10994-026-07161-4" rel="noopener noreferrer">https://doi.org/10.1007/s10994-026-07161-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10994-026-07161-4" rel="noopener noreferrer">10.1007/s10994-026-07161-4</a></p>
<p><strong>Keywords:</strong> sentiment analysis, LoRA, parameter-efficient fine-tuning, transformers, AdaLoRA, natural language processing, BERT, RoBERTa, DeBERTa, emotion detection, aspect-based sentiment analysis, machine learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">214019</post-id>	</item>
		<item>
		<title>AI Learns to Read Kurdish News Stance with Just 2,174 Articles</title>
		<link>https://scienmag.com/ai-learns-to-read-kurdish-news-stance-with-just-2174-articles/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Fri, 25 Sep 2026 01:29:43 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[automatic stance classification in Kurdish]]></category>
		<category><![CDATA[benchmarking]]></category>
		<category><![CDATA[class imbalance]]></category>
		<category><![CDATA[data augmentation]]></category>
		<category><![CDATA[data engineering for low-resource languages]]></category>
		<category><![CDATA[empirical benchmark for Kurdish NLP]]></category>
		<category><![CDATA[fine-tuning]]></category>
		<category><![CDATA[fine-tuning language models for Kurdish]]></category>
		<category><![CDATA[KuBERT]]></category>
		<category><![CDATA[Kurdish news bias detection]]></category>
		<category><![CDATA[Kurdish news sentiment analysis]]></category>
		<category><![CDATA[Kurdish news stance detection]]></category>
		<category><![CDATA[LoRa]]></category>
		<category><![CDATA[low-resource language NLP]]></category>
		<category><![CDATA[low-resource languages]]></category>
		<category><![CDATA[macro F1]]></category>
		<category><![CDATA[misinformation and polarized discourse analysis]]></category>
		<category><![CDATA[multilingual NLP development in Iraq]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[QLoRA]]></category>
		<category><![CDATA[Sorani Kurdish]]></category>
		<category><![CDATA[Sorani Kurdish natural language processing]]></category>
		<category><![CDATA[stance detection]]></category>
		<category><![CDATA[stance detection models for Kurdish]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=213859</guid>

					<description><![CDATA[Researchers report the first trained stance-detection models and benchmark for Sorani Kurdish, using data augmentation and QLoRA fine-tuning of KuBERT to reach a macro F1 of 0.58 on a leakage-free test protocol.]]></description>
										<content:encoded><![CDATA[<p>Researchers in Sulaimani, in the Kurdistan Region of Iraq, have built the first trained stance-detection models and empirical benchmark for Sorani Kurdish news, one of the most widely spoken varieties of Kurdish yet long absent from the map of modern natural language processing. Their study, published in the International Journal of Data Science and Analytics, shows how far careful data engineering and efficient fine-tuning of a language model can go when a language has almost no annotated data to learn from. The work was carried out by Hawar Hussein Yaba of the Kurdistan Technical Institute together with Rebwar M. Nabi and Rebaz M. Nabi of Sulaimani Polytechnic University and the Raparin Technical and Vocational Institute.</p>
<p>Stance detection is the task of automatically determining whether a piece of text is in favor of, against, or neutral toward a target such as a political claim, a public figure, or a news event. It is a cornerstone technology for studying misinformation, polarized discourse, and the shape of public debate online. For high-resource languages like English, mature models and large benchmark datasets exist, and recent research has pushed toward multimodal and zero-shot approaches. For Sorani Kurdish, however, the authors report that no trained stance-detection model or benchmark had previously been published at all, despite the recent release of a small annotated dataset known as the Bochun Kurdish Stance Detection Dataset. That mismatch between an available resource and any applied modeling is the gap the new study set out to close.</p>
<p>The starting point was a public dataset of 2,174 annotated Sorani news articles, hosted on Mendeley Data. That number is tiny by the standards of modern machine learning, where models routinely train on tens of thousands or millions of labeled examples. Worse, the dataset suffered from severe class imbalance, meaning the three stance categories were far from equally represented, a condition that notoriously causes classifiers to ignore minority classes and inflate their apparent accuracy. The team therefore framed their central question as how to build reliable stance detectors under both extreme data scarcity and skewed label distributions.</p>
<p>Their answer combined two lines of attack. The first was a contextual data augmentation pipeline that expanded the training corpus into a strictly filtered, class-balanced set. At its heart lies masked-token substitution powered by KuBERT, a BERT language model pre-trained specifically for central Kurdish. In this technique, words in training sentences are replaced with a mask token, and the language model predicts plausible substitutes that fit the surrounding context, generating new sentences that stay grammatically and semantically close to the originals. This is a low-resource adaptation of contextual augmentation, an approach introduced at NAACL in 2018, and the team paired it with controlled oversampling so that each stance class contributed a balanced share of training examples. Crucially, the augmentation was subjected to strict filtering, and human Kurdish-language annotators validated a sample of the generated text to confirm that the procedure preserved the original labels, an assumption the authors say underpins the entire approach.</p>
<p>The second line of attack was the fine-tuning strategy itself. Rather than training models from scratch, the researchers adapted KuBERT, the Kurdish BERT model released in 2024, under three distinct regimes. Full-parameter updating adjusts every weight in the network and typically demands the most compute and memory. Low-rank adaptation, or LoRA, freezes the base weights and instead learns small low-rank matrices injected into the transformer&#8217;s layers, cutting the number of trainable parameters dramatically. QLoRA goes a step further by quantizing the frozen base model to 4-bit precision before applying LoRA, a configuration popularized by the 2023 QLoRA paper for efficient fine-tuning of large language models. Comparing these three strategies on the same data provided a rare empirical head-to-head in a genuinely low-resource setting.</p>
<p>Evaluation methodology received as much attention as the models themselves. The team ran all experiments across five random seeds, a safeguard against the luck of any single training run, and scored everything on a held-out test set drawn exclusively from real, unaugmented articles. This distinction matters: evaluating on synthetic text can silently leak the very patterns augmentation introduces, so testing only on genuine journalism gives a more honest picture of real-world performance. Macro-averaged F1 served as the primary metric because it weights each stance class equally and thus exposes failure on minority classes, with accuracy, weighted F1, the Matthews Correlation Coefficient, and Cohen&#8217;s kappa reported alongside it. MCC, in particular, is prized for giving a truthful single number even when class distributions are unbalanced.</p>
<p>The results tell a clear story about why task-specific adaptation matters. A zero-shot KuBERT baseline, applied to stance detection without any fine-tuning, scored a macro F1 of roughly 0.20 with an MCC near negative 0.03, meaning it performed barely better than random guessing. After fine-tuning on the original, imbalanced dataset, models reached a macro F1 of around 0.48 with an MCC of about 0.24, more than doubling the baseline and confirming that domain adaptation is essential even when data is scarce. But the most striking gains appeared on the class-balanced augmented condition, evaluated under a corrected, leakage-free protocol. There, the best configuration, KuBERT fine-tuned with QLoRA, achieved a mean macro F1 of 0.58 and an MCC of 0.37 across seeds, with the strongest single run reaching a macro F1 of 0.60.</p>
<p>Perhaps the most instructive finding came from the per-class and ablation analysis, which disentangled where the improvement actually came from. The authors report that most of the gain was driven by correcting class imbalance rather than by the diversity of augmented examples alone. In other words, balancing the classes was the dominant lever, with contextual augmentation contributing as the mechanism that made balancing possible without exhausting the real data. This nuance carries a practical lesson for anyone building classifiers on small, skewed datasets in low-resource languages: the expensive machinery of augmentation and parameter-efficient fine-tuning pays off most when paired with a disciplined focus on label distribution and honest evaluation.</p>
<p>The choice of QLoRA as the winning configuration also has practical implications beyond accuracy. Because QLoRA trains only tiny adapter modules on top of a 4-bit quantized base model, it slashes memory requirements, putting fine-tuning of pretrained language models within reach of research groups without access to large-scale computing infrastructure. For language communities outside the technological mainstream, that accessibility may prove as consequential as the benchmark numbers themselves. The team&#8217;s augmentation pipeline, fine-tuning scripts, model configuration files, and per-seed evaluation results are available from the corresponding author on reasonable request, and the underlying Bochun dataset is public, lowering the barrier for follow-up work.</p>
<p>The broader significance of the study lies in what it demonstrates for the roughly tens of millions of Sorani speakers whose media landscape has been effectively invisible to computational text analysis. Reliable stance detection could enable systematic study of how misinformation spreads through Kurdish news and social media, inform fact-checking efforts, and support media-monitoring tools tuned to the region&#8217;s discourse. The authors are explicit that their contribution is a first competitive baseline rather than a solved problem: a macro F1 of 0.58, while a substantial step above chance, still leaves considerable room for improvement, and future work will likely explore larger corpora, additional architectures, and zero-shot techniques now emerging for stance detection more broadly. But as a proof of concept, the study shows that a small annotated dataset, a Kurdish language model, and a leakage-free, class-aware evaluation protocol can together close a meaningful portion of the gap between low-resource languages and the cutting edge of natural language understanding.</p>
<p><strong>Subject of Research:</strong> Stance detection for low-resource Sorani Kurdish news using data augmentation and parameter-efficient fine-tuning of a Kurdish BERT model</p>
<p><strong>Article Title:</strong> Towards robust stance detection: data augmentation, model fine-tuning, and empirical benchmarking</p>
<p><strong>Article References:</strong> Yaba, H. H., Nabi, R. M., &amp; Nabi, R. M. (2026). Towards robust stance detection: data augmentation, model fine-tuning, and empirical benchmarking. <em>International Journal of Data Science and Analytics, 22</em>(1), Article 310. <a href="https://doi.org/10.1007/s41060-026-01266-8" rel="noopener noreferrer">https://doi.org/10.1007/s41060-026-01266-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s41060-026-01266-8" rel="noopener noreferrer">10.1007/s41060-026-01266-8</a></p>
<p><strong>Keywords:</strong> stance detection, Sorani Kurdish, KuBERT, data augmentation, QLoRA, LoRA, fine-tuning, low-resource languages, natural language processing, class imbalance, macro F1, benchmarking</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">213859</post-id>	</item>
		<item>
		<title>AI System Turns Sketches Drawn in Mid-Air Into Clay-Style Art</title>
		<link>https://scienmag.com/ai-system-turns-sketches-drawn-in-mid-air-into-clay-style-art/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 22:35:24 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[3D sketch gesture transformation]]></category>
		<category><![CDATA[air drawing]]></category>
		<category><![CDATA[air-doodle to art conversion]]></category>
		<category><![CDATA[air-drawing recognition]]></category>
		<category><![CDATA[BiLSTM attention]]></category>
		<category><![CDATA[clay-style image synthesis]]></category>
		<category><![CDATA[creative AI]]></category>
		<category><![CDATA[deep learning for freeform sketching]]></category>
		<category><![CDATA[generative adversarial networks]]></category>
		<category><![CDATA[gesture-based image creation]]></category>
		<category><![CDATA[human-computer interaction]]></category>
		<category><![CDATA[image-to-image translation]]></category>
		<category><![CDATA[innovation in freehand digital art tools]]></category>
		<category><![CDATA[live air-drawing capture technology]]></category>
		<category><![CDATA[LoRa]]></category>
		<category><![CDATA[MediaPipe]]></category>
		<category><![CDATA[multi-stage AI framework for gesture to image]]></category>
		<category><![CDATA[natural human-computer interaction for creative expression]]></category>
		<category><![CDATA[Pix2Pix]]></category>
		<category><![CDATA[sketch recognition]]></category>
		<category><![CDATA[Stable Diffusion XL]]></category>
		<category><![CDATA[stylized claymation art generation]]></category>
		<category><![CDATA[temporal deep learning]]></category>
		<category><![CDATA[temporal sketch recognition systems]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=208387</guid>

					<description><![CDATA[Researchers have built Inklude, a two-stage AI framework that recognizes sketches drawn in mid-air and transforms them into stylized clay-like images in real time without text prompts.]]></description>
										<content:encoded><![CDATA[<p>Imagine a child waving a finger through the air and watching a rough, invisible doodle bloom into a polished claymation-style picture on screen. That vision now has a concrete technical foundation. Researchers have unveiled Inklude, a two-stage deep learning framework that captures free-form drawings made in three-dimensional space, recognizes what the user intended to draw, and then transforms the gesture into a stylized, clay-like image, all without a single text prompt or a physical drawing surface. The work, published in Machine Learning with Applications, is among the first attempts to unify live air-drawing acquisition, temporal sketch recognition, and recognition-conditioned image synthesis in one deployment-oriented pipeline.</p>
<p>The problem the researchers set out to solve is deceptively simple to describe but hard to engineer. Children, in particular, often conceive vivid mental images yet lack the fine motor control and drawing experience needed to render them on paper, creating a frustrating gap between creative intent and visual execution. Most modern generative systems, including text-to-image giants like DALL·E and Stable Diffusion, sidestep the body entirely: they rely on typed descriptions or pre-existing raster images. Inklude instead treats drawing as what it fundamentally is, a motion process, capturing fingertip trajectories in the air and preserving the temporal structure of every stroke rather than collapsing the result into a static pixel grid.</p>
<p>At the heart of the system is a trajectory encoding the authors call the Motion Language Matrix, or MLM. Each air-drawn sketch is modeled as a sequence of spatial coordinates paired with a pen-state variable indicating whether the fingertip is actively drawing or transitioning between strokes. Consecutive positional displacements are computed, normalized, and padded or truncated to a fixed length of 128 time steps, producing a compact 128-by-3 matrix in which the three channels represent relative horizontal motion, relative vertical motion, and stroke continuity. Crucially, the team subjected this encoding to unusually honest scrutiny: a controlled, element-wise comparison revealed that MLM is byte-identical to the standard normalized offset-plus-pen-state and stroke-3 representations already familiar from the Sketch-RNN literature. The researchers explicitly label MLM an implementation name rather than a novel contribution, a level of transparency that stands out in a field often criticized for rebranding existing techniques.</p>
<p>Recognition is performed by a hybrid Conv1D–BiLSTM–Attention network. A one-dimensional convolutional layer first extracts localized motion descriptors, capturing short-range directional transitions, curvature changes, and stroke continuity patterns. A bidirectional long short-term memory layer then models dependencies across the entire trajectory in both temporal directions, a design choice that helps interpret incomplete or evolving sketches. Finally, an attention mechanism learns to weight the most semantically informative temporal frames, allowing the classifier to emphasize discriminative stroke transitions while suppressing noisy or redundant movements. Evaluated on 225,000 sketches spanning 45 child-relevant categories from Google&#8217;s Quick, Draw! dataset, using a rigorous five-rotation protocol in which every sketch served exactly once as a test sample, the classifier achieved a mean Top-1 accuracy of 87.66 percent with a tight standard deviation of 0.13 percent. Interestingly, the same controlled comparison showed that simple absolute coordinates with pen state actually outperformed the offset-based encoding by 1.70 percentage points, a finding the authors report without equivocation.</p>
<p>The second stage of Inklude handles stylized generation, and here the team confronted a genuine data bottleneck: no public benchmark exists for supervised sketch-to-clay translation. Their solution was to construct one. Using Stable Diffusion XL enhanced with a Low-Rank Adaptation module fine-tuned on 500 curated claymation reference images, they generated paired clay-style targets for 16 object categories, from palm trees and birthday cakes to cats and snowmen, yielding 16,000 sketch-image pairs in total. Human validators checked a random 10 percent subset per category for semantic correctness and stylistic consistency. On top of this synthetic corpus, the researchers trained category-specific Pix2Pix conditional generative adversarial networks, each pairing a U-Net generator with skip connections and a PatchGAN discriminator that judges realism on local 70-by-70 pixel patches. The training objective jointly balances adversarial realism, L1 reconstruction, and VGG-based perceptual loss.</p>
<p>The evaluation of the generation stage produced a nuance worth savoring. Measured against the synthetic SDXL+LoRA targets, Pix2Pix posted modest reconstruction scores, with an SSIM of 0.55 and an FID of 280, while unpaired methods like CycleGAN and MUNIT scored far higher on similarity to reference pixels. Yet when twenty blinded human raters judged single images for semantic recognizability and perceived clay-style quality, Pix2Pix came out on top, earning a mean recognizability rating of 4.13 out of 5 against 3.73 for CycleGAN and a dismal 1.02 for SketchyGAN. A reference-free CLIP analysis corroborated the human verdicts on the clay-versus-sketch margin. The lesson is methodologically important: in stylized synthesis, pixel-level fidelity to a synthetic target can be a poor proxy for what people actually perceive, and the authors wisely treat these endpoints as separate descriptive signals rather than forcing them into a single ranking.</p>
<p>The team also confronted the messiness of real-world deployment head-on. Twenty adult participants each drew the same 16 categories three times through a MediaPipe-based air-drawing interface, generating 960 evaluation trials with a frozen classifier and frozen generators. The results were sobering: Top-1 recognition accuracy dropped to 57.3 percent on real air trajectories, compared with 97.1 percent on matched Quick, Draw! data. Compounding the problem, only 16 of the classifier&#8217;s 45 output categories had trained generators, so 37.4 percent of all trials routed to labels with no available model. An oracle-versus-predicted routing analysis quantified the consequence: estimated intended-category success fell from 95.0 percent under oracle routing to 55.0 percent when the system followed its own predictions. The authors identify this recognition domain gap and limited generator coverage, not generator fidelity, as the dominant determinants of system-level success, and they recommend confidence-aware abstention and user confirmation as practical remedies.</p>
<p>On the latency front, the numbers are encouraging for interactivity. In a repeated CPU benchmark on an Apple M4 machine, the sum of six instrumented post-load stages, spanning encoding, classification, routing, rasterization, generation, and post-processing, averaged 155.4 milliseconds for the 637 trials that reached an available generator, with generation itself consuming roughly 82 percent of that time. The authors are careful to scope the claim precisely: camera acquisition, MediaPipe tracking, trajectory loading, and unmeasured interstage overhead were excluded, so this is not a full end-to-end wall-clock measurement, and it characterizes only the tested hardware. Still, a sub-200-millisecond processing window for the computational core suggests that gesture-driven creative loops are feasible on consumer devices without resorting to heavyweight diffusion inference at run time, which was precisely the deployment motivation for choosing lighter conditional GANs over sketch-conditioned diffusion alternatives.</p>
<p>Cross-dataset testing on the SEVA benchmark, which contains roughly 90,000 sketches of 128 concepts produced under varying time constraints, showed the temporal architecture retaining its relative lead, with the Conv1D–BiLSTM–Attention model reaching 62.00 percent Top-1 accuracy ahead of hybrid RNN-CNN, Transformer, and BiLSTM baselines, though all models suffered in this harder setting dominated by organic, blob-like categories. The authors are candid about the remaining limitations: only 16 categories are supported, extending to Quick, Draw!&#8217;s full 345-class vocabulary would be impractical under the current class-specific design, the clay style is the sole artistic modality, and the 20-participant study involved adults rather than the children the system ultimately aims to serve. Future work points toward universal class-conditional generators, style-conditional architectures spanning watercolor and pixel art, model compression for mobile deployment, and user-centered studies in educational settings.</p>
<p>What makes Inklude compelling beyond its specific numbers is the philosophy it embodies. Rather than asking users to translate imagination into language for a prompt box, the system meets them in the embodied, gesture-driven space where creativity actually begins. The honest accounting of where the pipeline breaks, real trajectories confuse the classifier, uncovered categories produce no output, and synthetic training targets complicate evaluation, makes the work a unusually transparent baseline for the emerging field of embodied generative AI. If the recognition gap can be closed and generator coverage expanded, the loop the researchers describe, motion to meaning to image in a fraction of a second, could reshape how children and novices experience the act of making art, turning the empty air itself into a canvas that understands what you meant to draw.</p>
<p><strong>Subject of Research:</strong> A conditional generative deep learning framework for real-time air-drawing recognition and stylized clay-style image synthesis</p>
<p><strong>Article Title:</strong> Inklude: A Conditional Generative Model for Stylized Air-Drawing Augmentation</p>
<p><strong>Article References:</strong> Singh, S., Kankanala, S. R., Chen, W., &amp; Masum, M. (2026). Inklude: A Conditional Generative Model for Stylized Air-Drawing Augmentation. <em>Machine Learning with Applications</em>, Article 101027. <a href="https://doi.org/10.1016/j.mlwa.2026.101027" rel="noopener noreferrer">https://doi.org/10.1016/j.mlwa.2026.101027</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.mlwa.2026.101027" rel="noopener noreferrer">10.1016/j.mlwa.2026.101027</a></p>
<p><strong>Keywords:</strong> air drawing, sketch recognition, generative adversarial networks, Pix2Pix, Stable Diffusion XL, LoRA, MediaPipe, temporal deep learning, BiLSTM attention, image-to-image translation, human-computer interaction, creative AI</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">208387</post-id>	</item>
		<item>
		<title>AI-Generated Faces Pass the Realness Test in New Emotion Research Toolset</title>
		<link>https://scienmag.com/ai-generated-faces-pass-the-realness-test-in-new-emotion-research-toolset/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 19:58:58 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[AI-generated faces]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[behavioral experiment validation]]></category>
		<category><![CDATA[behavioral research]]></category>
		<category><![CDATA[body weight]]></category>
		<category><![CDATA[ComfyUI]]></category>
		<category><![CDATA[dynamic facial expression synthesis]]></category>
		<category><![CDATA[ecological validity in psychology]]></category>
		<category><![CDATA[emotion perception]]></category>
		<category><![CDATA[emotion perception research]]></category>
		<category><![CDATA[emotion recognition]]></category>
		<category><![CDATA[ethical considerations in face datasets]]></category>
		<category><![CDATA[facial affect datasets]]></category>
		<category><![CDATA[facial expression stimuli]]></category>
		<category><![CDATA[facial expressions]]></category>
		<category><![CDATA[generative AI]]></category>
		<category><![CDATA[generative AI in psychological studies]]></category>
		<category><![CDATA[LoRa]]></category>
		<category><![CDATA[photorealism]]></category>
		<category><![CDATA[photorealistic artificial intelligence]]></category>
		<category><![CDATA[psychology methods]]></category>
		<category><![CDATA[Py-Feat]]></category>
		<category><![CDATA[stimulus validation]]></category>
		<category><![CDATA[validation framework for AI-generated images]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=207691</guid>

					<description><![CDATA[Researchers at the University of Haifa created and validated photorealistic AI-generated facial expression stimuli that over 2,000 participants could not reliably distinguish from real photographs, establishing a dual computational and human validation framework for emotion research.]]></description>
										<content:encoded><![CDATA[<p>For more than four decades, psychological research on how people read emotions from faces has depended on photographs of human actors. From the classic Pictures of Facial Affect compiled by Paul Ekman and Wallace Friesen in 1976 to modern datasets such as the Karolinska Directed Emotional Faces and the Amsterdam Dynamic Facial Expression Set, scientists have relied on posed expressions captured in controlled studio conditions. Now a team at the University of Haifa has demonstrated a rigorous new alternative: facial expression stimuli created entirely with generative artificial intelligence, validated so thoroughly that more than 2,000 study participants could not reliably tell them apart from real photographs.</p>
<p>The research, published in Behavior Research Methods, describes both a method for producing photorealistic AI-generated emotional expressions and a comprehensive validation framework for testing whether such images can stand in for photographs in behavioral experiments. The work, led by Shlomo Hareli and Shlomo David, addresses a long-standing tension in emotion research between experimental control and ecological validity. Traditional photographic datasets are constrained by the actors available, the expressions they can convincingly produce, and the ethical complications of using images of real people. Computational alternatives such as 3D morphable models and FACS-based animation tools offer more control but require expensive specialized software, months of technical training, and often produce faces that fall squarely into the uncanny valley.</p>
<p>Generative AI, the researchers argue, relaxes this trade-off. The method is built on the open-source platform ComfyUI and the Flux.1-dev diffusion model, enhanced with Low-Rank Adaptation, or LoRA, fine-tuning. The authors trained custom LoRA models for each combination of emotion and gender using 15 images per combination drawn, with permission, from the validated FACES database of real facial expressions. Each model was trained for 16 epochs on a single consumer-grade laptop GPU, requiring roughly six hours per emotion-gender pair, with a fixed random seed to guarantee that other laboratories can reproduce the results exactly. Prompt engineering then controlled body weight, clothing, lighting, background, and gaze direction, with weighted attention values emphasizing specific features during generation.</p>
<p>Crucially, the pipeline does not stop at generation. Every image was screened by Py-Feat, a Python library for facial expression analysis, using a strict criterion that the intended emotion had to reach at least a 95 percent classification likelihood before the image was accepted. Images that failed were iteratively refined with a node-based expression editor until they passed. The final Study 1 dataset contained 72 images, evenly balanced across gender, three body weight categories, and three emotional expressions, with a mean intended-emotion confidence of 98.38 percent. All workflows, models, datasets, and analysis scripts were released openly on the Open Science Framework, a move the authors say is central to making the method genuinely reproducible.</p>
<p>The first human validation study recruited 606 US participants through Prolific, each viewing a single image. The results confirmed the manipulations across the board: happy faces were rated significantly happier than neutral ones, which were rated happier than sad ones, and perceived weight rose in clear steps from the thin through the average to the higher-weight categories. Perhaps most strikingly, when asked whether the person in the photo was real or AI-generated, on a scale where 3 meant they could not decide, no condition was clearly identified as artificial. Even the least convincing cell, higher-weight male faces, scored only around 3.25, barely past the uncertainty midpoint.</p>
<p>The age judgments revealed an unexpected form of validation. Happy faces were perceived as younger than sad or neutral ones, and perceived age increased systematically with body weight, patterns that replicate well-documented findings from research on real human photographs. Rather than revealing artificial confounds in the stimuli, these effects suggest the AI generation process captured realistic correlations between facial morphology, body weight, and perceived age, enhancing rather than undermining the ecological validity of the images.</p>
<p>Study 2 extended the framework in two important ways. First, it added anger as a fourth emotion. Second, it adopted a same-character design, in which a single AI-generated identity appears across all four emotional expressions, eliminating potential confounds from differing facial features between conditions. This posed a technical challenge: maintaining identity while altering expression. The team solved it with a cyclic workflow combining an expression editor, LoRA-based inpainting with the identity-preserving PuLID technique, and a face-swapping node, validated computationally with FaceShape software that confirmed 100 percent within-character similarity, benchmarked against the ADFES database. A total of 96 images depicting 24 distinct characters passed both identity and emotion checks, and 736 new participants confirmed that all four emotions and all weight categories were perceived as intended.</p>
<p>A third study, Study 2b, pushed validation beyond simple emotion recognition. With 731 additional participants rating the same 96 images, the researchers examined valence, intensity, naturalness, authenticity, and non-focal emotions. The stimuli behaved like real-expression photographs on every dimension: happiness was rated positive, sadness and anger negative, and neutral faces affectively neutral yet not expressively empty, with measurable intensity. The ordering of naturalness and authenticity ratings, happiness and neutrality highest, anger lowest, mirrored the genuineness hierarchy previously established for posed expressions in real databases. Male faces were rated as angrier than female ones, consistent with known gender stereotypes in emotion perception, and body weight influenced sadness and fear ratings, showing that the stimuli carry meaningful social-category information rather than functioning as abstract emotion displays.</p>
<p>The authors are careful about the limits of their achievement. Photorealism as judged by human observers does not mean AI-generated and real faces are equivalent in every respect; computational algorithms can still detect subtle artifacts invisible to people, and the current framework addresses only static images, whereas dynamic expressions are known to aid emotion recognition. The pipeline also still demands considerable technical skill, and the authors call for future template interfaces to make it accessible without workflow-level expertise. Nevertheless, they conclude that AI-generated stimuli can now serve as a scientifically validated, ethically attractive alternative to photographs, particularly for research on stigmatized characteristics such as body weight, and that the dual computational and human validation framework they established provides a template any laboratory can adapt as generative technology continues to evolve.</p>
<p><strong>Subject of Research:</strong> Creation and validation of photorealistic AI-generated facial expression stimuli for emotion perception research</p>
<p><strong>Article Title:</strong> Creating and validating photorealistic AI-generated facial expression stimuli for emotion research</p>
<p><strong>Article References:</strong> Hareli, S., &amp; David, S. (2026). Creating and validating photorealistic AI-generated facial expression stimuli for emotion research. <em>Behavior Research Methods, 58</em>(11), Article 296. <a href="https://doi.org/10.3758/s13428-026-03156-0" rel="noopener noreferrer">https://doi.org/10.3758/s13428-026-03156-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.3758/s13428-026-03156-0" rel="noopener noreferrer">10.3758/s13428-026-03156-0</a></p>
<p><strong>Keywords:</strong> artificial intelligence, facial expressions, emotion perception, stimulus validation, body weight, generative AI, LoRA, ComfyUI, Py-Feat, behavioral research, photorealism, psychology methods</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">207691</post-id>	</item>
	</channel>
</rss>
