<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>AI-based app analysis &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/ai-based-app-analysis/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 11 Oct 2026 03:01:41 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>AI-based app analysis &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Android Malware Caught by Teaching AI to Read Apps as Pictures</title>
		<link>https://scienmag.com/android-malware-caught-by-teaching-ai-to-read-apps-as-pictures/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sun, 11 Oct 2026 03:01:41 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI-based app analysis]]></category>
		<category><![CDATA[Android malware detection]]></category>
		<category><![CDATA[converting app features to images]]></category>
		<category><![CDATA[cybersecurity]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning for mobile security]]></category>
		<category><![CDATA[ensemble methods]]></category>
		<category><![CDATA[feature fusion]]></category>
		<category><![CDATA[grayscale images]]></category>
		<category><![CDATA[image-based Android threat identification]]></category>
		<category><![CDATA[innovative malware detection techniques]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in cybersecurity]]></category>
		<category><![CDATA[majority voting]]></category>
		<category><![CDATA[mobile security]]></category>
		<category><![CDATA[QR code images]]></category>
		<category><![CDATA[recursive feature elimination]]></category>
		<category><![CDATA[signature-free malware detection]]></category>
		<category><![CDATA[static and dynamic app analysis]]></category>
		<category><![CDATA[vision transformer]]></category>
		<category><![CDATA[vision transformer for malware detection]]></category>
		<category><![CDATA[visual detection of malicious software]]></category>
		<category><![CDATA[visual pattern recognition in cybersecurity]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=260986</guid>

					<description><![CDATA[Researchers in Turkey have built a multimodal Vision Transformer framework that converts Android app features into grayscale and QR code images, achieving 98.72 percent accuracy in malware detection through feature fusion and majority-voting ensembles.]]></description>
										<content:encoded><![CDATA[<p>Android devices have become the default computing platform for billions of people, and that success has made them the single most attractive target for malicious software in the mobile world. Traditional defenses, which rely on signature-based detection, essentially keep a library of known threats and flag anything that matches. That approach falters the moment attackers tweak a few lines of code to produce a new variant that no longer matches any stored signature. A research team in Turkey, led by Erdal Başaran and Ömer Okucu of Ağrı İbrahim Çeçen University together with Yusuf Alaca of Hitit University, has now unveiled a detection framework that takes an unusually visual route to the problem: it turns the telltale features of an Android application into images and lets a Vision Transformer, one of the most powerful architectures in modern computer vision, learn to spot the difference between benign software and malware.</p>
<p>The core idea behind the framework, published in Cluster Computing, is to convert both static and dynamic analysis features of an application into two distinct two-dimensional image formats. The first is a conventional 2D grayscale image, in which extracted feature vectors are reshaped into pixel matrices so that patterns of malicious behavior appear as visual textures. The second is a QR code image representation, a more unusual encoding that arranges feature information into the blocky, high-contrast structure familiar from quick-response codes. The researchers trained separate Vision Transformer models on each image type. A Vision Transformer works by slicing an image into small patches, treating each patch as a token, and then using self-attention mechanisms to model the relationships between all patches simultaneously. This allows the model to capture long-range spatial dependencies that convolutional networks, which process images through local filters, can miss.</p>
<p>Once the two ViT models had been trained, the team did not simply take their final classification verdicts. Instead, they extracted the spatial pooling features from the deeper layers of each network, the compact numerical summaries that the transformers build up as they interpret the images. These feature sets, one derived from grayscale images and one from QR code images, were then combined into a single fused representation. The logic is that the two image encodings emphasize different aspects of an application&#8217;s behavior, so a classifier that sees both should have a richer picture than one that sees either alone. Feature fusion of this kind has become a recurring theme in malware research precisely because no single view of a program, whether permissions, API calls, or runtime behavior, tells the whole story.</p>
<p>The fused feature vector, however, contained a mixture of informative and redundant attributes, and feeding everything into a classifier can dilute the signal. To address this, the researchers applied Recursive Feature Elimination, an iterative selection technique that trains a model, ranks the features by importance, discards the weakest ones, and repeats the process until only the most discriminative subset remains. This RFE-optimized selection step proved to be one of the decisive ingredients of the framework. By pruning away noise before classification, it sharpened the boundaries between benign and malicious samples and reduced the computational burden of the final decision stage.</p>
<p>For the final classification, the team deployed a majority-voting ensemble, a strategy with a long pedigree in machine learning. Rather than trusting a single algorithm, several different machine learning classifiers each cast a vote on whether a sample is malicious, and the majority decision prevails. Ensembles of this kind tend to be more robust than individual models because the errors of one classifier can be corrected by the others, provided their mistakes are not perfectly correlated. Combined with the fused and filtered features, the voting ensemble delivered the framework&#8217;s headline result: an overall detection accuracy of 98.72 percent, outperforming the standalone Vision Transformer models and every single-modality configuration the authors tested.</p>
<p>The ablation-style comparison embedded in the results carries an important message for the field. Neither the grayscale pathway nor the QR code pathway alone matched the multimodal system, and the gains from fusing features and applying RFE-based selection were described by the authors as substantial contributions to the overall performance. In other words, the improvement did not come from a single clever trick but from the deliberate stacking of complementary techniques: two image representations, two transformer encoders, feature fusion, disciplined feature selection, and ensemble voting. Each layer of the pipeline removes a different weakness of the layer before it.</p>
<p>The study situates itself within a rapidly growing literature on image-based malware detection. Previous work has explored converting bytecode, permission lists, and network traffic into images for convolutional neural networks, and recent efforts have introduced transformer architectures such as ViTDroid for attending to malicious behavior in Android binaries. Others have fused multivariate features or applied reinforcement learning to feature selection. The Turkish team&#8217;s contribution is to bring these threads together in a single multimodal pipeline and to demonstrate, with a publicly available dataset, that the combination outperforms its individual components. The dataset used in the study can be accessed through the official UNB CIC website and Kaggle, a decision that supports reproducibility and allows other groups to benchmark against the reported results.</p>
<p>The practical implications are considerable. Because the framework relies on static and dynamic analysis features rather than on signatures of known malware families, it has the potential to generalize to new and obfuscated threats that signature databases have never seen. The QR code representation in particular is a novel twist, building on earlier work by one of the co-authors on cyber attack detection using QR code images with lightweight deep learning models. Encoding security-relevant features into a format designed for machine readability appears to produce distinctive visual signatures of malicious behavior that transformers can exploit effectively. For mobile security vendors, the findings suggest that multimodal image representations could become a standard component of next-generation detection engines, especially as transformer models continue to be optimized for deployment on resource-constrained platforms.</p>
<p>Challenges remain before such systems reach production. Image-based approaches depend on the quality and completeness of the underlying feature extraction, and adversaries who understand the encoding scheme may attempt to craft features that evade visual detection. The computational cost of running two Vision Transformer models per application also needs to be weighed against the latency requirements of app store scanning and on-device security tools. Nevertheless, the 98.72 percent accuracy reported by Başaran, Alaca, and Okucu represents a compelling data point in the ongoing effort to stay ahead of Android malware, and it underscores a broader trend in cybersecurity: the most effective defenses increasingly come not from looking harder at code, but from teaching machines to see it in entirely new ways.</p>
<p><strong>Subject of Research:</strong> Multimodal Vision Transformer-based Android malware detection using 2D grayscale and QR code image representations</p>
<p><strong>Article Title:</strong> Multimodal vision transformer framework for android malware detection via 2D grayscale and QR code ımage representations, feature fusion, and RFE-optimized majority voting</p>
<p><strong>Article References:</strong> Başaran, E., Alaca, Y., &amp; Okucu, Ö. (2026). Multimodal vision transformer framework for android malware detection via 2D grayscale and QR code ımage representations, feature fusion, and RFE-optimized majority voting. <em>Cluster Computing, 29</em>(15), Article 834. <a href="https://doi.org/10.1007/s10586-026-06667-9" rel="noopener noreferrer">https://doi.org/10.1007/s10586-026-06667-9</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10586-026-06667-9" rel="noopener noreferrer">10.1007/s10586-026-06667-9</a></p>
<p><strong>Keywords:</strong> Android malware detection, Vision Transformer, QR code images, grayscale images, feature fusion, Recursive Feature Elimination, majority voting, machine learning, cybersecurity, deep learning, ensemble methods, mobile security</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">260986</post-id>	</item>
	</channel>
</rss>
