<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>loss function &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/loss-function/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 10:34:48 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>loss function &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Curved math gives activity-recognition AI a surprising accuracy boost</title>
		<link>https://scienmag.com/curved-math-gives-activity-recognition-ai-a-surprising-accuracy-boost/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 10:34:48 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[activity recognition accuracy improvement]]></category>
		<category><![CDATA[activity recognition with sensor and video data]]></category>
		<category><![CDATA[camera-based human activity monitoring]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[edge computing]]></category>
		<category><![CDATA[geometric awareness in neural networks]]></category>
		<category><![CDATA[GPT-2]]></category>
		<category><![CDATA[human activity recognition]]></category>
		<category><![CDATA[human activity recognition using curved geometric models]]></category>
		<category><![CDATA[loss function]]></category>
		<category><![CDATA[loss functions incorporating geometric constraints]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning benchmarks for activity recognition]]></category>
		<category><![CDATA[mathematical approaches to activity recognition]]></category>
		<category><![CDATA[multimodal fusion]]></category>
		<category><![CDATA[neural network parameter spaces on curved surfaces]]></category>
		<category><![CDATA[Riemannian manifolds in machine learning]]></category>
		<category><![CDATA[Riemannian optimization]]></category>
		<category><![CDATA[sensor data]]></category>
		<category><![CDATA[sensor-based human activity detection]]></category>
		<category><![CDATA[Stiefel manifold]]></category>
		<category><![CDATA[transformers]]></category>
		<category><![CDATA[wearable sensors]]></category>
		<category><![CDATA[wearable sensors for activity recognition]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=222054</guid>

					<description><![CDATA[Researchers have introduced a Riemannian manifold-based loss function that pushes multimodal Transformer models to state-of-the-art accuracy in human activity recognition while remaining efficient enough for real-time edge deployment.]]></description>
										<content:encoded><![CDATA[<p>Teaching computers to recognize what a person is doing—walking, sitting, swinging a tennis racket, or preparing a meal—has long been one of the central goals of machine learning research. A new study published in Machine Learning with Applications by Farnaz Soleimani, Ghazaleh Khodabandelou, Abdelghani Chibani, and Yacine Amirat takes a fresh swing at this problem, and the twist is mathematical: instead of treating the internal parameters of a neural network as ordinary numbers floating in flat Euclidean space, the researchers constrain key parts of the model to live on curved geometric surfaces known as Riemannian manifolds. The result is a loss function that injects geometric awareness into the training process, and it delivers consistent accuracy gains across three widely used human activity recognition benchmarks.</p>
<p>The field of human activity recognition, or HAR, has evolved dramatically since its beginnings in the early 2000s, when researchers first strapped accelerometers to volunteers and applied classical algorithms such as decision trees and k-nearest neighbors to the resulting signals. Wearable sensor systems offer privacy and continuous tracking but often struggle with limited contextual information and sensitivity to sensor placement. Camera-based systems, by contrast, capture rich visual detail through RGB video, depth maps, and skeletal data, achieving high accuracy at the cost of privacy concerns, fixed fields of view, and vulnerability to lighting and background changes. Each modality has its own blind spots: skeleton data encodes human motion but omits the objects a person interacts with, while RGB captures both the person and the scene but falters when illumination shifts. Combining modalities has therefore become the dominant strategy for building systems that can cope with the messiness of the real world.</p>
<p>Deep learning accelerated this trend considerably. Convolutional and recurrent architectures, including CNN-LSTM hybrids, automated feature extraction and pushed recognition accuracy well beyond what handcrafted statistical features could achieve. Yet as datasets grew in scale and diversity, these models revealed intrinsic limitations. Their narrow receptive fields and strictly sequential processing make it difficult to capture long-range temporal dependencies and the complex interactions between modalities. Recognition accuracy, studies have shown, increases substantially with longer context windows, which is precisely where recurrent networks struggle. Transformers, with their self-attention mechanism, have emerged as the natural successor, dynamically focusing on the most salient features across entire sequences and processing them in parallel rather than one step at a time.</p>
<p>The new work builds on this Transformer backbone but attacks a different part of the pipeline: the loss function that guides training. The authors&#8217; multimodal architecture processes four data streams—RGB video, depth maps, Kinect skeletal coordinates, and inertial sensor readings—through modality-specific encoders. Visual streams pass through 3D ResNet-18 backbones that extract spatiotemporal features, skeletal data flows through a graph convolutional network that respects the joint-and-bone structure of the human body, and inertial signals are handled by an LSTM. The resulting features are projected into a common 128-dimensional embedding space, enriched with learnable positional encodings, and fused by a Transformer encoder with two layers and four attention heads before a classification layer produces the final activity prediction.</p>
<p>The innovation lies in what happens during optimization. Standard training treats all parameters as free vectors updated by gradient descent in flat space. The proposed Riemannian loss adds a regularization term that measures the gradient of the cross-entropy loss with respect to a metric tensor, encouraging the model&#8217;s fusion projection matrices and classification weights to remain on a Stiefel manifold—the set of matrices with orthonormal columns. In practice, the Euclidean gradient computed by backpropagation is projected onto the tangent space of the manifold, and parameters are updated via a QR-based retraction that keeps them feasible. Only about three percent of the network&#8217;s parameters, those in the fusion and classification layers, are constrained this way; everything else trains with standard Adam. This two-track scheme keeps the method compatible with ordinary deep learning frameworks and requires no second-order gradient computation.</p>
<p>The empirical results are striking. On the UCI-HAR smartphone dataset, the Transformer&#8217;s accuracy rose from 97.75 percent with a standard loss to 99.80 percent with the Riemannian loss. On the more challenging, class-imbalanced Opportunity dataset, accuracy climbed from 93.01 to 95.10 percent. On the multimodal UTD-MHAD dataset, which contains synchronized RGB, depth, skeleton, and inertial recordings of 27 actions performed by eight subjects, fused multimodal accuracy improved from 98.84 to 99.42 percent. Notably, among single modalities, depth data achieved the highest unimodal accuracy at 97.68 percent while inertial data lagged at 85.55 percent, underscoring how much complementary information the fusion stage has to work with—and how much the geometric regularizer helps exploit it.</p>
<p>Crucially, the authors went beyond the favorable evaluation protocols that can inflate reported accuracy. Under stratified random splits, windows from the same person can appear in both training and test sets, allowing the model to exploit within-subject regularities in posture and gait. When the researchers repeated their experiments with subject-independent splits, in which test subjects are entirely unseen during training, accuracy dropped as expected—but the Riemannian loss preserved a clear advantage, holding 95.18 percent versus 92.03 percent for the standard loss. An eight-fold leave-one-subject-out cross-validation confirmed the pattern, with mean accuracies of 94.9 plus or minus 0.6 percent for the proposed method against 92.1 plus or minus 0.8 percent for the baseline. In zero-shot cross-dataset transfer from UTD-MHAD to the NTU-RGB+D 120 benchmark, the Riemannian model reached 83.9 percent against 78.6 percent, and on free-living inertial data from the RealWorld-HAR dataset it achieved 86.4 percent F1, outperforming a CNN-LSTM baseline at 82.7 percent.</p>
<p>Statistical rigor backs the headline numbers. Across ten random seeds on a fixed stratified partition, the Riemannian variant averaged 99.04 plus or minus 1.11 percent accuracy, beating both the standard cross-entropy baseline at 97.77 plus or minus 1.31 percent and an L2-regularized variant at 96.91 plus or minus 1.74 percent. One-sided Wilcoxon signed-rank tests yielded p-values of 0.002 and below 0.001 with large effect sizes, and the results survived Holm correction. An ablation study reinforced the point: swapping the manifold-aware term for a conventional L2 penalty dropped accuracy from 99.42 to 99.08 percent, showing that explicitly respecting the geometry of the data provides regularization that flat-space penalties cannot match. Reducing the Transformer&#8217;s depth or attention width also degraded performance, confirming that both architectural capacity and geometric constraints matter.</p>
<p>Perhaps the most surprising demonstration of generality came from an appendix experiment with GPT-2, an autoregressive language model with a fundamentally different architecture from the multimodal Transformer. When fine-tuned on sensor windows converted for sequence modeling, GPT-2&#8217;s accuracy on UCI-HAR rose from 84.03 to 87.60 percent with the Riemannian loss, and on UTD-MHAD multimodal fusion jumped from 78.49 to 85.67 percent. The loss function, in other words, is model-agnostic: geometric regularization helps whether the backbone is an encoder-style Transformer or a generative language model, hinting at a broader role for manifold-aware optimization as large language models enter the activity recognition arena.</p>
<p>Deployment considerations matter just as much as accuracy for real-world systems, and here the news is also good. Because only the small fusion matrices live on the manifold, the computational overhead is modest: training throughput on an RTX A6000 dropped about 11 percent relative to Adam, with under one percent extra memory, yet the Riemannian optimizer reached 95 percent of its final accuracy in 25 percent fewer epochs, keeping wall-clock training time comparable. On edge hardware, the isolated Transformer blocks run at 54 frames per second on a Jetson Xavier NX and sustain 32 frames per second on a Jetson Orin Nano with INT8 quantization, comfortably meeting the 25 to 30 frames per second of typical camera streams. The authors envision applications ranging from fall detection and routine monitoring for elderly care to workplace safety and rehabilitation feedback, while cautioning that fully unconstrained in-the-wild deployments—with occluded views, adversarial lighting, and shifting domains—will demand stronger adaptation techniques. For now, the message is clear: sometimes the shortest path to better AI is not a straight line at all, but a geodesic.</p>
<p><strong>Subject of Research:</strong> Riemannian optimization for multimodal human activity recognition with Transformers</p>
<p><strong>Article Title:</strong> Activity recognition via Riemannian optimization and multimodal Transformers</p>
<p><strong>Article References:</strong> Soleimani, F., Khodabandelou, G., Chibani, A., &amp; Amirat, Y. (2026). Activity recognition via Riemannian optimization and multimodal Transformers. <em>Machine Learning with Applications, 26</em>, Article 101021. <a href="https://doi.org/10.1016/j.mlwa.2026.101021" rel="noopener noreferrer">https://doi.org/10.1016/j.mlwa.2026.101021</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.mlwa.2026.101021" rel="noopener noreferrer">10.1016/j.mlwa.2026.101021</a></p>
<p><strong>Keywords:</strong> human activity recognition, Riemannian optimization, Transformers, multimodal fusion, deep learning, wearable sensors, Stiefel manifold, GPT-2, edge computing, machine learning, sensor data, loss function</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">222054</post-id>	</item>
	</channel>
</rss>
