<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>deep video transformers in industrial monitoring &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/deep-video-transformers-in-industrial-monitoring/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 09 Oct 2026 13:33:14 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>deep video transformers in industrial monitoring &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Fuzzy Logic Gives AI Video Auditors the Confidence to Explain Themselves</title>
		<link>https://scienmag.com/fuzzy-logic-gives-ai-video-auditors-the-confidence-to-explain-themselves/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Fri, 09 Oct 2026 13:33:14 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[action segmentation]]></category>
		<category><![CDATA[action segmentation in untrimmed videos]]></category>
		<category><![CDATA[AI video auditing]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[continuous video analysis for compliance]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep video transformers in industrial monitoring]]></category>
		<category><![CDATA[explainable AI]]></category>
		<category><![CDATA[explainable AI for video analysis]]></category>
		<category><![CDATA[fuzzy logic]]></category>
		<category><![CDATA[fuzzy logic in artificial intelligence]]></category>
		<category><![CDATA[human-readable AI explanations]]></category>
		<category><![CDATA[industrial automation]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models for audit reports]]></category>
		<category><![CDATA[machine learning for industrial workflow analysis]]></category>
		<category><![CDATA[MS-TCN++]]></category>
		<category><![CDATA[retrieval-augmented generation]]></category>
		<category><![CDATA[retrieval-augmented generation in AI]]></category>
		<category><![CDATA[TimeSformer]]></category>
		<category><![CDATA[transparent AI decision-making]]></category>
		<category><![CDATA[trust in AI video systems]]></category>
		<category><![CDATA[video auditing]]></category>
		<category><![CDATA[VideoMAE]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=254121</guid>

					<description><![CDATA[Researchers at the University of Essex and British Telecom have built an AI video auditing system that combines fuzzy logic, dual-stream video transformers and retrieval-augmented language models to segment industrial workflows and generate calibrated, human-readable audit reports.]]></description>
										<content:encoded><![CDATA[<p>Artificial intelligence has become remarkably good at watching videos, but it remains notoriously bad at explaining what it saw. A new study published in the Journal of Ambient Intelligence and Humanized Computing tackles that gap head-on, presenting a framework that watches long, unedited recordings of industrial work and produces audit reports a human inspector can actually read, question and trust. The research, led by Yahia Mady and Hani Hagras of the University of Essex together with Hugo Leon-Garza and Anasol Pena-Rios of British Telecom, weaves together three strands of modern machine learning: deep video transformers, fuzzy logic, and large language models connected to a knowledge base through retrieval-augmented generation. The result is a system that does not merely label what happens on screen, but tells you how sure it is, in plain language, and why.</p>
<p>The problem the researchers set out to solve is called action segmentation, and it is far harder than the clip-recognition tasks that dominate video AI benchmarks. Instead of classifying a short, pre-trimmed snippet containing one activity, an auditing system must take a continuous, untrimmed recording and assign a label to every moment of it, deciding precisely where one workflow step ends and the next begins. In real industrial settings that boundary problem is brutal. Workers&#8217; hands overlap between tasks, lighting shifts, cameras move, some steps look almost identical to others, and rare actions are vastly outnumbered by common ones. Early convolutional networks and recurrent models captured local temporal patterns but struggled with long-range context and error propagation, while transformer architectures improved temporal reasoning at the cost of computational expense and even less transparency about their decisions.</p>
<p>The new framework&#8217;s answer begins with how it looks at the video. Each long recording is chopped into overlapping sliding windows of 64 frames, and every window is described by two frozen pre-trained encoders working in parallel. A TimeSformer stream applies divided space-time attention, factorising self-attention into separate temporal and spatial operations so the model can track both what objects look like and how they move over long ranges. A VideoMAE stream, trained through masked autoencoding in which 90 percent or more of the spatiotemporal tokens are hidden and must be reconstructed, contributes motion-aware features learned without labels. The two 768-dimensional descriptors are concatenated into a single 1536-dimensional vector per window, and the full sequence of these fused descriptors is passed to MS-TCN++, a multi-stage temporal convolutional network whose stacked dilated-convolution stages iteratively refine predictions, suppress over-segmentation errors and sharpen boundaries across temporal scales spanning up to 1024 windows.</p>
<p>The first and most distinctive contribution of fuzzy logic appears during training. Conventional segmentation systems give every window a single crisp label, usually the dominant action, which corrupts supervision near step transitions where two actions genuinely overlap. Instead, the team built a type-1 Mamdani fuzzy system that converts ground-truth annotations into boundary-aware soft labels. Each window is summarised by the local proportion of frames belonging to its two top candidate actions and by the global frequency of those actions across the whole video. Triangular membership functions map these four inputs to linguistic terms, and a compact rule base of 27 rules, deliberately reduced from the 81 logically possible antecedent combinations because most are structurally infeasible within fixed-length windows, assigns degrees of membership such as Possible, Likely or Very Likely. Windows straddling a transition therefore receive graded, honest supervision rather than a falsely confident single label.</p>
<p>The second fuzzy system operates at inference time and is where the framework earns its explainability credentials. Raw softmax probabilities, the numbers deep networks normally output, rank predictions reasonably well but are not calibrated: a stated probability of 0.85 does not correspond to an 85 percent chance of being correct, and the number says nothing about how the model reached it. The researchers&#8217; Decision Confidence system instead takes three inputs: the average peak softmax probability across the window, the gap between the top-1 and top-2 probabilities, and the dominance of the predicted class across the entire video. Each input is partitioned into fuzzy sets such as Low, Medium and High, and a complete rule base of 27 rules maps combinations to a five-level linguistic output running from Very Low to Very High confidence. Centre-of-sets defuzzification yields a scalar, but crucially the system can also report the linguistic band and the exact rules that fired.</p>
<p>Experiments were conducted on a demanding but deliberately modest scale: 86 untrimmed user-contributed videos from the Gadgets domain of the COIN dataset, all depicting the complete procedure of making an RJ-45 network cable, from stripping insulation and arranging wires to crimping the connector. The videos vary in lighting, camera placement and background clutter, split into 61 training, 7 validation and 18 held-out test videos across five action classes plus background. Seven configurations were compared, including ablations that swapped the dual-stream backbone for single streams, replaced MS-TCN++ with a simple per-window MLP head, and substituted hard labels for fuzzy ones, plus an external ASFormer transformer baseline trained on identical cached features, splits and seeds. The full proposed configuration achieved the best window-level macro F1 score of 0.8241, comfortably ahead of the ASFormer baseline at 0.7624, with both encoder streams shown to contribute and cross-window temporal context emerging as the dominant factor.</p>
<p>Perhaps the most scientifically honest finding concerns what fuzzy labelling did not do. The accuracy difference between fuzzy and hard labels, 0.8241 versus 0.8146, fell within seed-to-seed variance, so the authors explicitly decline to claim any recognition improvement. The fuzzy layer&#8217;s real contribution, they argue, lies elsewhere: in calibrated, rule-traceable confidence. On 1164 test windows, the fuzzy Decision Confidence matched raw softmax and softmax-margin baselines on error detection as measured by AUROC, but substantially outperformed them on calibration, measured by expected calibration error. Accuracy climbed from roughly 0.42 in the lowest confidence bin to about 0.95 in the highest, and the system&#8217;s tendency toward underconfidence is a deliberate safety feature, routing borderline cases to human review rather than letting errors slip through. Because each score arrives with a linguistic band and an explicit rule trace, a reviewer can set a review threshold in terms the system itself can explain.</p>
<p>The final stage turns predictions into prose. A retrieval-augmented generation pipeline embeds each predicted step and searches a FAISS index, built with the bge-small-en-v1.5 embedding model, over a knowledge base of eleven documents derived from a public WikiHow article describing the cable-making procedure, standing in for the standard operating documents a deploying organisation would supply. A Qwen2-7B large language model, served in quantised GGUF format with a 32,768-token context and a low decoding temperature of 0.1, then integrates the predicted steps with their retrieved references to produce a structured audit report enumerating completed, missing and misordered steps with natural-language justifications anchored in the retrieved evidence. Evaluated against manually annotated deviations across all 86 videos, supplemented with 72 controlled reordering cases because natural misorderings appeared in only three recordings, the reporting module achieved near-total recall, detecting omissions and misorderings with high fidelity while erring conservatively toward over-reporting missing steps rather than concealing them.</p>
<p>The authors are candid about the limits of the work. The evaluation covers a single task from a single dataset domain with a small action vocabulary and an 18-video test set, transfer to other domains remains untested, the deviation ground truth was annotated by a single author, and no separate human rating of report fluency or groundedness was collected. The implementation, developed with British Telecommunications under a non-disclosure agreement, cannot be released publicly, and deployment cost-effectiveness was not measured. Yet the architectural direction is clear and consequential. The team plans to adopt adaptive fuzzy systems, fuse multimodal data, and exploit vision-language models fine-tuned efficiently through techniques such as LoRA to build auditors that generalise across domains while keeping their reasoning inspectable.</p>
<p>What makes this work resonate beyond its niche is the broader message about trustworthy AI in high-stakes workplaces. In manufacturing, healthcare and telecommunications, the difference between a useful auditing system and a dangerous one is rarely raw accuracy; it is whether a human supervisor can understand, interrogate and override the machine&#8217;s judgement. By replacing opaque probability scalars with calibrated confidence expressed through human-readable rules, and by grounding every generated claim in retrievable procedural documents, the Essex and BT team demonstrates that the path to explainable industrial AI may run not through ever-larger black boxes, but through a thoughtful marriage of fuzzy logic&#8217;s interpretability, deep learning&#8217;s perceptual power, and generative models&#8217; ability to speak our language. The videos keep getting longer; now, at least, the explanations can keep up.</p>
<p><strong>Subject of Research:</strong> A fuzzy logic and retrieval-augmented generative AI framework for explainable automated action segmentation and auditing of industrial videos</p>
<p><strong>Article Title:</strong> A fuzzy logic based generative explainable AI for automated video auditing</p>
<p><strong>Article References:</strong> Mady, Y., Hagras, H., Leon-Garza, H., &amp; Pena-Rios, A. (2026). A fuzzy logic based generative explainable AI for automated video auditing. <em>Journal of Ambient Intelligence and Humanized Computing</em>. <a href="https://doi.org/10.1007/s12652-026-05136-w" rel="noopener noreferrer">https://doi.org/10.1007/s12652-026-05136-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s12652-026-05136-w" rel="noopener noreferrer">10.1007/s12652-026-05136-w</a></p>
<p><strong>Keywords:</strong> fuzzy logic, explainable AI, video auditing, action segmentation, retrieval-augmented generation, large language models, VideoMAE, TimeSformer, MS-TCN++, deep learning, industrial automation, computer vision</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">254121</post-id>	</item>
	</channel>
</rss>
