<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>scalable video analysis systems &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/scalable-video-analysis-systems/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Tue, 06 Oct 2026 12:37:40 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>scalable video analysis systems &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New AI Pipeline Makes Searching Hours of Surveillance Video as Easy as Typing a Sentence</title>
		<link>https://scienmag.com/new-ai-pipeline-makes-searching-hours-of-surveillance-video-as-easy-as-typing-a-sentence/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Tue, 06 Oct 2026 12:37:40 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI in public safety and security]]></category>
		<category><![CDATA[AI-powered video retrieval]]></category>
		<category><![CDATA[BoT-SORT]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[deep learning for video search]]></category>
		<category><![CDATA[face recognition]]></category>
		<category><![CDATA[large-scale surveillance data processing]]></category>
		<category><![CDATA[modular AI architecture for industrial video analysis]]></category>
		<category><![CDATA[multimedia content indexing]]></category>
		<category><![CDATA[multimodal embeddings]]></category>
		<category><![CDATA[multimodal image and text matching]]></category>
		<category><![CDATA[natural language video querying]]></category>
		<category><![CDATA[person identification in surveillance footage]]></category>
		<category><![CDATA[person re-identification]]></category>
		<category><![CDATA[RSTPReID]]></category>
		<category><![CDATA[scalable video analysis systems]]></category>
		<category><![CDATA[SigLIP2]]></category>
		<category><![CDATA[snapshot-based person search]]></category>
		<category><![CDATA[surveillance video]]></category>
		<category><![CDATA[surveillance video search]]></category>
		<category><![CDATA[text-based person search]]></category>
		<category><![CDATA[vector database]]></category>
		<category><![CDATA[video retrieval]]></category>
		<category><![CDATA[YOLO11]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=241354</guid>

					<description><![CDATA[Researchers have built a modular AI pipeline that indexes hours of video and retrieves specific people from text, image, or face queries, achieving record mean Average Precision on the RSTPReid benchmark.]]></description>
										<content:encoded><![CDATA[<p>Every second, more than six hours of video land on YouTube, while city-wide camera networks churn out petabytes of footage each day. Buried inside those archives are people: suspects to locate, missing persons to find, customers to track across stores. The problem is that nobody has time to watch it all. A team of researchers at the University of Cagliari and the data engineering firm AgileLab has now built a system that promises to change that, letting users find any individual in a mountain of video simply by typing a description like &#8220;a young man with glasses and an olive-green parka&#8221; or by uploading a single snapshot. The work, published in Multimedia Tools and Applications, reports state-of-the-art results on a standard person-retrieval benchmark while keeping the whole architecture modular enough to scale to industrial workloads.</p>
<p>The core insight of the new study is that the bottleneck in video search is no longer the accuracy of individual AI models but the lack of a coherent system that stitches them together. Deep multimodal models such as CLIP and its successors have made it possible to align images and text in a shared mathematical space, so that a sentence and a photograph of the same scene end up close together as high-dimensional vectors. Yet, as the authors argue, most research focuses on squeezing extra accuracy out of a single model while ignoring the messy engineering questions: how to chunk hours of raw footage, how to avoid indexing the same person thousands of times, how to keep track of who appeared when and where, and how to serve queries in milliseconds. Their answer is a two-part pipeline that cleanly separates offline indexing from online retrieval.</p>
<p>On the indexing side, the system ingests raw video files and first splits them into roughly ten-minute chunks using FFmpeg, preserving the original encoding so each segment can be processed independently and in parallel. A frame-sampling module then skips redundant frames, since typical surveillance footage recorded at 30 or 60 frames per second offers far more temporal detail than person search actually needs. The sampled frames flow into YOLO11, the latest generation of Ultralytics&#8217; real-time object detector, chosen for its favorable speed-accuracy trade-off, paired with the BoT-SORT tracker, which combines motion cues with appearance features to assign each detected person a persistent identity across frames. On the MOT17 benchmark, BoT-SORT reports 80.5 MOTA and 80.2 IDF1, meaning it rarely confuses one pedestrian with another, a property the authors consider essential for generating temporally coherent metadata.</p>
<p>Detection alone would still flood the database with near-duplicate images, so the pipeline adds a quality-filtering stage that is among the paper&#8217;s most practical contributions. Every cropped person image receives a score combining the detector&#8217;s confidence with an edge penalty: bounding boxes that touch the frame border, indicating a partially visible subject, are down-weighted. A perceptual hashing step then compares consecutive crops within the same tracked trajectory, discarding any whose hash differs by fewer than a threshold number of bits, a signal that the image is essentially identical to the previous one. What survives is a compact, diverse gallery of high-quality person crops rather than a bloated collection of blurry, truncated, or repetitive snapshots.</p>
<p>Each surviving crop is then transformed into a semantic fingerprint by SigLIP2, a multilingual vision-language encoder that extends the original SigLIP architecture with a sigmoid-based training loss, self-supervised objectives, and careful data curation. Because the model was trained on paired images and text, its visual embeddings live in the same latent space as its text embeddings, which is precisely what allows a typed sentence to retrieve a photograph. The system also computes a second, complementary embedding for each person using InsightFace: the RetinaFace detector localizes the face and its landmarks even under challenging poses and lighting, and the ArcFace recognizer converts the aligned face patch into a highly discriminative identity vector. Both embeddings, along with rich metadata such as source video, timestamps, track identity, and bounding-box coordinates, are stored in a Qdrant vector database, so every search result can be traced back to the exact frame in the original footage.</p>
<p>At query time, the retrieval pipeline mirrors the indexing representation. A user can submit an image, a natural-language description, or both; the same SigLIP2 encoders embed the inputs, guaranteeing that queries and database entries are directly comparable. When both modalities are present, a tunable balance parameter blends the image and text vectors, letting users slide between purely visual similarity and purely linguistic matching. A face-based mode restricts the search to identity embeddings when a face is detectable, and a negative-prompt field excludes unwanted attributes. The back-end runs an approximate nearest-neighbor search with cosine similarity, then consolidates results so that only the single best match per tracked object per video is returned, each accompanied by the retrieved frame, a similarity score, and metadata including camera identifier, first-appearance timestamp, and duration in the scene.</p>
<p>The quantitative centerpiece of the paper is an evaluation on RSTPReid, a benchmark of 20,505 pedestrian images covering 4,101 individuals captured by 15 cameras, each image paired with two text descriptions. The researchers fine-tuned the SigLIP2-SO400M-Patch14-384 checkpoint on a curated mixture of person-centric image-text datasets, including CUHK-PEDES, ICFG-PEDES, RSTPReid itself, an attribute-rich multiview re-identification set, and the video-based TVPReid corpus, totaling roughly 316,000 training pairs. Training used the native contrastive loss on four NVIDIA A100 GPUs on the Cineca Leonardo supercomputer, with early stopping based on validation mean Average Precision. The result: a mean Average Precision of 0.68, comfortably ahead of the best competing value of 0.54 among recent text-based person search methods such as OCDL, UP-Person, RDE, CTGI, WoRA, MARS, MRA, and CONQUER, while Recall@1 remained competitive within one standard deviation of the leaders.</p>
<p>Notably, the gains did not come from inventing new architectures. When the model was fine-tuned only on the three standard datasets used by the comparison methods, it already achieved the best mAP, indicating that the SigLIP2 backbone itself provides a stronger shared embedding space for matching descriptions to pedestrian attributes. Adding the larger, more diverse corpus then improved all metrics further. Ablation experiments on a subset of the Wildtrack multi-camera surveillance dataset underscored the value of the system&#8217;s design choices: removing face embeddings dropped Recall@1 from about 51 percent to 33 percent, while replacing quality-scored crops with uniformly sampled ones collapsed mAP from 0.66 to 0.39, showing that both identity cues and careful crop selection matter substantially.</p>
<p>The scalability analysis offers a sober, engineering-minded picture. Processing 128 ten-minute videos, about 256 gigabytes of footage, took roughly 90 minutes on a four-GPU server, with tracking saturating the CPU and encoding bottlenecked by memory traffic and model replication across worker processes rather than raw computation. Extrapolating linearly, the authors estimate the current setup could index about a terabyte of video in six hours, though they are candid that GPU utilization during encoding remains low and that a centralized inference server would be needed for truly efficient large deployments. They also acknowledge limitations: the system is person-centric, does not yet model complex temporal activities or events, and the auxiliary training data included AI-generated descriptions that were only spot-checked rather than fully validated by humans.</p>
<p>Even with those caveats, the study reads as a template for how modern AI systems should be built: not as isolated models chasing leaderboard numbers, but as orchestrated pipelines where detection, tracking, filtering, embedding, metadata, and search reinforce one another. The authors envision extending the framework beyond people to vehicles and other objects, adding temporal and multi-view reasoning, and eventually integrating large language models and knowledge graphs for conversational, explainable video search. If surveillance archives, streaming platforms, and personal devices keep growing at their current pace, systems that turn unwatchable oceans of footage into a simple search box may soon become as unremarkable, and as indispensable, as the web search bar itself.</p>
<p><strong>Subject of Research:</strong> A scalable semantic pipeline for video indexing and text-based person retrieval using multimodal vision-language embeddings</p>
<p><strong>Article Title:</strong> A scalable and semantic pipeline for efficient video indexing and person retrieval</p>
<p><strong>Article References:</strong> Hmaidan, R., Milardo, S., Donato, I., Greco, D., Ingargiola, A., &amp; Reforgiato Recupero, D. (2026). A scalable and semantic pipeline for efficient video indexing and person retrieval. <em>Multimedia Tools and Applications, 85</em>(9), Article 727. <a href="https://doi.org/10.1007/s11042-026-21882-7" rel="noopener noreferrer">https://doi.org/10.1007/s11042-026-21882-7</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11042-026-21882-7" rel="noopener noreferrer">10.1007/s11042-026-21882-7</a></p>
<p><strong>Keywords:</strong> video retrieval, person re-identification, multimodal embeddings, SigLIP2, YOLO11, BoT-SORT, face recognition, vector database, surveillance video, computer vision, text-based person search, RSTPReid</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">241354</post-id>	</item>
	</channel>
</rss>
