<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>computer-generated training datasets &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/computer-generated-training-datasets/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Tue, 06 Oct 2026 17:20:35 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>computer-generated training datasets &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Virtual Orchards: Synthetic Data Closes the Gap for AI Fruit Detection</title>
		<link>https://scienmag.com/virtual-orchards-synthetic-data-closes-the-gap-for-ai-fruit-detection/</link>
		
		<dc:creator><![CDATA[Alan Morgan]]></dc:creator>
		<pubDate>Tue, 06 Oct 2026 17:20:35 +0000</pubDate>
				<category><![CDATA[Agriculture]]></category>
		<category><![CDATA[agricultural computer vision]]></category>
		<category><![CDATA[annotation geometry]]></category>
		<category><![CDATA[apple detection]]></category>
		<category><![CDATA[apple detection in orchards]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[automated fruit annotation]]></category>
		<category><![CDATA[benchmarking apple-detection datasets]]></category>
		<category><![CDATA[Blender]]></category>
		<category><![CDATA[challenges in real-world agricultural image collection]]></category>
		<category><![CDATA[computer-generated training datasets]]></category>
		<category><![CDATA[dataset benchmarking]]></category>
		<category><![CDATA[domain randomisation]]></category>
		<category><![CDATA[generalization of AI models in agriculture]]></category>
		<category><![CDATA[impact of data quality on fruit detection models]]></category>
		<category><![CDATA[improving AI accuracy with synthetic images]]></category>
		<category><![CDATA[neural network orchard imaging]]></category>
		<category><![CDATA[object detection]]></category>
		<category><![CDATA[precision agriculture]]></category>
		<category><![CDATA[procedural synthetic dataset creation]]></category>
		<category><![CDATA[sim-to-real transfer]]></category>
		<category><![CDATA[synthetic data]]></category>
		<category><![CDATA[synthetic data for agricultural robotics]]></category>
		<category><![CDATA[Unreal Engine 5]]></category>
		<category><![CDATA[YOLO]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=242083</guid>

					<description><![CDATA[A systematic study shows that carefully engineered synthetic orchard imagery, refined to match real annotation geometry and mixed with a small fraction of real photographs, can train apple-detection models that rival those trained on fully real datasets.]]></description>
										<content:encoded><![CDATA[<p>Training an artificial intelligence to spot apples in an orchard sounds simple enough: point a camera at a tree, label the fruit, and let a neural network learn. In practice, it is one of the most stubborn bottlenecks in agricultural robotics. Collecting thousands of orchard photographs is laborious enough, but annotating every visible apple with a bounding box is worse, and the resulting datasets are often riddled with inconsistencies that quietly sabotage the models trained on them. A new study published in Smart Agricultural Technology argues that the solution may not lie in collecting more real images, but in manufacturing better ones inside a computer.</p>
<p>The research, conducted by Jhonny Hueller, Massimo Vecchio, and Fabio Antonelli, takes an unusually rigorous approach to a question that has mostly been answered piecemeal. Rather than demonstrating yet another synthetic-data pipeline for a single task, the team systematically benchmarked all ten publicly available apple-detection datasets, measured how well models trained on each one generalised to a new real-world test set, and then used those findings to iteratively refine a family of fully procedural synthetic datasets. The central question was not whether synthetic images can be generated at scale, but under what conditions they actually teach a detector something transferable to the real world.</p>
<p>The benchmark itself produced a striking result. When the researchers trained the same YOLOv26 Medium detector, with identical hyperparameters and a five-fold cross-validation scheme, on each public dataset and evaluated it on their own newly collected OpenIoT test set of 104 smartphone photographs containing 6,573 annotated apples, three datasets stood far above the rest: MinneApple, MetaFruit, and APPLE MOTS, with peak F1 scores around 0.58 to 0.64. Seven other collections languished far below, some with peak F1 scores under 0.35. Crucially, performance did not correlate with dataset size. Some of the largest collections performed worst, while a manual audit revealed that the failures traced back to low resolution, poor lighting, incomplete labels, duplicated imagery, and, perhaps most importantly, inconsistent annotation geometry, such as bounding boxes that were oversized or drawn from a limited set of predefined shapes.</p>
<p>That last finding shaped the entire synthetic-data strategy. The team built their virtual orchards in Blender, using an open-source procedural tree generator to create apple trees from random seeds and parameter ranges, and then rendered scenes in Unreal Engine 5, chosen for its photorealistic real-time rendering. Because every object in a virtual scene has known geometry, ground-truth annotations come for free and are pixel-perfect by construction. They produced three synthetic collections of 3,840 images each. Synthetic A captured broad scene diversity. Synthetic B was redesigned after analysing the best public datasets, aligning camera angles, optics, and fruit-generation parameters so that bounding-box sizes matched the statistical profile of MinneApple and MetaFruit. Synthetic C added finer geometric refinements, including aspect-ratio alignment and camera-pose regularisation.</p>
<p>The results of this iterative refinement were decisive. Models trained purely on Synthetic A already beat all seven lower-quality public datasets, but, tellingly, adding more synthetic images made performance worse rather than better, with peak F1 scores actually rising as the dataset was downsampled from 3,840 to 240 images. Synthetic B, with its aligned annotation geometry, delivered a consistent jump of roughly five to eight percentage points in F1 score and mean average precision across all size tiers. Synthetic C added a further two to three points, reaching a peak F1 of about 0.52. The team quantified the geometric alignment using the Wasserstein distance between bounding-box-area distributions and found a moderate negative correlation with detection performance, suggesting that how tightly annotations hug the fruit matters as much as how many images you render.</p>
<p>Even so, purely synthetic training never quite matched the best real data. The gap to MinneApple remained statistically significant in bootstrap comparisons over 10,000 resamples. The authors attribute the residual difference to the inherent limits of simulation: a finite library of tree and fruit models, simplified botanical structure, procedurally placed fruit that ignores real branch correlations, subtle rendering artefacts invisible to humans, and missing confounds such as haze, lens aberration, and sensor noise. Deep networks, it turns out, notice things our eyes do not.</p>
<p>The most practically important finding came when the researchers mixed real images into the synthetic training sets. With a conservative blend of one real image for every four synthetic ones, performance leapt dramatically. The best hybrid, built on Synthetic C, reached a peak F1 of 0.652, statistically indistinguishable from the model trained entirely on MinneApple. A systematic sweep of real-data fractions showed that the benefit follows a curve of diminishing returns: the first five to ten percent of real images accounts for most of the improvement, and below roughly five percent the loss relative to a fully mixed training set becomes statistically significant. In other words, a training corpus that is overwhelmingly synthetic, salted with a small handful of real photographs, can perform within a few F1 points of a fully real-data pipeline while eliminating the vast majority of annotation labour.</p>
<p>The study also probed a question that often goes unexamined: does dataset redundancy matter? Using perceptual hashing and feature embeddings from a frozen DINOv2 vision transformer to detect near-duplicate images, the team found that redundancy varied wildly across datasets, from essentially zero to nearly complete duplication in APPLE MOTS, where 97 percent of frames were near-identical at the structural level. Yet redundancy showed little correlation with final performance. High-quality and low-quality collections alike appeared at both extremes. Within the scope of fruit detection, at least, annotation geometry appears to matter far more than whether images repeat.</p>
<p>The authors are careful about the boundaries of their claims. The absolute performance figures come from a single test set of 104 images photographed in one orchard on one day with one smartphone, and every model tested was a YOLOv26 Medium detector. The real-data fraction at which the sim-to-real gap closes may be a property of this particular pipeline rather than a universal constant, and whether the ordering of training sets survives with transformer-based detectors or different cultivars remains untested. They suggest two complementary paths forward: embracing domain randomisation, which deliberately randomises textures, colours, and lighting rather than chasing photorealism, or pushing realism further with modern 3D reconstruction techniques such as Gaussian splatting and neural radiance fields.</p>
<p>What the study establishes, with unusual methodological care, is an ordering: carefully engineered synthetic datasets beat poorly curated real ones, refined annotation geometry narrows the gap to the best real data, and a small dose of reality finishes the job. For a field where every labelled apple costs human time and every harvesting robot must cope with changing light, weather, and growth stages, that is a genuinely useful recipe. The vision of machines that learn to farm inside a simulator before ever touching soil moved a measurable step closer.</p>
<p><strong>Subject of Research:</strong> Synthetic dataset design for apple fruit detection with computer vision in agriculture</p>
<p><strong>Article Title:</strong> Designing effective synthetic datasets for fruit detection in agriculture</p>
<p><strong>Article References:</strong> Hueller, J., Vecchio, M., &amp; Antonelli, F. (2026). Designing effective synthetic datasets for fruit detection in agriculture. <em>Smart Agricultural Technology, 15</em>, Article 102605. <a href="https://doi.org/10.1016/j.atech.2026.102605" rel="noopener noreferrer">https://doi.org/10.1016/j.atech.2026.102605</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.atech.2026.102605" rel="noopener noreferrer">10.1016/j.atech.2026.102605</a></p>
<p><strong>Keywords:</strong> synthetic data, apple detection, agricultural computer vision, YOLO, sim-to-real transfer, Unreal Engine 5, Blender, object detection, annotation geometry, dataset benchmarking, precision agriculture, domain randomisation</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">242083</post-id>	</item>
	</channel>
</rss>
