<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Python software &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/python-software/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Wed, 23 Sep 2026 23:02:58 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>Python software &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New Python Library Puts Conflicting Evidence on the Same Page</title>
		<link>https://scienmag.com/new-python-library-puts-conflicting-evidence-on-the-same-page/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Wed, 23 Sep 2026 23:02:58 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[belief functions]]></category>
		<category><![CDATA[belief functions in decision making]]></category>
		<category><![CDATA[belief modeling in expert systems]]></category>
		<category><![CDATA[conflict redistribution]]></category>
		<category><![CDATA[conflicting information resolution in computing]]></category>
		<category><![CDATA[decision support]]></category>
		<category><![CDATA[Dempster–Shafer theory]]></category>
		<category><![CDATA[Dempster–Shafer Theory applications]]></category>
		<category><![CDATA[Dezert–Smarandache theory]]></category>
		<category><![CDATA[evidence fusion]]></category>
		<category><![CDATA[evidence theory for hypothesis management]]></category>
		<category><![CDATA[evidencelib]]></category>
		<category><![CDATA[handling ambiguous medical diagnoses]]></category>
		<category><![CDATA[information fusion]]></category>
		<category><![CDATA[multi-source data conflict resolution]]></category>
		<category><![CDATA[open-source evidence fusion library]]></category>
		<category><![CDATA[open-source library]]></category>
		<category><![CDATA[Python libraries for evidence-based reasoning]]></category>
		<category><![CDATA[Python software]]></category>
		<category><![CDATA[sensor data integration tools]]></category>
		<category><![CDATA[sensor fusion]]></category>
		<category><![CDATA[uncertain evidence modeling in Python]]></category>
		<category><![CDATA[uncertainty quantification]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=211062</guid>

					<description><![CDATA[Researchers have released evidencelib, an open-source Python library that unifies Dempster–Shafer and Dezert–Smarandache evidence modeling, high-conflict fusion rules, and decision support in a single dependency-free package.]]></description>
										<content:encoded><![CDATA[<p>A team of Polish researchers has released an open-source Python library that tackles one of the quietest but most consequential problems in modern computing: what to do when different sources of information disagree. The library, called evidencelib, was described in the journal SoftwareX by Szymon Śniegowski, Adrianna Świder, Andrii Shekhovtsov, and Wojciech Sałabun. It offers a single, coherent toolkit for modeling uncertain evidence and fusing it into decisions, even when the underlying sources contradict each other outright or describe hypotheses that overlap rather than exclude one another.</p>
<p>The mathematical backbone of the library is Dempster–Shafer Theory, sometimes called evidence theory or the theory of belief functions. Unlike classical probability, which requires every ounce of belief to be pinned to individual hypotheses, Dempster–Shafer Theory lets a source declare that its confidence covers a set of possibilities. A doctor reading a patient&#8217;s symptoms may be able to say the evidence points to a group of possible diseases without being able to split that support among them individually. This ability to represent partial knowledge and outright ignorance explicitly, rather than being forced to divide belief artificially, is what has made the theory attractive for sensor fusion, medical expert systems, target recognition in defense applications, and robotics.</p>
<p>Classical Dempster–Shafer Theory, however, rests on an assumption that rarely holds perfectly in the messy real world: that the possible hypotheses are mutually exclusive and exhaust all options. Categories can overlap, and information sources can generate extreme conflict. To handle overlapping hypotheses, researchers developed Dezert–Smarandache Theory, an extension built on a structure called the hyper-power set, which includes not just unions of hypotheses but also their intersections. A second, more notorious problem arises with Dempster&#8217;s original combination rule, which normalizes conflict away. In Zadeh&#8217;s famous counterexample, two sources strongly support two different hypotheses, and after normalization a third hypothesis that both sources consider nearly impossible ends up absorbing all the belief. That kind of counterintuitive result spurred a family of alternative fusion rules, including Yager&#8217;s rule, the transferable belief model of Smets, the Dubois–Prade rule, and the proportional conflict redistribution rules PCR5 and PCR6.</p>
<p>Building such machinery from scratch for every study is time-consuming and error-prone, and the authors found that existing tools each covered only fragments of the picture. The two existing Python packages and two R packages focus on classical Dempster–Shafer Theory, while the two MATLAB frameworks date from 2008 and 2010. According to a feature comparison conducted by the authors, no other modern tool combines classical DST, free and constrained Dezert–Smarandache models, a symbolic proposition parser, the hyper-power set, the DSmH fusion rule, and PCR5 within one interface. evidencelib, currently at version 1.2.0 under the MIT license, fills that gap. Its computational core is pure Python with zero runtime dependencies, requiring only Python 3.10 or newer, with optional Matplotlib support for visualization. It is installable from the Python Package Index, documented online, and validated by a test suite running in continuous integration on Python 3.10 through 3.14 with strict static type checking and an enforced coverage threshold of at least 90 percent.</p>
<p>Architecturally, the library is organized into five interconnected components. The Frame component defines the frame of discernment in one of three model variants: classical DST, the free DSm model in which all hypotheses may intersect, or hybrid models in which selected intersections are explicitly forbidden by constraints. A parser converts textual propositions into symbolic objects without executing arbitrary code, and the Proposition component encodes hypotheses as integer bitmasks over Venn regions, so that union, intersection, and containment reduce to single integer operations. The MassFunction component does the heavy lifting, offering belief, plausibility, commonality, and conflict measures; uncertainty metrics such as Deng entropy, fractal-based belief entropies, information volume, nonspecificity, and strife; the full menu of fusion rules; and round-trip import and export through dictionaries, JSON, and CSV, plus publication-ready LaTeX table output.</p>
<p>Two worked examples in the paper show what the library can do. The first is a weld inspection decision with three mutually exclusive outcomes: accept the component, repair it, or reject it. Visual inspection strongly supports acceptance with a mass of 0.70, while ultrasonic testing strongly supports repair with 0.65, producing a conjunctive conflict of 0.455 — a deliberately high value that exposes the differences between conflict-handling strategies. Smets&#8217; rule parks the conflicting mass on the empty set, Dempster&#8217;s rule normalizes it away, Yager&#8217;s rule transfers it to total ignorance, Dubois–Prade shunts it to the union of the contested hypotheses, and PCR5 redistributes it proportionally. The choice of rule changes the operational outcome: under Dempster&#8217;s rule the pignistic probability of acceptance is 0.539, clearing an illustrative decision threshold of 0.50, but under Yager&#8217;s rule it falls to 0.445, which under the example&#8217;s policy would trigger an additional inspection. The same evidence, the same hypotheses — yet a different action, purely because of how conflict is managed.</p>
<p>The second example leaps into territory classical theory cannot reach: diagnosing a pump that may suffer electrical, mechanical, and hydraulic faults simultaneously. In the free DSm model, these fault types may co-occur, and the hyper-power set for three hypotheses contains 19 propositions. Fusing readings from motor-current, vibration, and pressure-and-flow analysis with the conjunctive DSmC rule assigns the largest mass, 0.2925, to the intersection of electrical and mechanical faults, revealing a dominant electromechanical state with an additional hydraulic contribution rather than forcing a single diagnosis. The generalized pignistic transformation then distributes probability over the seven disjoint Venn regions, with the three-way fault intersection receiving 0.425.</p>
<p>The pump case also demonstrates the hybrid model&#8217;s power. If maintenance knowledge rules out a purely electrical fault co-occurring with a purely hydraulic one, a single constraint deletes those regions, shrinking the hyper-power set from 19 propositions to 13. Refusing the same data with the DSmH rule, which redistributes the support of forbidden intersections to the disjunctions of the involved propositions, sharpens the diagnosis: the electromechanical region&#8217;s pignistic probability rises from 0.257 to 0.521, and the Deng entropy of the fused assignment drops from 5.8371 to 4.5523 — a quantified measure of how a structural constraint reduces total uncertainty. Because the sources and masses are identical, every difference is attributable solely to the model constraint, making the library a natural laboratory for testing how modeling assumptions shape conclusions.</p>
<p>The tool has real limits, and the authors are candid about them. The free DSm hyper-power set grows explosively: with six hypotheses it contains 7,828,353 propositions, so complete free-model enumeration is recommended only up to five hypotheses, with larger problems requiring constrained hybrid models or classical DST. Venn diagram visualization caps out at three hypotheses, the implementation processes finite batches rather than streams, and the dependency-free pure-Python core, while maximally portable, avoids NumPy vectorization and is therefore not built for high-throughput workloads. The library also takes mass functions as input; deriving those masses from raw sensor data and modeling source reliability remain the user&#8217;s responsibility.</p>
<p>Even so, the significance is considerable. By turning the choice of hypothesis model and combination rule into a parameter of analysis rather than a fixture of hand-written code, evidencelib lets researchers directly study how modeling assumptions and conflict handling affect belief measures, uncertainty metrics, decision rankings, and final actions. For anyone building sensor networks, diagnostic systems, or any pipeline where imperfect sources must be reconciled, it promises to shorten the path from idea to reproducible, publishable result. The authors plan adaptive conflict-redistribution rules and more efficient fusion for large numbers of sources in future releases, and hint at an optional accelerated backend — keeping the core lean while the evidence gets heavier.</p>
<p><strong>Subject of Research:</strong> Evidence modeling and fusion under uncertainty using Dempster–Shafer and Dezert–Smarandache theories</p>
<p><strong>Article Title:</strong> evidencelib: A python library for evidence modeling and fusion under uncertainty</p>
<p><strong>Article References:</strong> Śniegowski, S., Świder, A., Shekhovtsov, A., &amp; Sałabun, W. (2026). evidencelib: A python library for evidence modeling and fusion under uncertainty. <em>SoftwareX, 36</em>, Article 103019. <a href="https://doi.org/10.1016/j.softx.2026.103019" rel="noopener noreferrer">https://doi.org/10.1016/j.softx.2026.103019</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> Not provided</p>
<p><strong>Keywords:</strong> evidencelib, Dempster–Shafer theory, Dezert–Smarandache theory, evidence fusion, uncertainty quantification, belief functions, conflict redistribution, sensor fusion, decision support, Python software, information fusion, open-source library</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">211062</post-id>	</item>
		<item>
		<title>New software brings rigorous nested validation to genomic prediction</title>
		<link>https://scienmag.com/new-software-brings-rigorous-nested-validation-to-genomic-prediction/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 22:35:10 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[CIMMYT]]></category>
		<category><![CDATA[CIMMYT genomic prediction tools]]></category>
		<category><![CDATA[crop breeding data analysis tools]]></category>
		<category><![CDATA[genomic estimated breeding values]]></category>
		<category><![CDATA[genomic prediction]]></category>
		<category><![CDATA[Genomic prediction software]]></category>
		<category><![CDATA[genomic selection]]></category>
		<category><![CDATA[hyperparameter tuning]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning for agriculture]]></category>
		<category><![CDATA[model benchmarking]]></category>
		<category><![CDATA[model evaluation in plant breeding]]></category>
		<category><![CDATA[nested cross-validation]]></category>
		<category><![CDATA[nested validation in genomic selection]]></category>
		<category><![CDATA[NV4GP]]></category>
		<category><![CDATA[open-source plant breeding software]]></category>
		<category><![CDATA[plant breeding]]></category>
		<category><![CDATA[Python software]]></category>
		<category><![CDATA[Python tools for genomic prediction]]></category>
		<category><![CDATA[quantitative genetics]]></category>
		<category><![CDATA[quantitative genetics validation methods]]></category>
		<category><![CDATA[reproducible research in crop science]]></category>
		<category><![CDATA[software for evaluating predictive models in agriculture]]></category>
		<category><![CDATA[wheat]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=203576</guid>

					<description><![CDATA[A new Python-based tool called NV4GP brings statistically rigorous nested cross-validation and independent validation to genomic prediction, allowing breeders without programming skills to fairly benchmark fifteen machine learning models.]]></description>
										<content:encoded><![CDATA[<p>Genomic prediction has quietly become one of the most consequential technologies in modern agriculture. By training statistical or machine learning models on genome-wide molecular markers paired with phenotypic records from a training population, breeders can forecast how candidate lines will perform before they are ever planted in a field. The approach, known in practice as genomic selection, has accelerated the development of improved crop varieties by allowing selection decisions to be made on predicted genomic estimated breeding values rather than on years of laborious field testing. Yet a persistent bottleneck has remained: rigorously evaluating which prediction model works best for a given trait and population, and tuning that model fairly, has demanded programming skills that many plant breeders and quantitative geneticists simply do not have. A newly released open-source tool called NV4GP, described in the journal SoftwareX, aims to close that gap.</p>
<p>NV4GP, short for Nested Validation for Genomic Prediction, was developed by Paolo Vitale of the International Maize and Wheat Improvement Centre, known as CIMMYT. The software is written entirely in Python and distributed under the permissive MIT licence through GitHub, with a permanent code archive on Zenodo. Its designers describe it as a user-friendly, efficient and reproducible framework for nested cross-validation-based model evaluation in genomic prediction, and its interface guides users through three sequential modules without requiring any command-line interaction or scripting. The motivation, the author explains, is that graduate students and early-career researchers entering plant breeding, biometrics or quantitative genetics often struggle to translate theoretical knowledge into practice because of limited coding experience, and accessible software can play a critical role in easing that transition.</p>
<p>The statistical heart of the software is its handling of hyperparameter tuning, a step that strongly influences model performance and generalization but is frequently mishandled in applied studies. Many prediction models, particularly non-parametric and artificial intelligence-based approaches, include one or more hyperparameters that are not estimated during model training. Nested cross-validation addresses this by defining an outer loop of conventional k-fold cross-validation to assess model performance, and then, within each outer-loop training set, running an inner-loop cross-validation that evaluates every combination of hyperparameters in a user-defined search grid. The combination with the lowest prediction error or the highest predictability is selected and used to fit the corresponding outer-loop model, ensuring unbiased performance estimates by preventing information leakage between model tuning and evaluation.</p>
<p>In an independent validation scenario, the software instead performs an internal k-fold cross-validation on the training population to select the best hyperparameters, then refits the model on the full training set and tests it on a genuinely separate dataset. Predictability is reported as Pearson&#8217;s correlation between predicted and observed values in the independent set. This strategy mirrors the real-world breeding context, in which breeders must use a training population to predict the performance of an unknown, genetically related breeding population. Independent validation has historically been under-applied in the genomic prediction literature, despite offering a fairer basis for model comparison than naive cross-validation schemes that inadvertently share information between training and testing data.</p>
<p>NV4GP implements fifteen prediction models spanning parametric, non-parametric and ensemble categories. The parametric set includes ridge regression, Bayesian ridge and LASSO, while the non-parametric range covers kernel ridge regression, support vector regression, elastic net, stochastic gradient descent, partial least squares, nearest neighbours, Gaussian process regression, decision trees, random forests, gradient boosting and multi-layer perceptrons, capped by a voting regressor ensemble. Each model carries its own hyperparameter grid; the multi-layer perceptron alone exposes twenty-one tunable settings, and gradient boosting and stochastic gradient descent each expose seventeen. All implementations draw on custom functions and the widely used Scikit-learn library, and a benchmark GBLUP model can be run in R through the BGLR package for comparison purposes.</p>
<p>Before any modelling begins, a marker filtering module lets users upload genotype data in numeric matrix or HapMap format, following IUPAC nucleotide nomenclature where applicable, and filter markers and genotypes by missing data, minor allele frequency and heterozygosity. Optional imputation by mean or major allele is available, and a detailed log reports exactly how many markers were removed by each filter and the per-marker quality statistics. In the validation module, users upload fully imputed marker matrices and phenotype files, which may be unbalanced because the software internally matches line identifiers and restricts analysis to common genotypes. Users specify the target trait, an optional logarithmic transformation, the hyperparameter grid, the number of cycles and folds for the outer and inner loops, a random seed for reproducibility, and the selection criterion: mean absolute error, mean squared error or predictability.</p>
<p>To demonstrate the software, the author ran five cycles of five-fold nested cross-validation, with four inner folds, on the wheat599 dataset of 599 individuals, 1,279 markers and four grain yield phenotypes recorded across environments. Predictability varied substantially across models: for the first target variable, support vector regression reached roughly 0.6 while decision trees scored 0.0; for another, the classical GBLUP benchmark led with values near 0.5, while Gaussian process regression lagged furthest behind; for a fourth, kernel ridge regression exceeded 0.5. Running times were equally revealing, ranging from about fifteen seconds for ridge regression to nearly twelve hours for the multi-layer perceptron on a conventional 64-bit laptop with 32 gigabytes of RAM. In an independent validation using published wheat data from two consecutive years and two simulated irrigation environments, predictability rankings shifted markedly between environments, ranging from slightly negative values for stochastic gradient descent to 0.32 for support vector regression in one environment and 0.16 for the multi-layer perceptron in the other.</p>
<p>The software was benchmarked against the GBLUP implementation in the BGLR R package using identical datasets and validation strategies, and produced comparable predictability estimates, supporting its correctness. Reproducibility is enforced through user-defined random seeds that fix the random state across all model components and fold partitioning, and the software was independently tested on Linux, Windows and MacOS, yielding identical results. Input validation checks flag common errors, such as misformatted marker files or incompatible cross-validation partitioning, before execution, returning informative error messages rather than silent failures. A result summary module processes output files and generates bar plots with standard error bars for test-set predictability, allowing direct visual comparison across models without additional data harmonization.</p>
<p>Compared with existing tools such as BGLR, rrBLUP, sommer, MegaLMM, ShinyGS, CHiDO and GS4PB, NV4GP is distinctive in combining a graphical interface with true nested cross-validation and independent validation, systematic grid-based hyperparameter tuning, and a broad roster of machine learning and deep learning models. The limitations are candidly acknowledged: exhaustive grid searches become computationally demanding as models, hyperparameter combinations or folds multiply, memory demands grow with large marker matrices, and the current version supports only a single trait or environment per run, deliberately excluding multi-environment and multi-trait modelling to keep the software usable on standard laptop hardware. The recommended ceiling for a conventional laptop is a search grid below fifty hyperparameter combinations on datasets of hundreds of lines with a few thousand markers.</p>
<p>The implications extend beyond convenience. By collapsing what previously required substantial bespoke programming into a uniform graphical workflow, and by standardizing reported metrics including training and test predictability, mean squared and absolute errors, percentage errors and the gap between training and test performance, NV4GP makes cross-study and cross-model comparisons of machine learning methods in genomic prediction more reproducible and more accessible. For breeders, the practical consequence is a shift in how model-selection decisions are made: rather than defaulting to a single familiar model because alternatives are impractical to implement, users can rapidly screen multiple model families and choose the empirically best performer for their specific trait and population, embedding statistically rigorous, leakage-free validation into routine genomic selection pipelines.</p>
<p><strong>Subject of Research:</strong> A Python software framework for nested cross-validation and independent validation in genomic prediction for breeding</p>
<p><strong>Article Title:</strong> NV4GP: Nested validation for genomic prediction</p>
<p><strong>Article References:</strong> Vitale, P. (2026). NV4GP: Nested validation for genomic prediction. <em>SoftwareX, 36</em>, Article 103021. <a href="https://doi.org/10.1016/j.softx.2026.103021" rel="noopener noreferrer">https://doi.org/10.1016/j.softx.2026.103021</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.softx.2026.103021" rel="noopener noreferrer">10.1016/j.softx.2026.103021</a></p>
<p><strong>Keywords:</strong> genomic prediction, genomic selection, nested cross-validation, hyperparameter tuning, machine learning, plant breeding, quantitative genetics, NV4GP, CIMMYT, wheat, Python software, model benchmarking</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">203576</post-id>	</item>
	</channel>
</rss>
