A traditional Chinese medicinal herb that has been prescribed for centuries to treat coughs, asthma, and inflammatory conditions is finally getting the genomic attention it deserves. Researchers at Guizhou University of Traditional Chinese Medicine have unveiled PeuDB, a comprehensive genomic and transcriptomic platform for Peucedanum praeruptorum Dunn, a plant better known in Chinese medicine as Qianhu. The resource, described in an open-access paper in BMC Genomics, integrates two independently annotated reference genome assemblies with expression data from 135 independent biological samples spanning nine separate BioProjects, creating what the team describes as a one-stop platform for gene-function analysis in a species whose pharmacologically active compounds have long outpaced our understanding of how they are actually made.
The significance of the work lies in the gap it addresses. The roots of P. praeruptorum are rich in coumarins and flavonoids, two classes of plant secondary metabolites with well-documented pharmacological activity, yet the genetic basis underlying the biosynthesis of these compounds has remained insufficiently characterized. For medicinal plant researchers, that is a familiar frustration: the chemistry is known, the clinical tradition is deep, but the genes and regulatory networks controlling production of the active ingredients stay hidden. PeuDB is designed to change that by giving scientists the tools to move from a compound of interest to candidate genes, and from candidate genes to testable hypotheses about biosynthetic pathways.
Technically, the platform is built on a dual-assembly foundation that reflects the current reality of plant genomics. It incorporates the PP_2024v2 telomere-to-telomere assembly, representing the most complete end-to-end reconstruction of the plant’s chromosomes, alongside the PP_2024v1 chromosome-level assembly generated independently. Maintaining both assemblies side by side is a deliberate choice rather than redundancy. Different research groups have already deposited data mapped against different references, and by supporting both, the database allows users to work within whichever coordinate system their data or collaborators use. To make this practical, PeuDB links 24,846 accepted cross-assembly gene pairs, enabling identifier conversion so that a gene discovered in one assembly can be traced to its counterpart in the other.
Functional annotation of predicted genes was carried out against a battery of established reference databases, including NR, Swiss-Prot, TAIR10, TrEMBL, GO, KEGG, and Pfam. This layered annotation strategy means users can query a gene and retrieve not only its likely protein family but also its predicted molecular function, biological process, cellular component, and the metabolic pathways it may participate in. The scale of the classification effort is substantial: the team catalogued six major gene-family categories, comprising 5,981 transcription factors and transcriptional regulators, 2,916 protein kinases, 1,940 carbohydrate-active enzymes, 1,879 transporters, 3,112 ubiquitin-related proteins, and 594 cytochrome P450 genes. That last number is particularly noteworthy for natural-products research, since cytochrome P450 enzymes are the workhorses of plant secondary metabolism and frequently catalyze the tailoring steps that give medicinal compounds their final, bioactive structures.
The transcriptomic component of the database required careful statistical treatment before it could support reliable conclusions. Expression data from nine BioProjects inevitably carry batch effects, systematic technical differences between sequencing runs, laboratories, and library preparations that can masquerade as biology. The researchers applied tissue-stratified batch correction using mean-only parametric ComBat, applied separately within root, leaf, and stem tissues, with BioProject as the batch variable and estimable within-project biological-condition contrasts built into the model. Tissues that appeared in only a single BioProject, namely callus and flower, were passed through unchanged. Principal component analysis diagnostics before and after correction showed a substantial reduction, though not complete elimination, of measurable project-associated variation, a level of transparency that the authors document in their supplementary tables rather than glossing over.
With corrected expression profiles in hand, the team constructed co-expression networks using two complementary filtering approaches: direct Pearson-correlation and mutual-rank filtering. The baseline thresholds, a Pearson correlation coefficient above 0.70 combined with a mutual rank below 30, identified 427,999 positive co-expression relationships, of which 255,633 belong to the PP_2024v2 assembly and 172,366 to PP_2024v1. Importantly, the researchers did not treat these thresholds as sacred. A sensitivity analysis evaluated all nine combinations of strict correlation cutoffs from 0.65 to 0.75 and mutual-rank cutoffs from 20 to 40, quantifying how network size and overlap respond to parameter choices. This kind of threshold-sensitivity reporting is increasingly seen as a hallmark of rigorous network biology, since it tells downstream users exactly how much confidence to place in any individual predicted relationship.
Beyond co-expression, PeuDB extends into protein-interaction prediction through a strategy borrowed from comparative genomics. The team projected evidence-filtered Arabidopsis interaction data onto P. praeruptorum using strict reciprocal-best-hit analysis, yielding 5,864 confidence-filtered interolog predictions, split as 3,027 for the T2T assembly and 2,837 for the chromosome-level assembly. The logic is that if two proteins are known to interact in the well-studied reference plant Arabidopsis thaliana, and their best-matching counterparts in P. praeruptorum are each other’s best matches in turn, the interaction is likely conserved. It is an inference, not a measurement, and the database is explicit about that distinction, presenting these predictions as computational hypotheses awaiting experimental validation.
To demonstrate the platform in action, the authors built two reproducible case-study workflows, one centered on the F6’H enzyme family and another on WRKY transcription factors. Both are directly relevant to the plant’s medicinal value: feruloyl 6-hydroxylase enzymes participate in coumarin biosynthesis, while WRKY transcription factors are candidate regulators of secondary-metabolite pathways. The case studies walk users through candidate prioritization while carefully distinguishing what the computational evidence supports from what still requires laboratory confirmation. Notably, the WRKY analysis showed that biosynthesis of unsaturated fatty acids remained significantly enriched across all nine threshold combinations tested, while other baseline-significant pathways and individual network edges proved threshold-dependent, a finding the authors report honestly rather than burying.
The platform itself runs on a LAMP architecture, the long-established combination of Linux, Apache, MySQL, and PHP, and offers a user-friendly web interface with an unusually complete analytical toolkit. Users can perform BLAST searches, browse genomes with JBrowse2, generate expression heatmaps, explore co-expression and predicted protein-interaction networks, run enrichment and differential-expression analyses, design primers, convert identifiers between assemblies, and conduct multiple-sequence alignments. For programmatic users, PeuDB provides a versioned, read-only REST API with OpenAPI/Swagger documentation, enabling automated retrieval of assembly information, gene annotations and sequences, raw TPM expression profiles, cross-assembly mappings, co-expression relationships, and predicted protein interactions. That API layer matters for reproducibility, allowing pipelines to query the database directly rather than relying on manual downloads that quickly go stale.
PeuDB is freely accessible at https://www.gzybioinformatics.cn/PeuDB, and its arrival reflects a broader shift in how medicinal plant science is done. The era in which a single reference genome and a handful of transcriptomes could anchor a research program is giving way to integrated platforms that manage multiple assemblies, large sample collections, and the statistical corrections needed to make heterogeneous public data usable. For P. praeruptorum, the practical payoff could be faster identification of the genes governing coumarin and flavonoid production, with implications for breeding, cultivation, and potentially metabolic engineering of a herb whose therapeutic reputation spans centuries. For the wider community of researchers working on under-characterized medicinal species, the database offers a template: assemble thoroughly, annotate broadly, correct for batch effects rigorously, report threshold sensitivity honestly, and always keep the line between computational prediction and experimental proof clearly visible.
Subject of Research: A genomic and transcriptomic database for gene-function and coumarin biosynthesis analysis in the medicinal plant Peucedanum praeruptorum
Article Title: PeuDB: a comprehensive genomic and transcriptomic platform for gene function analysis in Peucedanum praeruptorum
Article References: Yang, J., Yang, H., Shu, G., Xiao, Q., & Yang, J. (2026). PeuDB: a comprehensive genomic and transcriptomic platform for gene function analysis in Peucedanum praeruptorum. BMC Genomics. https://doi.org/10.1186/s12864-026-13302-9
Image Credits: AI Generated
DOI: 10.1186/s12864-026-13302-9
Keywords: Peucedanum praeruptorum, PeuDB, genomics, transcriptomics, co-expression network, coumarin biosynthesis, cytochrome P450, telomere-to-telomere assembly, gene annotation, protein interaction prediction, medicinal plants, biological database
Cite Scienmag News
Juliet Wilcox. (October 3, 2026). New Genomic Database PeuDB Aims to Decode the Medicinal Chemistry of a Classic Chinese Herb. Scienmag. https://scienmag.com/new-genomic-database-peudb-aims-to-decode-the-medicinal-chemistry-of-a-classic-chinese-herb/
Juliet Wilcox. "New Genomic Database PeuDB Aims to Decode the Medicinal Chemistry of a Classic Chinese Herb." Scienmag, 3 October 2026, https://scienmag.com/new-genomic-database-peudb-aims-to-decode-the-medicinal-chemistry-of-a-classic-chinese-herb/. Accessed 3 October 2026.
Juliet Wilcox. "New Genomic Database PeuDB Aims to Decode the Medicinal Chemistry of a Classic Chinese Herb." Scienmag. October 3, 2026. https://scienmag.com/new-genomic-database-peudb-aims-to-decode-the-medicinal-chemistry-of-a-classic-chinese-herb/

