Sunday, September 20, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Cancer

AI Reads the Charts: How Open-Source OCR Could Slash Cancer Data Delays

September 20, 2026
in Cancer
Nathaniel Bowman
By Nathaniel Bowman Scienmag Editorial Profile - Precision Oncology
Reading Time: 5 mins read
0
AI Reads the Charts: How Open-Source OCR Could Slash Cancer Data Delays

AI Reads the Charts: How Open-Source OCR Could Slash Cancer Data Delays

AI Reads the Charts: How Open-Source OCR Could Slash Cancer Data Delays

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Every year, thousands of patients with early-stage breast cancer receive a genomic test result that quietly shapes the rest of their treatment. The Oncotype DX recurrence score, a number between 0 and 100 derived from a 21-gene expression assay, helps clinicians decide whether adjuvant chemotherapy is warranted. Yet despite informing real clinical decisions at the bedside, these results often take 12 to 18 months to surface in the cancer registries that power population-level research, quality improvement, and real-world evidence generation. A new study published in Cancer Causes & Control suggests that a well-configured open-source optical character recognition, or OCR, pipeline can close much of that latency gap, reading scanned reports faster and, in some respects, more accurately than the human abstraction workflows they are meant to complement.

The research, led by Qianyun Luo and Schelomo Marmor at the University of Minnesota with colleagues including Rui Zhang, Nikitha Vobugari, and Jane Y. C. Hui, tackled a deceptively simple question: can machines reliably read genomic test results that already exist inside the electronic health record? The obstacle is not that the data are missing. It is that they are trapped. Commercial genomic assays typically arrive at hospitals as scanned PDF documents or image-based files, stored in the EHR as unstructured binary objects rather than discrete, machine-readable fields. Before those results can populate a cancer registry or research database, a human must open each report, locate the recurrence score, and transcribe it. That manual workflow is slow, expensive, and vulnerable to transcription errors, and it is the principal reason genomic biomarkers lag so far behind the clinical care they inform.

To test whether automation could help, the team assembled a retrospective validation set of 675 unique Oncotype DX reports from a Midwestern U.S. health system, drawn from records stored between 2016 and 2022 within Epic, the most widely used EHR platform in the United States. Because the reports entered the record as scanned documents rather than structured data, OCR was the only viable route to automated extraction. Two independent reviewers manually abstracted all 675 recurrence scores in roughly two hours, establishing a reference standard against which both the machines and the existing cancer registry abstraction could be judged.

The study compared three fully open-source OCR approaches chosen deliberately for their cost, transparency, and reproducibility, qualities the authors argue are decisive for adoption by cancer registries and resource-constrained institutions. Tesseract, a mature engine built on a long short-term memory-based recognition architecture, runs efficiently on ordinary CPUs. EasyOCR represents a more modern deep-learning pipeline, coupling Character Region Awareness for Text Detection, known as CRAFT, with neural sequence recognition, and it benefits from GPU acceleration. The third approach, a hybrid pipeline the researchers call H-OCR, primarily used EasyOCR while falling back to Tesseract whenever the primary engine returned no text, detected no circled score, or failed to produce an integer value. This fallback design was intended to capture the complementary strengths of each engine: EasyOCR’s superior localization of the circled recurrence score annotations that clinicians routinely mark on these reports, and Tesseract’s faster, sometimes sharper character recognition.

The engineering details matter because Oncotype DX reports, while standardized, contain a subtle challenge. The recurrence score is typically circled by hand or by the reporting laboratory, and identifying which number on the page carries that annotation is part of the recognition problem. The researchers rendered PDF pages at 300 DPI, converted them to grayscale, and applied image-processing tricks tuned to each engine: Tesseract was configured with OCR Engine Mode 3 and Page Segmentation Mode 10, with Hough circle detection flagging candidate circled scores, while EasyOCR used CRAFT-based detection after images were scaled to 220 percent, Gaussian blurred, adaptively thresholded, and screened with a circularity threshold above 0.7 to isolate circled regions.

The hybrid pipeline won decisively. Across all 675 reports, H-OCR achieved 97 percent agreement with the manually abstracted reference standard, with a precision of 0.997, a recall of 0.972, and an F1 score of 0.984. Standalone EasyOCR reached 93 percent agreement with an F1 of 0.968, while Tesseract posted the highest precision of the single engines, 0.990, but lower recall at 0.934 and an F1 of 0.961. Strikingly, the conventional cancer registry abstraction matched 91 percent of manual scores, with an F1 of 0.983, meaning the automated pipeline actually outperformed the human-run process it could potentially replace or augment. Speed differences were dramatic in a different dimension: the full dataset was processed in about 1.2 hours by Tesseract, 4.8 hours by EasyOCR, and 4.9 hours by the hybrid pipeline, compared with the months-long lag that typically separates a clinical result from its registry debut.

Accuracy alone is not the whole story, because not every misread number carries clinical weight. The team therefore examined whether extraction errors crossed TAILORx risk boundaries, the clinically meaningful thresholds that separate low risk, defined as scores of 0 to 15, from intermediate risk, 16 to 25, and high risk, 26 to 100. Most discordances under the hybrid pipeline did not cross a risk category, although a subset did, and all category-crossing errors underestimated the true risk level. In registry abstraction, eight low-risk cases should have been intermediate and seven intermediate cases should have been high; under H-OCR, six low-risk cases should have been high and three should have been intermediate. The failure modes were largely mechanical, traced to misread or missing leading digits and poor fax or scan quality, with no consistent pattern suggesting a systematic blind spot.

A second layer of analysis probed whether registry discordance, the mismatch between registry-reported and manually abstracted scores, clustered among particular patient groups. Running multivariable logistic regression across 472 unique patients, the researchers tested age, race, histology, tumor grade and size, lymph node status, progesterone receptor status, chemotherapy exposure, lymphovascular invasion, rural-urban classification, and neighborhood socioeconomic deprivation. Almost none of these factors predicted discordance. The lone significant predictor was unknown progesterone receptor status, with an adjusted odds ratio of 2.88, suggesting that residual registry errors stem more from incomplete documentation and abstraction challenges than from any particular clinical subgroup. The authors read this as an encouraging sign: the benefits of automated genomic capture are likely to apply broadly across diverse patient populations rather than concentrating in any one demographic or disease profile.

The study also offered a practical mechanism for quality control. Both OCR engines emit confidence scores alongside their output, and the researchers found these scores strongly discriminated correct from incorrect extractions. Tesseract’s median confidence was 83.5 for concordant reads versus 45.0 for discordant ones, yielding an area under the ROC curve of 0.90, while EasyOCR’s confidence showed an AUC of 0.92. Within the hybrid pipeline, the 629 EasyOCR-generated extractions that were concordant carried a median confidence of 1.00, and confidence discrimination held even among the 46 Tesseract fallback extractions, with an AUC of 0.81. The authors argue that future implementations could exploit these engine-specific confidence thresholds to flag uncertain reads for targeted human review, further reducing the risk of clinically consequential misclassification while preserving the speed advantage of automation.

The implications extend well beyond a single genomic assay. Timely, structured genomic biomarkers underpin comparative effectiveness research, quality measurement, precision oncology, and the training of artificial intelligence and machine learning models that depend on current oncology data. Delayed and incomplete genomic information, the authors note, limits the ability of learning health systems to generate near real-time evidence from routine practice. Because Tesseract and EasyOCR are open-source and can run inside HIPAA-compliant environments, the approach represents a realistic, low-cost pathway for health systems, particularly lower-resourced cancer centers, to modernize their genomic data infrastructure without proprietary software or specialized hardware. The researchers are candid about limitations: the study took place at a single health system, evaluated only the highly standardized Oncotype DX report format, and did not prospectively measure how the pipeline would perform when embedded in live registry workflows. Somatic next-generation sequencing reports, with their far greater heterogeneity across vendors and formats, remain a harder and more important target for future work. Still, the demonstration stands as a concrete proof of concept that the machines can read what the clinicians already see, and that closing the data latency gap in oncology may be less a problem of missing information than of finally teaching software to look in the right place.

Subject of Research: Automated extraction of genomic biomarkers from unstructured clinical documents using open-source optical character recognition to reduce data latency in real-world oncology research

Article Title: Bridging the data latency gap: automated extraction of genomic biomarkers from unstructured clinical documents to support real-world oncology data

Article References: Bridging the data latency gap: automated extraction of genomic biomarkers from unstructured clinical documents to support real-world oncology data. (n.d.). https://doi.org/10.1007/s10552-026-02252-y

Image Credits: AI Generated

DOI: 10.1007/s10552-026-02252-y

Keywords: optical character recognition, genomic biomarkers, Oncotype DX, breast cancer, cancer registries, real-world data, electronic health records, precision oncology, data latency, EasyOCR, Tesseract, real-world evidence

Cite Scienmag News

Nathaniel Bowman. (September 20, 2026). AI Reads the Charts: How Open-Source OCR Could Slash Cancer Data Delays. Scienmag. https://scienmag.com/ai-reads-the-charts-how-open-source-ocr-could-slash-cancer-data-delays/

Nathaniel Bowman. "AI Reads the Charts: How Open-Source OCR Could Slash Cancer Data Delays." Scienmag, 20 September 2026, https://scienmag.com/ai-reads-the-charts-how-open-source-ocr-could-slash-cancer-data-delays/. Accessed 20 September 2026.

Nathaniel Bowman. "AI Reads the Charts: How Open-Source OCR Could Slash Cancer Data Delays." Scienmag. September 20, 2026. https://scienmag.com/ai-reads-the-charts-how-open-source-ocr-could-slash-cancer-data-delays/

Tags: AI in medical data processingautomated reading of medical reportsbreast cancerbreast cancer genomic testing delayscancer registriescancer research data accuracydata latencyEasyOCRelectronic health record data extractionelectronic health recordsgenomic biomarkersgenomic data extractionmachine learning in oncologyOncotype DXopen-source OCR for cancer report analysisopen-source tools for healthcare dataoptical character recognitionoptical character recognition in healthcareprecision oncologyreal-time cancer registry datareal-world dataReal-world evidencereducing healthcare data latencyTesseract
Share26Tweet16
Previous Post

Neural Geometry Field Learns Curvature-Aware Oversampling for Imbalanced Data

Next Post

Massive EHR Analysis Reveals Why Atopic Dermatitis Patients Flood Emergency Rooms Instead of Dermatology Clinics

Related Posts

Cholesterol Enzyme DHCR24 Emerges as Driver and Biomarker of Endometrial Cancer
Cancer

Cholesterol Enzyme DHCR24 Emerges as Driver and Biomarker of Endometrial Cancer

September 20, 2026
HER2-Low Breast Cancer Mirrors HER2-0 in Metastasis Timing but Confers a Survival Edge
Cancer

HER2-Low Breast Cancer Mirrors HER2-0 in Metastasis Timing but Confers a Survival Edge

September 20, 2026
Light and Radiation Combo Doubles Survival in Rat Bladder Cancer Model
Cancer

Light and Radiation Combo Doubles Survival in Rat Bladder Cancer Model

September 20, 2026
New Nomogram Predicts Which Thyroid Cancer Patients Will Fail Radioactive Iodine Therapy
Cancer

New Nomogram Predicts Which Thyroid Cancer Patients Will Fail Radioactive Iodine Therapy

September 20, 2026
Scientists Decode the Genetic Secrets of Eye Cancer in Hereford Cattle
Cancer

Scientists Decode the Genetic Secrets of Eye Cancer in Hereford Cattle

September 20, 2026
Y Chromosome Loss in Aging Men May Signal Cancer Before Tumors Form
Cancer

Y Chromosome Loss in Aging Men May Signal Cancer Before Tumors Form

September 20, 2026
Next Post
Massive EHR Analysis Reveals Why Atopic Dermatitis Patients Flood Emergency Rooms Instead of Dermatology Clinics

Massive EHR Analysis Reveals Why Atopic Dermatitis Patients Flood Emergency Rooms Instead of Dermatology Clinics

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Cholesterol Enzyme DHCR24 Emerges as Driver and Biomarker of Endometrial Cancer
  • Antidepressants and Placebo Rewire the Brain Within Two Weeks, Machine Learning Study Reveals
  • Massive EHR Analysis Reveals Why Atopic Dermatitis Patients Flood Emergency Rooms Instead of Dermatology Clinics
  • AI Reads the Charts: How Open-Source OCR Could Slash Cancer Data Delays

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading