<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>cancer registries &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/cancer-registries/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 24 Sep 2026 23:27:05 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>cancer registries &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Why Europe&#8217;s Cancer Registries Are Too Frail to Track the Fight Against Cancer</title>
		<link>https://scienmag.com/why-europes-cancer-registries-are-too-frail-to-track-the-fight-against-cancer/</link>
		
		<dc:creator><![CDATA[Nathaniel Bowman]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 23:27:05 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[Beating Cancer Plan]]></category>
		<category><![CDATA[cancer data collection]]></category>
		<category><![CDATA[cancer incidence]]></category>
		<category><![CDATA[cancer incidence and survival statistics]]></category>
		<category><![CDATA[cancer registries]]></category>
		<category><![CDATA[cancer registry coverage challenges]]></category>
		<category><![CDATA[cancer research and epidemiology]]></category>
		<category><![CDATA[cancer surveillance]]></category>
		<category><![CDATA[cancer surveillance and monitoring]]></category>
		<category><![CDATA[cancer survival]]></category>
		<category><![CDATA[data quality in cancer registries]]></category>
		<category><![CDATA[digitalisation of cancer registries]]></category>
		<category><![CDATA[ECIS]]></category>
		<category><![CDATA[Europe]]></category>
		<category><![CDATA[European Cancer Information System]]></category>
		<category><![CDATA[European cancer registries]]></category>
		<category><![CDATA[European Health Data Space]]></category>
		<category><![CDATA[European Network of Cancer Registries]]></category>
		<category><![CDATA[health data infrastructure]]></category>
		<category><![CDATA[health policy]]></category>
		<category><![CDATA[health policy for cancer control]]></category>
		<category><![CDATA[population-based cancer control]]></category>
		<category><![CDATA[population-based cancer registries]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=213303</guid>

					<description><![CDATA[A commentary in The Lancet Regional Health – Europe warns that Europe's population-based cancer registries lack the funding, regulation and infrastructure needed to monitor the continent's cancer burden, and sets out policy reforms at EU, national and registry level.]]></description>
										<content:encoded><![CDATA[<p>Every cancer diagnosis in Europe should, in principle, leave a trace in a population-based cancer registry. These registries are the quiet machinery of cancer control: they systematically and continuously collect data on every cancer case occurring within a defined population, following international standards designed to guarantee harmonisation, data quality and complete coverage. Operating at regional or national level, they tell governments how many people are diagnosed with cancer each year, how many are still alive after treatment, and how many are living with the disease. Together with mortality statistics, the three indicators they produce—incidence, survival and prevalence—are considered the essential elements of population-based cancer control. Without them, policymakers are effectively navigating one of the continent&#8217;s biggest health challenges with the lights switched off.</p>
<p>A new commentary published in The Lancet Regional Health – Europe argues that this machinery is in trouble. Written by Gijs Geleijnse of the Netherlands Comprehensive Cancer Organisation and colleagues from cancer registries and research institutions across the continent, the piece lays out a stark diagnosis: despite decades of investment in healthcare digitalisation and the ongoing preparation of the European Health Data Space, the infrastructure underpinning Europe&#8217;s cancer registries is frail. Coverage of the continent remains suboptimal, the timeliness of data publication varies widely from country to country, and registries themselves repeatedly cite resource limitations as the main barrier to investing in innovation. The result is a system that cannot reliably answer even basic questions about the state of cancer in Europe today.</p>
<p>The stakes are rising fast. Through the European Cancer Information System, known as ECIS, the registries of the European Network of Cancer Registries reveal cancer inequalities within and between European geographies, showing where progress is being made and where attention is urgently required. But the financial context is shifting beneath them. Cancer spending in Europe is expected to rise by 59 percent by 2050, according to an analytical report from the OECD and the European Commission. Registries are uniquely positioned to guide how that money is spent—supporting the effective allocation of resources and the implementation and evaluation of cancer prevention, early detection and quality of care—yet only if the data they deliver are complete, comparable and current.</p>
<p>One of the most striking asymmetries highlighted in the commentary is regulatory. Many European countries have national legislation on cancer registration, but there is no European regulation governing PBCR-based statistics. That stands in sharp contrast to mortality statistics, which are governed by explicit EU regulation and delivered by national statistics bureaus and Eurostat, ensuring timely and comparable collection of cause-of-death data across the Union. In other words, Europe legally guarantees that it knows how many people die of cancer and where, but not that it knows how many are diagnosed, how they are treated, or whether they survive. For a continent that has made cancer control a flagship policy priority, the gap is difficult to justify.</p>
<p>The technical picture is equally uneven. Europe&#8217;s 192 population-based cancer registries follow the registration guidelines issued by the European Commission&#8217;s Joint Research Centre together with the ENCR, but their capacity to collect timely, high-quality data varies enormously. Some registries struggle to record basic clinical elements such as stage at diagnosis and the treatments patients actually received—variables that are indispensable for measuring early detection programmes and the quality of care. A survey of registries published in the International Journal of Cancer in 2025 documented this global capacity gap, and recent work mapping European registries has shown that coverage, data availability and the ability to generate real-world evidence differ substantially across the continent.</p>
<p>The timing of the warning matters. The Joint Action CancerWatch, running from 2025 to 2028 with the participation of 92 organisations from 29 countries, aims specifically to improve the timeliness and quality of the registry data feeding into ECIS. The project exists precisely because geographic coverage is incomplete and timely indicators are limited on the platform. Meanwhile, the European Court of Auditors has called for a monitoring framework for the European Commission&#8217;s Beating Cancer Plan, noting in a 2026 special report that the wide-ranging plan faces an uncertain future. ECIS was established to monitor the cancer burden in Europe, but the current limitations in registry data infrastructure undermine its ability to fulfil that role. A monitoring plan without a functioning measurement system, the authors imply, is a promise without a receipt.</p>
<p>To close the gap, the commentary sets out a layered set of policy options. At EU level, the authors call for formal recognition of population-based cancer registries as core components of public health systems, an endorsement that would echo the Council Recommendation on strengthening prevention through early detection, which gave EU cancer screening programmes a firm regulatory footing. The European Commission, they argue, should publish explicit criteria for the data quality and timeliness required for registry data on ECIS, so that the quality and progress of cancer registration in member states can be monitored and managed transparently. Under the European Health Data Space Regulation, they further propose a data usage fee for registries that prioritise delivering data and statistics to ECIS—turning registries from passive data suppliers into recognised, resourced participants in the European data economy.</p>
<p>The integration agenda goes further. Registries, the authors contend, should be built into cancer screening programmes and comprehensive cancer centres as an integral element rather than an afterthought, and future EU project proposals on these topics should require the involvement and adequate resourcing of registries for planning, monitoring and evaluation. At member state level, each country should designate an organisation responsible for delivering national registry data to ECIS, acting as the link between EU bodies and regional registries and embedded in the design and monitoring of the national Cancer Mission Hubs. Sustainable funding is only part of the answer: national innovation funds earmarked for artificial intelligence and digital sovereignty should, the authors argue, also support innovation within registries, which are precisely the kind of high-value, privacy-preserving data infrastructure those funds are meant to cultivate.</p>
<p>At the registry level itself, the recommendations are more sober but no less important. Registries should obtain a formal mandate from their governments as a core component of public health and cancer control, giving them the legal standing to negotiate data access and funding. They should also allocate resources to increase the efficiency of data collection and publication—modernising the often manual, fragmented workflows that delay the arrival of statistics by years. The technical direction of travel is clear: automated extraction from electronic health records, standardised coding, and harmonised quality assurance could shorten the lag between diagnosis and data, provided the underlying legal and financial foundations are secure.</p>
<p>The commentary closes with a sentence that doubles as its thesis: we need to count every cancer patient because every patient counts. Timely, rich and high-quality population-based cancer data, the authors argue, are essential for robust monitoring and for evidence-informed European, national and local cancer policies. As Europe prepares to spend hundreds of billions more on cancer care over the coming decades, the unglamorous work of counting cases, staging tumours and tracking survival may prove to be the highest-yield investment of all. The alternative—policy made in the dark—would be far more expensive, and far less equitable, than the registries themselves.</p>
<p><strong>Subject of Research:</strong> Strengthening population-based cancer registries in Europe to improve cancer control and monitoring</p>
<p><strong>Article Title:</strong> Every cancer patient counts: strengthening European population-based cancer registries to improve cancer control</p>
<p><strong>Article References:</strong> Geleijnse, G., Backes, C., Chirlaque, M. D., Sloep, M., Rodon Navarro, E., Šekerija, M., van Eycken, L., &amp; Ursin, G. (2026). Every cancer patient counts: strengthening European population-based cancer registries to improve cancer control. <em>The Lancet Regional Health &#8211; Europe, 70</em>, Article 101855. <a href="https://doi.org/10.1016/j.lanepe.2026.101855" rel="noopener noreferrer">https://doi.org/10.1016/j.lanepe.2026.101855</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.lanepe.2026.101855" rel="noopener noreferrer">10.1016/j.lanepe.2026.101855</a></p>
<p><strong>Keywords:</strong> cancer registries, population-based cancer registries, European Cancer Information System, ECIS, European Network of Cancer Registries, cancer surveillance, Beating Cancer Plan, European Health Data Space, cancer incidence, cancer survival, health policy, Europe</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">213303</post-id>	</item>
		<item>
		<title>AI Reads the Charts: How Open-Source OCR Could Slash Cancer Data Delays</title>
		<link>https://scienmag.com/ai-reads-the-charts-how-open-source-ocr-could-slash-cancer-data-delays/</link>
		
		<dc:creator><![CDATA[Nathaniel Bowman]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 22:40:25 +0000</pubDate>
				<category><![CDATA[Cancer]]></category>
		<category><![CDATA[AI in medical data processing]]></category>
		<category><![CDATA[automated reading of medical reports]]></category>
		<category><![CDATA[breast cancer]]></category>
		<category><![CDATA[breast cancer genomic testing delays]]></category>
		<category><![CDATA[cancer registries]]></category>
		<category><![CDATA[cancer research data accuracy]]></category>
		<category><![CDATA[data latency]]></category>
		<category><![CDATA[EasyOCR]]></category>
		<category><![CDATA[electronic health record data extraction]]></category>
		<category><![CDATA[electronic health records]]></category>
		<category><![CDATA[genomic biomarkers]]></category>
		<category><![CDATA[genomic data extraction]]></category>
		<category><![CDATA[machine learning in oncology]]></category>
		<category><![CDATA[Oncotype DX]]></category>
		<category><![CDATA[open-source OCR for cancer report analysis]]></category>
		<category><![CDATA[open-source tools for healthcare data]]></category>
		<category><![CDATA[optical character recognition]]></category>
		<category><![CDATA[optical character recognition in healthcare]]></category>
		<category><![CDATA[precision oncology]]></category>
		<category><![CDATA[real-time cancer registry data]]></category>
		<category><![CDATA[real-world data]]></category>
		<category><![CDATA[Real-world evidence]]></category>
		<category><![CDATA[reducing healthcare data latency]]></category>
		<category><![CDATA[Tesseract]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=203624</guid>

					<description><![CDATA[A University of Minnesota study found that a hybrid open-source OCR pipeline extracted Oncotype DX breast cancer recurrence scores from scanned clinical reports with 97 percent agreement with manual abstraction, outperforming cancer registry abstraction while processing 675 reports in hours instead of the 12 to 18 months typical of registry workflows.]]></description>
										<content:encoded><![CDATA[<p>Every year, thousands of patients with early-stage breast cancer receive a genomic test result that quietly shapes the rest of their treatment. The Oncotype DX recurrence score, a number between 0 and 100 derived from a 21-gene expression assay, helps clinicians decide whether adjuvant chemotherapy is warranted. Yet despite informing real clinical decisions at the bedside, these results often take 12 to 18 months to surface in the cancer registries that power population-level research, quality improvement, and real-world evidence generation. A new study published in Cancer Causes &amp; Control suggests that a well-configured open-source optical character recognition, or OCR, pipeline can close much of that latency gap, reading scanned reports faster and, in some respects, more accurately than the human abstraction workflows they are meant to complement.</p>
<p>The research, led by Qianyun Luo and Schelomo Marmor at the University of Minnesota with colleagues including Rui Zhang, Nikitha Vobugari, and Jane Y. C. Hui, tackled a deceptively simple question: can machines reliably read genomic test results that already exist inside the electronic health record? The obstacle is not that the data are missing. It is that they are trapped. Commercial genomic assays typically arrive at hospitals as scanned PDF documents or image-based files, stored in the EHR as unstructured binary objects rather than discrete, machine-readable fields. Before those results can populate a cancer registry or research database, a human must open each report, locate the recurrence score, and transcribe it. That manual workflow is slow, expensive, and vulnerable to transcription errors, and it is the principal reason genomic biomarkers lag so far behind the clinical care they inform.</p>
<p>To test whether automation could help, the team assembled a retrospective validation set of 675 unique Oncotype DX reports from a Midwestern U.S. health system, drawn from records stored between 2016 and 2022 within Epic, the most widely used EHR platform in the United States. Because the reports entered the record as scanned documents rather than structured data, OCR was the only viable route to automated extraction. Two independent reviewers manually abstracted all 675 recurrence scores in roughly two hours, establishing a reference standard against which both the machines and the existing cancer registry abstraction could be judged.</p>
<p>The study compared three fully open-source OCR approaches chosen deliberately for their cost, transparency, and reproducibility, qualities the authors argue are decisive for adoption by cancer registries and resource-constrained institutions. Tesseract, a mature engine built on a long short-term memory-based recognition architecture, runs efficiently on ordinary CPUs. EasyOCR represents a more modern deep-learning pipeline, coupling Character Region Awareness for Text Detection, known as CRAFT, with neural sequence recognition, and it benefits from GPU acceleration. The third approach, a hybrid pipeline the researchers call H-OCR, primarily used EasyOCR while falling back to Tesseract whenever the primary engine returned no text, detected no circled score, or failed to produce an integer value. This fallback design was intended to capture the complementary strengths of each engine: EasyOCR&#8217;s superior localization of the circled recurrence score annotations that clinicians routinely mark on these reports, and Tesseract&#8217;s faster, sometimes sharper character recognition.</p>
<p>The engineering details matter because Oncotype DX reports, while standardized, contain a subtle challenge. The recurrence score is typically circled by hand or by the reporting laboratory, and identifying which number on the page carries that annotation is part of the recognition problem. The researchers rendered PDF pages at 300 DPI, converted them to grayscale, and applied image-processing tricks tuned to each engine: Tesseract was configured with OCR Engine Mode 3 and Page Segmentation Mode 10, with Hough circle detection flagging candidate circled scores, while EasyOCR used CRAFT-based detection after images were scaled to 220 percent, Gaussian blurred, adaptively thresholded, and screened with a circularity threshold above 0.7 to isolate circled regions.</p>
<p>The hybrid pipeline won decisively. Across all 675 reports, H-OCR achieved 97 percent agreement with the manually abstracted reference standard, with a precision of 0.997, a recall of 0.972, and an F1 score of 0.984. Standalone EasyOCR reached 93 percent agreement with an F1 of 0.968, while Tesseract posted the highest precision of the single engines, 0.990, but lower recall at 0.934 and an F1 of 0.961. Strikingly, the conventional cancer registry abstraction matched 91 percent of manual scores, with an F1 of 0.983, meaning the automated pipeline actually outperformed the human-run process it could potentially replace or augment. Speed differences were dramatic in a different dimension: the full dataset was processed in about 1.2 hours by Tesseract, 4.8 hours by EasyOCR, and 4.9 hours by the hybrid pipeline, compared with the months-long lag that typically separates a clinical result from its registry debut.</p>
<p>Accuracy alone is not the whole story, because not every misread number carries clinical weight. The team therefore examined whether extraction errors crossed TAILORx risk boundaries, the clinically meaningful thresholds that separate low risk, defined as scores of 0 to 15, from intermediate risk, 16 to 25, and high risk, 26 to 100. Most discordances under the hybrid pipeline did not cross a risk category, although a subset did, and all category-crossing errors underestimated the true risk level. In registry abstraction, eight low-risk cases should have been intermediate and seven intermediate cases should have been high; under H-OCR, six low-risk cases should have been high and three should have been intermediate. The failure modes were largely mechanical, traced to misread or missing leading digits and poor fax or scan quality, with no consistent pattern suggesting a systematic blind spot.</p>
<p>A second layer of analysis probed whether registry discordance, the mismatch between registry-reported and manually abstracted scores, clustered among particular patient groups. Running multivariable logistic regression across 472 unique patients, the researchers tested age, race, histology, tumor grade and size, lymph node status, progesterone receptor status, chemotherapy exposure, lymphovascular invasion, rural-urban classification, and neighborhood socioeconomic deprivation. Almost none of these factors predicted discordance. The lone significant predictor was unknown progesterone receptor status, with an adjusted odds ratio of 2.88, suggesting that residual registry errors stem more from incomplete documentation and abstraction challenges than from any particular clinical subgroup. The authors read this as an encouraging sign: the benefits of automated genomic capture are likely to apply broadly across diverse patient populations rather than concentrating in any one demographic or disease profile.</p>
<p>The study also offered a practical mechanism for quality control. Both OCR engines emit confidence scores alongside their output, and the researchers found these scores strongly discriminated correct from incorrect extractions. Tesseract&#8217;s median confidence was 83.5 for concordant reads versus 45.0 for discordant ones, yielding an area under the ROC curve of 0.90, while EasyOCR&#8217;s confidence showed an AUC of 0.92. Within the hybrid pipeline, the 629 EasyOCR-generated extractions that were concordant carried a median confidence of 1.00, and confidence discrimination held even among the 46 Tesseract fallback extractions, with an AUC of 0.81. The authors argue that future implementations could exploit these engine-specific confidence thresholds to flag uncertain reads for targeted human review, further reducing the risk of clinically consequential misclassification while preserving the speed advantage of automation.</p>
<p>The implications extend well beyond a single genomic assay. Timely, structured genomic biomarkers underpin comparative effectiveness research, quality measurement, precision oncology, and the training of artificial intelligence and machine learning models that depend on current oncology data. Delayed and incomplete genomic information, the authors note, limits the ability of learning health systems to generate near real-time evidence from routine practice. Because Tesseract and EasyOCR are open-source and can run inside HIPAA-compliant environments, the approach represents a realistic, low-cost pathway for health systems, particularly lower-resourced cancer centers, to modernize their genomic data infrastructure without proprietary software or specialized hardware. The researchers are candid about limitations: the study took place at a single health system, evaluated only the highly standardized Oncotype DX report format, and did not prospectively measure how the pipeline would perform when embedded in live registry workflows. Somatic next-generation sequencing reports, with their far greater heterogeneity across vendors and formats, remain a harder and more important target for future work. Still, the demonstration stands as a concrete proof of concept that the machines can read what the clinicians already see, and that closing the data latency gap in oncology may be less a problem of missing information than of finally teaching software to look in the right place.</p>
<p><strong>Subject of Research:</strong> Automated extraction of genomic biomarkers from unstructured clinical documents using open-source optical character recognition to reduce data latency in real-world oncology research</p>
<p><strong>Article Title:</strong> Bridging the data latency gap: automated extraction of genomic biomarkers from unstructured clinical documents to support real-world oncology data</p>
<p><strong>Article References:</strong> Bridging the data latency gap: automated extraction of genomic biomarkers from unstructured clinical documents to support real-world oncology data. (n.d.). <a href="https://doi.org/10.1007/s10552-026-02252-y" rel="noopener noreferrer">https://doi.org/10.1007/s10552-026-02252-y</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10552-026-02252-y" rel="noopener noreferrer">10.1007/s10552-026-02252-y</a></p>
<p><strong>Keywords:</strong> optical character recognition, genomic biomarkers, Oncotype DX, breast cancer, cancer registries, real-world data, electronic health records, precision oncology, data latency, EasyOCR, Tesseract, real-world evidence</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">203624</post-id>	</item>
	</channel>
</rss>
