<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>reference mapping &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/reference-mapping/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 08 Oct 2026 12:51:29 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>reference mapping &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>ArchMap brings no-code single-cell reference mapping to the web browser</title>
		<link>https://scienmag.com/archmap-brings-no-code-single-cell-reference-mapping-to-the-web-browser/</link>
		
		<dc:creator><![CDATA[Drew Townsend]]></dc:creator>
		<pubDate>Thu, 08 Oct 2026 12:51:29 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[accessible single-cell genomics]]></category>
		<category><![CDATA[ArchMap]]></category>
		<category><![CDATA[automated cell mapping]]></category>
		<category><![CDATA[bioinformatics software]]></category>
		<category><![CDATA[cell type annotation]]></category>
		<category><![CDATA[cell-type annotation platform]]></category>
		<category><![CDATA[CELLxGENE]]></category>
		<category><![CDATA[collaborative genomics platforms]]></category>
		<category><![CDATA[computational biology visualization tools]]></category>
		<category><![CDATA[democratizing single-cell analysis]]></category>
		<category><![CDATA[Human Cell Atlas]]></category>
		<category><![CDATA[Human Cell Atlas reference integration]]></category>
		<category><![CDATA[Idiopathic pulmonary fibrosis]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[no-code bioinformatics tools]]></category>
		<category><![CDATA[reference atlases in genomics]]></category>
		<category><![CDATA[reference mapping]]></category>
		<category><![CDATA[scArches]]></category>
		<category><![CDATA[scVI]]></category>
		<category><![CDATA[Single-Cell Genomics]]></category>
		<category><![CDATA[single-cell reference mapping]]></category>
		<category><![CDATA[transfer learning]]></category>
		<category><![CDATA[uncertainty estimation in single-cell data]]></category>
		<category><![CDATA[web-based single-cell data analysis]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=247786</guid>

					<description><![CDATA[A free web platform called ArchMap now allows researchers to map single-cell datasets onto published reference atlases and annotate cell types without any coding expertise.]]></description>
										<content:encoded><![CDATA[<p>Single-cell biology has been waiting for its own equivalent of the first reference genome, and a team of computational biologists argues that the missing piece was not another algorithm but a doorway. In a Brief Communication published in Nature Genetics, researchers led by Mohammad Lotfollahi, Malte D. Luecken and Fabian J. Theis introduce ArchMap, a free, web-based platform that lets scientists map new single-cell datasets onto published reference atlases without writing a single line of code. The tool, available at ArchMap.bio, packages query-to-reference mapping, automated cell-type annotation, uncertainty estimation and collaborative analysis into a single browser interface, and its developers hope it will democratize a capability that has so far been locked behind programming skills, expensive computing hardware and, in some cases, paywalls.</p>
<p>The significance of reference mapping in single-cell genomics is difficult to overstate. Large-scale integrated atlases, such as the Human Lung Cell Atlas and other tissue references accepted by the Human Cell Atlas consortium, have become central to understanding variation between cells and to establishing consensus on cell-type nomenclature. The standard way to exploit these resources is transfer learning: a researcher takes a freshly sequenced dataset, the query, and projects it onto a prebuilt reference model, inheriting the reference&#8217;s cell-type labels and biological context in much the same way that genomics researchers align new reads to the reference genome. But until now, doing this required fluency in Python, familiarity with complex machine learning pipelines and, increasingly, access to graphics processing units. Even for those with the expertise, assembling the necessary software environment is time-consuming and error-prone.</p>
<p>Existing tools have chipped away at this barrier without removing it. Azimuth offers a graphical interface for projecting new data onto atlases, but it restricts users to mapping methods from Seurat v4 and lacks collaborative features, preventing free, no-code comparison of different mapping approaches. FASTGenomics once provided access to the Human Lung Cell Atlas but is no longer freely available. The scGPT Hub and CellTypist offer graphical routes to reference mapping and cell-type annotation respectively, yet both carry caveats: scGPT is a foundation model pretrained on diverse datasets and fine-tuned on the user&#8217;s data, which can introduce biases and residual batch effects from unrelated datasets, and CellTypist alone cannot remove batch effects, so a separate integration method is still required. Foundation models have also been shown to perform worse than scVI-style variational inference methods for integration tasks while being more computationally expensive to fine-tune.</p>
<p>ArchMap was built around four guiding principles: maximizing accessibility, protecting user data privacy, ensuring utility, quality and reproducibility, and enabling collaboration. Its cloud-based setup eliminates software installations and dependency headaches entirely, and it grants automatic access to CELL×GENE Annotate for downstream exploration, effectively democratizing the full analysis pipeline from upload to visualization. Because researchers typically use the platform early in data processing, when privacy concerns are highest, all user data are encrypted during upload, download and storage on the team&#8217;s Google Cloud implementation, and projects deleted by users are automatically purged from cloud storage after three days. Expert users who prefer full control can host the web application themselves or run the mapping pipeline locally using provided tutorials, at an estimated cost of around 200 euros per month in cloud services for multiple daily mappings.</p>
<p>Under the hood, ArchMap&#8217;s mapping engine is scArches, a transfer learning framework that extends conditional neural-network-based integration methods to new query data. The platform currently supports atlases built with scVI, scANVI and scPoli, with scVI and scANVI consistently ranking among the top-performing integration methods in independent benchmarks. The number of training epochs for mapping was fixed at 100 after systematic benchmarking across healthy and diseased query datasets showed that reconstruction loss plateaus and most integration metrics stop improving beyond that point. Crucially, the platform does not retrain reference models itself, because doing so would void the extensive performance evaluations and community-based annotations that consortium-approved atlases have already undergone.</p>
<p>Once a query is mapped, ArchMap offers a choice of pretrained classifiers for cell-type label transfer, including out-of-the-box XGBoost and K-nearest neighbors implementations as well as the native scANVI and scPoli classifiers. Using 80-20 train-test splits, the team found that KNN performed best across all atlases and label granularities. The platform also computes mapping quality metrics that go beyond what existing tools provide. Euclidean uncertainty, derived from a weighted KNN classifier over the reference embedding, flags cells whose transferred labels are unreliable, with a suggested cutoff of 0.5 above which cells are labeled unknown. A second metric, Mahalanobis uncertainty, accounts for correlations between latent features by rescaling the embedding, and the team recommends it for identifying disease-associated cell states. A presence score, computed by diffusing weighted-neighbor counts through the reference graph with a random walk with restart, quantifies which reference cell states are actually represented in the query, helping users judge whether a given atlas is suitable for their data.</p>
<p>To demonstrate the workflow, the team mapped a single-cell dataset from patients with idiopathic pulmonary fibrosis onto the Human Lung Cell Atlas. Mapping quality was high, with 96.85 percent gene overlap and full preservation of Leiden clusters, scoring five out of five. Just under 54 percent of query cells found an anchor, a reference cell in a similar biological state, while 8.99 percent of cells carried uncertainty scores above 0.5 and were classified as unknown, a pattern the authors note is expected when mapping diseased tissue onto a healthy reference. Examining the results in the integrated CELL×GENE viewer, they observed striking variability in uncertainty scores among adventitial fibroblasts, a cell type implicated in pulmonary fibrosis, and found that high-uncertainty cells were strongly associated with diseased donor labels.</p>
<p>Differential expression analysis within the same interface sharpened the biological story. Low-uncertainty adventitial fibroblasts reflected homeostatic processes, whereas high-uncertainty cells displayed fibrotic signatures, including elevated expression of CTHRC1, a well-known marker of fibrotic fibroblasts, while low-uncertainty cells showed higher SERPINF1 expression. The result shows that ArchMap can surface disease-associated cell states and candidate driver genes directly within its interface, without any local analysis. The authors are careful to add a caution: high uncertainty alone is not proof of a novel or diseased state, since residual batch effects or low-quality cells with few genes or high mitochondrial content can also inflate the score, so high-uncertainty cells should be validated manually, for example through differential expression testing.</p>
<p>The platform is also designed as a community repository. Atlas builders can upload their own references, which are initially marked as in revision while automated quality checks benchmark their integration method against scVI, scPoli with and without prototype learning, and principal component analysis using scIB metrics. Atlases that pass review can be published publicly or kept private, and compatibility with the scvi-hub repository extends coverage to atlases without formal publications, provided they include a preprint or documented model evaluation. Collaboration features allow secure sharing of projects within or across teams and institutions, and joint annotation through CELL×GENE. Current server limits cap each mapping at 200,000 cells, though users can split larger queries into batches and concatenate results with provided Colab notebooks, an approach the benchmarks show yields identical performance to mapping everything at once.</p>
<p>Looking ahead, the developers see ArchMap growing alongside the atlas ecosystem itself. As new generations of atlases incorporate multimodal data from CITE-seq and spatial transcriptomics, the platform is expected to extend toward multimodal reference mapping using frameworks such as TotalVI and MultiMIL. The team also envisions support for multiple lineage atlases, allowing users to map onto specific subsets of a reference, and the evolution of its repository functions into an open-source atlas version control system as existing atlases are updated. For now, the message to the community is simple: a wet-lab biologist with a properly formatted h5ad file, raw counts, gene identifiers and a batch key can go from raw query to annotated, visualized, quality-checked mapping in a browser session. In a field where the gap between data generation and biological interpretation has often been bridged only by computational specialists, ArchMap aims to make the reference atlas as routine a tool as the reference genome became for genomics.</p>
<p><strong>Subject of Research:</strong> A no-code web platform for reference-based mapping and annotation of single-cell transcriptomics datasets</p>
<p><strong>Article Title:</strong> ArchMap is a web-based platform for reference-based analysis of single-cell datasets</p>
<p><strong>Article References:</strong> Lotfollahi, M., Bright, C., Skorobogat, R., Dehkordi, M. M., George, X., Richter, S., Shitov, V. A., Topalova, A., Melcher, H., Nussbaumer, N., Fomm, A., Luecken, M. D., &amp; Theis, F. J. (2026). ArchMap is a web-based platform for reference-based analysis of single-cell datasets. <em>Nature Genetics</em>. <a href="https://doi.org/10.1038/s41588-026-02756-y" rel="noopener noreferrer">https://doi.org/10.1038/s41588-026-02756-y</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s41588-026-02756-y" rel="noopener noreferrer">10.1038/s41588-026-02756-y</a></p>
<p><strong>Keywords:</strong> single-cell genomics, reference mapping, ArchMap, Human Cell Atlas, transfer learning, cell-type annotation, scArches, scVI, machine learning, bioinformatics software, CELLxGENE, idiopathic pulmonary fibrosis</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">247786</post-id>	</item>
	</channel>
</rss>
