<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>multi-jurisdictional crime data standardization &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/multi-jurisdictional-crime-data-standardization/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 11 Sep 2026 01:20:27 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>multi-jurisdictional crime data standardization &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New semi-supervised tool streamlines coding of multi-jurisdiction criminal justice data</title>
		<link>https://scienmag.com/new-semi-supervised-tool-streamlines-coding-of-multi-jurisdiction-criminal-justice-data/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Fri, 11 Sep 2026 01:20:23 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[automated offense code classification]]></category>
		<category><![CDATA[cost-effective crime data automation]]></category>
		<category><![CDATA[cost-effective NLP tools for police records]]></category>
		<category><![CDATA[crime data coding automation]]></category>
		<category><![CDATA[criminal justice data standardization]]></category>
		<category><![CDATA[development of OTAC crime coding tool]]></category>
		<category><![CDATA[improvement of national crime statistics accuracy]]></category>
		<category><![CDATA[improving accuracy of criminal offense data]]></category>
		<category><![CDATA[linguistic normalization of free-text crime descriptions]]></category>
		<category><![CDATA[machine learning for law enforcement data]]></category>
		<category><![CDATA[machine learning for legal and law enforcement data]]></category>
		<category><![CDATA[multi-jurisdictional crime data coding tools]]></category>
		<category><![CDATA[multi-jurisdictional crime data standardization]]></category>
		<category><![CDATA[natural language processing in criminal justice]]></category>
		<category><![CDATA[offense description standardization]]></category>
		<category><![CDATA[RTI International crime data research]]></category>
		<category><![CDATA[RTI International criminal justice research]]></category>
		<category><![CDATA[semi-supervised machine learning in criminal justice]]></category>
		<category><![CDATA[semi-supervised natural language processing for crime data]]></category>
		<category><![CDATA[semi-supervised natural language processing for criminal justice data]]></category>
		<category><![CDATA[transformer-based modeling in criminal justice]]></category>
		<category><![CDATA[transformer-based NLP models for law enforcement]]></category>
		<category><![CDATA[validation of automated crime classification systems]]></category>
		<guid isPermaLink="false">https://scienmag.com/new-semi-supervised-tool-streamlines-coding-of-multi-jurisdiction-criminal-justice-data/</guid>

					<description><![CDATA[Every day, across thousands of police departments, courthouses, and correctional agencies in the United States, officers and clerks type short free-text descriptions of alleged crimes into case management systems: &#8220;assault w/ weapon,&#8221; &#8220;BURGLARY RES DWELLING NITE,&#8221; &#8220;poss meth amt undetermined.&#8221; These fragments are the raw material of national crime statistics, yet they are notoriously inconsistent, [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>Every day, across thousands of police departments, courthouses, and correctional agencies in the United States, officers and clerks type short free-text descriptions of alleged crimes into case management systems: &#8220;assault w/ weapon,&#8221; &#8220;BURGLARY RES DWELLING NITE,&#8221; &#8220;poss meth amt undetermined.&#8221; These fragments are the raw material of national crime statistics, yet they are notoriously inconsistent, riddled with local abbreviations, typos, and jurisdiction-specific shorthand. A new study published in the American Journal of Criminal Justice shows that modern natural language processing can tame this linguistic chaos, converting free-text offense descriptions into standardized offense codes with robust accuracy, precision, and recall, and doing so cheaply enough that the researchers trained their models on an ordinary laptop.</p>
<p>The work, led by Matthew DeMichele, Ian A. Silver, and Alexander J. Preiss of RTI International&#8217;s Center for Legal Systems Research, together with Peter Baumgartner, describes the development and evaluation of a semi-supervised automated tool for coding multi-jurisdictional criminal justice data. The tool, known as the Offense Text Auto Classifier (OTAC), builds on an earlier RTI initiative called the Rapid Offense Text Autocoder, extending that foundation with newer transformer-based modeling tools, expanded training data, and enhanced validation strategies. The central problem the team attacked is one of the oldest and most stubborn barriers in criminology: the absence of a common language for crime.</p>
<p>Unlike healthcare, which long ago adopted standardized vocabularies such as diagnostic classification systems that allow hospitals and researchers to compare cases across the country, the American criminal justice system lacks a unified offense taxonomy. Each state defines crimes differently, and each agency records them in its own idiom. The National Incident-Based Reporting System and the National Crime Victimization Survey provide broad national pictures, but both have well-documented limitations, and efforts such as the National Academies&#8217; two-volume report on modernizing crime statistics have repeatedly called for better measurement infrastructure. When data from two jurisdictions are pooled, researchers must either laboriously hand-code offense narratives or accept that apples are being counted alongside oranges. Hand-coding is expensive, slow, error-prone, and impossible to replicate exactly, and clerical errors in the justice system have had real consequences, including cases in which a bail hearing went unrecorded and a prisoner was released early because of a clerical mistake.</p>
<p>The consequences of fragmented data ripple outward. Jurisdictions cannot reliably compare themselves to peers. Policymakers cannot answer basic questions about whether a pretrial reform is working. Agencies spend enormous sums on manual data preparation; studies of mass incarceration costs have emphasized how much of the system&#8217;s budget is consumed by labor that better data infrastructure could streamline. The authors frame the problem in terms borrowed from medicine: what criminal justice needs is its own version of real-world evidence, the framework that transformed decision-making in healthcare by systematically harvesting data from actual clinical practice rather than from narrow trials. Realizing that vision requires offense data that are consistent, comparable, and available quickly enough to inform ongoing decisions.</p>
<p>The technical core of the new study is a comparison of machine learning approaches for mapping raw offense text to structured offense codes. The researchers worked with data drawn from two study jurisdictions that differ in scale and regional context: a large metropolitan county in the Northeast with roughly 1.2 million residents and a mid-sized metropolitan county in the Midwest with about 560,000 residents. Both are urban jurisdictions with crime levels broadly consistent with similarly situated metropolitan counties, and the contrast between them was deliberate, since any auto-coding tool that only works on one county&#8217;s dialect is useless for national harmonization.</p>
<p>Before any modeling could begin, the team subjected the free text to extensive preprocessing. Text normalization cleaned the strings by standardizing case, spacing, and punctuation; tokenization split raw text into numerically encoded units, whole words, subwords, or characters, each mapped to unique numerical identifiers from a predefined vocabulary so that models could process the input in structured form. Deduplication removed redundant records so that repeated identical offense strings would not inflate training sets, and rigorous quality control procedures audited the resulting labels. Label noise is a known killer of supervised text classifiers, and the team drew on techniques from the machine learning literature, including confident learning methods for estimating uncertainty in dataset labels, to make sure that the human-coded ground truth used to train and evaluate the models was itself trustworthy.</p>
<p>The modeling strategy deliberately spanned a range of architectures and data-efficiency regimes. On one end sat FastText, an efficient text classification method that enriches word vectors with subword information, allowing it to make reasonable guesses about rare or misspelled words by breaking them into character n-grams, an appealing property given the typo-rich reality of offense narratives. On the other end sat transformer-based models, the architecture that has redefined natural language processing since the introduction of self-attention mechanisms. Transformers replace the sequential recurrence of older recurrent neural networks with attention mechanisms that let a model weigh the relevance of every token to every other token in an input, capturing context that bag-of-words approaches miss. The team used DistilRoberta, a distilled variant of the robustly optimized BERT pretraining approach, retaining much of the accuracy of larger models at a fraction of the computational cost.</p>
<p>A key innovation was the use of parameter-efficient fine-tuning. Retraining large language models normally demands expensive hardware, but the researchers applied LoRA, low-rank adaptation of large language models, to fine-tune DistilRoberta with far fewer trainable parameters. This mattered enormously for practicality: all models in the study were trained on an M1 MacBook Pro, with FastText and SetFit models running on the CPU and DistilRoberta on the laptop&#8217;s built-in GPU via the Metal Performance Shaders backend. The ability to retrain models easily, without large computational expense or restrictive hardware requirements, was an explicit design goal, since a harmonization tool that only a well-funded lab can operate defeats its own purpose. SetFit, a few-shot learning framework that produces strong classifiers from dozens rather than thousands of labeled examples, rounded out the portfolio, addressing the chronic scarcity of hand-coded offense data. A hierarchical multilayer perceptron, an artificial neural network using multiple layers of interconnected units to classify complex inputs, was also part of the evaluation landscape.</p>
<p>Evaluation of the models revealed that the transformer-based approaches achieved robust accuracy, precision, and recall, the standard metrics of classification performance, across the multi-jurisdictional test data. The findings underscore the potential of pretrained models to handle the linguistic diversity and inconsistency that characterizes justice system data. Pretraining on massive general corpora, followed by targeted fine-tuning on a modest number of locally coded examples, an approach grounded in transfer learning theory, lets the model bring general language knowledge to a domain where labeled data are scarce and idiosyncratic. Few-shot learning complements this by allowing the system to adapt to a new jurisdiction&#8217;s vocabulary with only dozens of examples instead of hundreds or thousands, which is exactly the economics that interagency data-sharing projects need.</p>
<p>The implications extend beyond academic convenience. First, auto-coding enables greater consistency and comparability across jurisdictions, reducing the errors that arise from fragmented coding practices and making pooled multi-site analyses statistically meaningful. Second, by automating a labor-intensive process, it expands research capacity and lowers costs, making high-quality data accessible to agencies and policymakers who cannot afford armies of hand-coders. Third, improved standardization supports real-time evidence generation, aligning criminal justice data with real-world evidence frameworks that have transformed decision-making in healthcare; a jail or pretrial agency could, in principle, monitor outcomes continuously rather than waiting years for retrospective studies. Finally, the authors argue that adopting offense auto-coding provides a foundation for interoperable, timely, and actionable data systems capable of supporting evidence-based interventions and equitable policy reforms.</p>
<p>The study aligns with parallel national efforts to modernize criminal justice data infrastructure, including initiatives to harmonize records across systems and federal guidance on standardizing incident data. The research was supported by Grant No. 2020-85-CX-K002 awarded by the Bureau of Justice Statistics, Office of Justice Programs, U.S. Department of Justice, though the authors note that the points of view expressed are theirs alone and do not necessarily represent the official position or policies of the Department of Justice. The authors report no competing interests, and the corresponding author is Matthew DeMichele of RTI International.</p>
<p>What makes the study notable is not merely that machine learning can classify text, which is now commonplace, but that it can do so under the constraints that actually govern criminal justice agencies: small labeled datasets, heterogeneous local vocabularies, tight budgets, and no supercomputers. A tool that runs on a laptop, learns a new jurisdiction&#8217;s dialect from a few dozen examples, and produces codes that researchers can audit and replicate changes the calculus of who can build evidence about crime and how quickly. If widely adopted, semi-supervised auto-coding could close the gap between the mountains of unstructured offense text generated daily and the standardized, linkable, analysis-ready data that modern criminology and evidence-based policy demand. The era in which every county&#8217;s crime data speaks its own private language may finally be drawing to a close.</p>
<div class="scienmag-article-metadata"><strong>Subject of Research:</strong> Automated natural language processing and semi-supervised machine learning for standardizing free-text offense descriptions into structured, comparable offense codes across criminal justice jurisdictions</p>
<p><strong>Article Title:</strong> Modernizing Criminal Justice Data: Developing a Semi-Supervised Tool to Code Multi-Jurisdictional Data</p>
<p><strong>Article References:</strong> DeMichele, M., Silver, I. A., Preiss, A. J., &amp; Baumgartner, P. (2026). Modernizing Criminal Justice Data: Developing a Semi-Supervised Tool to Code Multi-Jurisdictional Data. <em>American Journal of Criminal Justice</em>. <a href="https://doi.org/10.1007/s12103-026-09915-1" target="_blank" rel="noopener noreferrer">https://doi.org/10.1007/s12103-026-09915-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s12103-026-09915-1" target="_blank" rel="noopener noreferrer">10.1007/s12103-026-09915-1</a></p>
<p><strong>Keywords:</strong> natural language processing, offense auto-coding, text classification, data standardization, transformer models, few-shot learning, criminal justice data, pretrial data, record linkage, LoRA fine-tuning, semi-supervised learning, multi-jurisdictional harmonization</p>
</div>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">192142</post-id>	</item>
	</channel>
</rss>
