<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>direct text input for language models &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/direct-text-input-for-language-models/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 08 Oct 2026 18:59:11 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>direct text input for language models &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Byteification retrofits language models to read raw text directly</title>
		<link>https://scienmag.com/byteification-retrofits-language-models-to-read-raw-text-directly/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 08 Oct 2026 18:59:11 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[byte-level modelling]]></category>
		<category><![CDATA[byte-level neural networks]]></category>
		<category><![CDATA[byteification]]></category>
		<category><![CDATA[character-level language modeling]]></category>
		<category><![CDATA[character-level understanding]]></category>
		<category><![CDATA[computational efficiency]]></category>
		<category><![CDATA[direct text input for language models]]></category>
		<category><![CDATA[DNA and protein sequence modeling]]></category>
		<category><![CDATA[efficient training of language models]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[latent tokenizer]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[model distillation]]></category>
		<category><![CDATA[multilingual text processing]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[raw text processing in language models]]></category>
		<category><![CDATA[scientific data analysis with language models]]></category>
		<category><![CDATA[subword tokenization]]></category>
		<category><![CDATA[subword tokenization limitations]]></category>
		<category><![CDATA[task arithmetic]]></category>
		<category><![CDATA[text preprocessing in NLP]]></category>
		<category><![CDATA[tokenization]]></category>
		<category><![CDATA[tokenization bottleneck]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=248849</guid>

					<description><![CDATA[Researchers have unveiled byteification, a two-stage conversion method that retrofits existing subword-based large language models into byte-level models matching their performance at under one percent of a typical pretraining budget.]]></description>
										<content:encoded><![CDATA[<p>Every large language model in production today begins its work with a step so routine that it is often taken for granted: before a single parameter is touched, the incoming text is chopped into discrete units called tokens. For nearly all leading systems, those tokens are words or fragments of words, produced by a subword tokenizer trained in advance. A new study published in Nature argues that this seemingly innocuous preprocessing stage is a hidden bottleneck, and it introduces a technique called byteification that retrofits existing subword-based models into models that operate directly on the raw bytes of text, with a fraction of the training cost of building such systems from scratch.</p>
<p>The case against subword tokenization is technical but consequential. Because a token such as a whole word is treated as an atomic unit, information about the individual characters inside it is largely obscured. That matters little for casual conversation, but it becomes a serious handicap for scientific data, where meaning often lives at the character level: computer code, DNA and protein sequences, chemical formulae and multilingual text with unusual orthographies all depend on fine-grained structure that subword vocabularies flatten. Subword tokenization also introduces subtler pathologies, including a bias towards particular responses depending on how a prompt happens to be segmented, a fixed vocabulary that in practice skews models towards English, and a rigid allocation of compute that spends the same effort on every token regardless of difficulty.</p>
<p>Byte-level models, which treat the UTF-8 encoding of text as their basic units, promise to solve these problems in principle. Every character, in every language, decomposes into the same 256 bytes, so no vocabulary is needed and no character information is lost. Yet despite years of promising research, byte-level models have never been adopted by any leading laboratory. The researchers behind the new work identify a structural reason for this mismatch between theory and practice: byte-level models have almost always been trained from random initialization and compared against subword models also trained from scratch, while the state of the art in subword model training, spanning data curation, architecture and post-training, evolves so quickly that byte-level efforts cannot keep pace.</p>
<p>Their solution is to stop competing on that uneven playing field and instead convert the winners. The team, led by Benjamin Minixhofer of the Allen Institute for AI and the University of Cambridge, together with colleagues at the University of Washington, Imperial College London and LMU Munich, calls the procedure byteification. Starting from fully open source models, including Olmo 3 7B, OLMo 2 1B, Qwen3 8B Base and Llama 3 8B, they produced byte-level counterparts named Bolmo 7B, Bolmo 1B, Bwen 8B and Blama 8B. The total conversion budget was 49.1 billion tokens, less than one percent of a typical pretraining run, yet the resulting models approach the performance of their subword parents and decisively outperform earlier byte-level systems of comparable size.</p>
<p>The architecture underlying the converted models belongs to a family the authors call latent tokenizer language models. A shallow local encoder first builds contextualized representations of each byte. A boundary predictor then decides where to place patch boundaries, grouping bytes into patches. These patch representations are pooled and passed through the deep global transformer, which does most of the computational work, before being depooled back to the byte level by a local decoder that predicts the next byte. Crucially, the global model is inherited intact from the source subword model, so the expensive knowledge embedded in its weights is preserved rather than relearned.</p>
<p>One of the study&#8217;s most elegant insights concerns a subtle asymmetry in how boundaries are placed. Subword tokenizers are not causal: they consult future characters when deciding where a token ends. A tokenizer might split the flowerbed after flower in one sentence but keep flowerb together in another, even though the text up to that point is identical, because what comes next differs. Earlier byte-level architectures, which must decide boundaries using only past context during generation, cannot replicate this behaviour, creating an expressivity gap. The team resolved it with a two-part scheme: during prefilling, their boundary predictor peeks one byte into the future, which proved sufficient to match subword behaviour, while during decoding a special boundary symbol fused into an expanded 512-entry byte vocabulary lets the model emit boundary decisions alongside each byte at essentially no extra cost.</p>
<p>The conversion itself proceeds in two stages. In the first, the global model is frozen while the new local components are trained to exactly recover the behaviour of the source model, using distillation losses that match pooled byte representations to subword embeddings and patch likelihoods to subword token likelihoods. The boundary predictor alone reaches over 99 percent accuracy in emulating the original tokenizer. In the second stage, the entire model is trained end to end so it can exploit genuinely byte-level information. Ablations showed that skipping the first stage hurts, particularly for smaller models, but also that the first stage serves a second purpose: it provides a fast feedback loop for testing whether a candidate architecture has enough capacity before committing to full training.</p>
<p>The performance results are striking. Bolmo 7B achieved a 16.5 percent absolute improvement in STEM tasks over BLT 7B, a leading byte-level model trained from scratch, and substantially outperformed its own subword parent on character-understanding benchmarks such as CUTE and the multilingual EXECUTE suite, aided by a small injection of synthetic character-manipulation training data. Bwen 8B, built on Qwen3, statistically significantly outperformed all earlier public byte-level models and came close to, and in some categories surpassed, its source model. The byteified models also showed higher pass@16 rates on code generation at lower pass@1, suggesting they produce more diverse candidate solutions than their subword counterparts, an observation the authors flag as requiring further study.</p>
<p>Two further results point toward practical advantages beyond raw accuracy. Because byte-level models are not locked into a fixed vocabulary, the researchers could raise the average number of bytes per patch to trade a smooth, controlled amount of performance for inference speed, avoiding the softmax bottleneck that eventually causes subword models transferred to larger SuperBPE vocabularies to become Pareto-dominated. And because byteification preserves the ecosystem around the source model, the team could transfer instruction-following abilities from a post-trained Olmo 3 checkpoint into Bolmo via task arithmetic, simply adding the weight difference between the post-trained and base transformer layers, with no additional training at all, lifting Bolmo to parity with the original post-trained model on the IFEval benchmark.</p>
<p>The authors position byteification not as a replacement for training byte-level models from scratch but as a complementary, cheap pathway that removes a long-standing performance barrier. If any state-of-the-art subword model can be converted to the byte level for a few percent of its original training cost, high-performing byte-level architectures can be identified quickly, and the English-centric biases, character-blindness and rigid compute allocation of subword tokenization may finally become optional rather than inevitable. For domains where meaning lives in the characters, from genomics to software engineering, the end-to-end byte-level future that tokenization researchers have long promised may have just moved considerably closer.</p>
<p><strong>Subject of Research:</strong> Converting subword-based large language models into byte-level models through a retrofitting method called byteification</p>
<p><strong>Article Title:</strong> Retrofitting language models to operate over bytes</p>
<p><strong>Article References:</strong> Minixhofer, B., Murray, T., Limisiewicz, T., Korhonen, A., Zettlemoyer, L., Smith, N. A., Ponti, E. M., Soldaini, L., &amp; Hofmann, V. (2026). Retrofitting language models to operate over bytes. <em>Nature</em>. <a href="https://doi.org/10.1038/s41586-026-11111-4" rel="noopener noreferrer">https://doi.org/10.1038/s41586-026-11111-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s41586-026-11111-4" rel="noopener noreferrer">10.1038/s41586-026-11111-4</a></p>
<p><strong>Keywords:</strong> large language models, byteification, tokenization, byte-level modelling, subword tokenization, latent tokenizer, machine learning, natural language processing, character-level understanding, model distillation, task arithmetic, computational efficiency</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">248849</post-id>	</item>
	</channel>
</rss>
