<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>large &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/large/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Tue, 22 Sep 2026 22:30:43 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>large &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Reads Imperial Archives: Language Models Reconstruct Qing Dynasty Craft Records</title>
		<link>https://scienmag.com/ai-reads-imperial-archives-language-models-reconstruct-qing-dynasty-craft-records/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 22:30:43 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[AI-driven historical knowledge extraction]]></category>
		<category><![CDATA[Chinese imperial workshop documentation]]></category>
		<category><![CDATA[classical Chinese archives]]></category>
		<category><![CDATA[classical Chinese text mining]]></category>
		<category><![CDATA[computational reconstruction of Qing craft records]]></category>
		<category><![CDATA[cultural heritage computing]]></category>
		<category><![CDATA[digital humanities]]></category>
		<category><![CDATA[digitization of Chinese historical archives]]></category>
		<category><![CDATA[entity normalization]]></category>
		<category><![CDATA[historical documentation of craftsmanship]]></category>
		<category><![CDATA[imperial archives analysis]]></category>
		<category><![CDATA[imperial artifact production]]></category>
		<category><![CDATA[imperial workshop supervision and logistics]]></category>
		<category><![CDATA[information extraction]]></category>
		<category><![CDATA[knowledge graph]]></category>
		<category><![CDATA[knowledge graph construction from historical records]]></category>
		<category><![CDATA[large]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models in digital humanities]]></category>
		<category><![CDATA[Qing court artifact commissioning]]></category>
		<category><![CDATA[Qing dynasty]]></category>
		<category><![CDATA[Qing dynasty craft production]]></category>
		<category><![CDATA[Yongzheng reign]]></category>
		<category><![CDATA[Zaobanchu]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=208347</guid>

					<description><![CDATA[Researchers used large language models to extract entities and relations from Qing dynasty workshop records and build a knowledge graph of Yongzheng-era imperial artifact production.]]></description>
										<content:encoded><![CDATA[<p>A new study published in Humanities and Social Sciences Communications describes how large language models were used to build a knowledge graph of imperial artifact production during the Yongzheng reign of the Qing dynasty, offering digital humanities researchers a fresh computational window onto one of the most intensively documented workshops in Chinese history. The work addresses a longstanding bottleneck in historical research on craft production: the vast, unstructured Chinese-language archives that record how the imperial court commissioned, supervised, and delivered objects ranging from porcelain and lacquerware to metalwork and textiles.</p>
<p>The imperial workshops of the Qing court, administered through bodies such as the Zaobanchu, generated enormous quantities of routine documentation. Memorials, requisitions, work logs, and delivery records capture the names of craftsmen, the materials they used, the deadlines imposed by the palace, and the judgments of supervising officials. For historians, this material is extraordinarily rich but also extraordinarily difficult to mine. The documents were written in classical Chinese with period-specific terminology, honorific conventions, and abbreviations, and the sheer volume of records makes manual extraction of structured facts a decades-long endeavor for even a well-funded team of archivists.</p>
<p>Knowledge graphs have become a standard tool for organizing such information. A knowledge graph represents entities, for example a craftsman, an object type, a workshop department, or a date, as nodes, and expresses relationships between them as typed edges, such as produced, supervised, requested, or made of. Once built, a graph allows researchers to ask structured questions: which craftsmen worked on both enamel and lacquer in a given year, which materials were most frequently requested for objects destined for particular palace halls, or how production patterns shifted after the death of one supervisor and the appointment of another. The difficulty has always been populating the graph, a task that traditionally required hand-coding rules or annotating large volumes of text by hand.</p>
<p>The researchers turned to large language models to automate this extraction. Modern LLMs, trained on enormous corpora of text, have demonstrated a strong ability to perform information extraction tasks with minimal task-specific training, a property that makes them attractive for low-resource domains such as classical Chinese archival documents where annotated training data is scarce. Rather than training a bespoke model from scratch, the approach described in the study leverages prompting strategies that instruct a general-purpose language model to identify entities and relations in archival passages and to normalize them into a consistent schema designed for the domain of imperial craft production.</p>
<p>Building a reliable pipeline required more than prompting alone. The authors confronted classic problems of LLM-based extraction: hallucinated entities, inconsistent naming of the same person or object across documents, ambiguity in classical Chinese phrasing, and the need to map colloquial or archaic terms onto a controlled vocabulary. Their framework therefore combines the generative strengths of language models with conventional knowledge engineering practices, including schema design, entity normalization and deduplication, and validation steps that check extracted records for internal consistency before they are committed to the graph. The result is a hybrid workflow in which the language model does the heavy lifting of reading unstructured text, while structured constraints keep the output tidy enough for scholarly use.</p>
<p>The choice of the Yongzheng reign as the target domain is historically motivated. Yongzheng, who reigned from 1722 to 1735, is known for close personal supervision of imperial manufacturing and for the high quality and distinctive aesthetic of objects produced under his direction. The workshop records from his reign are comparatively well preserved and internally consistent, which makes them an excellent testbed for computational methods: the density of documented production events provides ample material for extraction, while the known historical context offers a way to sanity-check the resulting graph against what historians already know about the period.</p>
<p>The resulting knowledge graph, once assembled, functions as a navigable map of the imperial production system. Nodes capture not only individual artifacts and their materials but also the institutional geography of the workshops, the networks of craftsmen, and the temporal rhythm of commissions. Such a resource supports quantitative analysis that would be impractical with the raw documents alone. Researchers can trace the lifecycle of an object from initial imperial request through design approval, material allocation, fabrication, revision, and final delivery, and can aggregate thousands of such lifecycles to reveal systemic patterns, for instance seasonal fluctuations in production or the concentration of particular techniques within particular workshop divisions.</p>
<p>Beyond Qing studies, the significance of the work lies in its demonstration of a generalizable method for the digital humanities. Many of the world&#8217;s historical archives remain locked in unstructured text, and the combination of large language models with knowledge graph construction offers a template that can be adapted to other corpora, languages, and domains. The study also contributes to an ongoing conversation about rigor in LLM-assisted scholarship. Because generative models can produce plausible but incorrect extractions, the authors emphasize evaluation of extraction quality and the importance of keeping human experts in the loop, particularly for contested or ambiguous readings of historical documents. The graph is best understood not as a replacement for close reading but as an index and hypothesis generator that directs expert attention to the most promising passages.</p>
<p>Limitations acknowledged in work of this kind are worth noting. Extraction accuracy for classical Chinese remains imperfect, entity resolution across hundreds of thousands of records is an open problem, and any schema imposes interpretive choices that shape downstream analysis. The authors position the current graph as a foundation to be extended and refined, with future directions including richer temporal modeling, integration with museum collection databases and surviving physical artifacts, and multimodal linkage between archival text and images of the objects themselves. As large language models continue to improve in multilingual and historical-language capability, pipelines of this type are likely to become a standard part of the historian&#8217;s toolkit.</p>
<p>For readers of science news, the study is a vivid example of artificial intelligence moving beyond chatbots and code generation into the archive. A language model trained largely on modern text was adapted to read eighteenth-century palace paperwork and return a structured, queryable model of an entire imperial manufacturing economy. In doing so, it turns stacks of memorials and requisition slips into something closer to a database of the Yongzheng court&#8217;s creative life, allowing scholars to see the craftsmen, materials, and decisions behind the treasures that now anchor museum collections worldwide, and pointing toward a future in which the deep past becomes searchable at scale.</p>
<p><strong>Subject of Research:</strong> Large language model-driven knowledge graph construction for Qing Yongzheng imperial artifact production</p>
<p><strong>Article Title:</strong> Large Language Model-driven knowledge graph construction for Qing Yongzheng imperial artifact production</p>
<p><strong>Article References:</strong> Large Language Model-driven knowledge graph construction for Qing Yongzheng imperial artifact production. (n.d.). <a href="https://doi.org/10.1038/s41599-026-09040-8" rel="noopener noreferrer">https://doi.org/10.1038/s41599-026-09040-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s41599-026-09040-8" rel="noopener noreferrer">10.1038/s41599-026-09040-8</a></p>
<p><strong>Keywords:</strong> large language models, knowledge graph, Qing dynasty, Yongzheng reign, imperial artifact production, digital humanities, information extraction, classical Chinese archives, Zaobanchu, entity normalization, cultural heritage computing, Large</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">208347</post-id>	</item>
		<item>
		<title>How Data Centers Are Reengineering AI Training for Larger Language Models</title>
		<link>https://scienmag.com/how-data-centers-are-reengineering-ai-training-for-larger-language-models/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sat, 29 Aug 2026 03:12:12 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[AI accelerators]]></category>
		<category><![CDATA[AI training hardware optimization]]></category>
		<category><![CDATA[AI training system reliability and recovery]]></category>
		<category><![CDATA[data center engineering for AI]]></category>
		<category><![CDATA[distributed training]]></category>
		<category><![CDATA[efficient]]></category>
		<category><![CDATA[efficient GPU utilization in deep learning]]></category>
		<category><![CDATA[fault tolerance]]></category>
		<category><![CDATA[fault-tolerant AI training systems]]></category>
		<category><![CDATA[GPU clusters]]></category>
		<category><![CDATA[high-performance computing in AI]]></category>
		<category><![CDATA[high-performance networking]]></category>
		<category><![CDATA[high-speed interconnects for AI clusters]]></category>
		<category><![CDATA[large]]></category>
		<category><![CDATA[large language model training infrastructure]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large-scale neural network training challenges]]></category>
		<category><![CDATA[memory optimization]]></category>
		<category><![CDATA[mixed precision]]></category>
		<category><![CDATA[model training efficiency metrics]]></category>
		<category><![CDATA[optimization of AI compute and storage resources]]></category>
		<category><![CDATA[parallelism]]></category>
		<category><![CDATA[scalable distributed AI training systems]]></category>
		<category><![CDATA[training]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=184390</guid>

					<description><![CDATA[A comprehensive survey explains how accelerators, networks, storage, parallelism and fault tolerance are being redesigned to train increasingly large language models efficiently.]]></description>
										<content:encoded><![CDATA[<p>Training a modern large language model is no longer simply a matter of assigning a neural network to a powerful computer. It is an extended, tightly coordinated operation involving thousands of accelerators, high-speed links, distributed storage, scheduling software and recovery systems. A survey by Jiangfei Duan, Shuo Zhang and colleagues maps the engineering behind that operation, showing how researchers are redesigning nearly every layer of the computing stack to keep models moving forward. The review focuses on three connected demands: scalability, efficiency and reliability. Together, these demands define whether a distributed training system can expand to very large clusters, use its hardware productively and survive the failures that become increasingly likely during jobs lasting weeks or months.</p>
<p>The scale of the challenge is illustrated by the training of LLaMA-3, which the survey reports took about 54 days using 16,000 H100-80GB GPUs on Meta’s production cluster. Even at that scale, the system reached a Model FLOPs Utilization, or MFU, of only about 38 to 41 percent. MFU measures how effectively available floating-point computing capacity is used during training; the unused portion can reflect communication delays, memory bottlenecks, synchronization, uneven workloads or inefficient kernels. The figures make clear why simply adding more processors does not automatically produce proportional gains. Every accelerator must receive data, exchange intermediate results or gradients and remain synchronized with its peers. A slow link, overloaded storage service or failed device can leave many other processors waiting.</p>
<p>Most of the models covered by the review use decoder-only Transformer architectures. Text is converted into tokens and then vectors, supplied with positional information and processed through repeated layers containing attention and feed-forward blocks. Attention calculates relationships between tokens by transforming them into query, key and value tensors and applying a weighted operation based on their similarities. In its conventional form, self-attention has computational and memory demands that grow quadratically with sequence length, making long contexts especially expensive. Architectural changes such as multi-query attention, grouped-query attention and multi-latent attention reduce some key-value storage pressures, while Mixture-of-Experts designs activate only a subset of feed-forward experts for each input. These model-level choices directly affect the hardware, communication patterns and memory strategies required for training.</p>
<p>The survey describes distributed training as a combination of complementary forms of parallelism. Data parallelism assigns different portions of a batch to different devices, then uses collective operations to aggregate gradients. Tensor parallelism divides the matrices inside individual layers and is generally best suited to tightly connected accelerators because intermediate activations must be exchanged frequently. Pipeline parallelism places consecutive groups of layers on different devices and passes activations between stages, making it useful across nodes with lower bandwidth but introducing idle “pipeline bubbles” while stages fill and drain. Sequence parallelism divides long token sequences among devices, reducing per-device activation and attention costs. In practice, these methods are combined into hybrid plans, often called three-dimensional parallelism when data, tensor and pipeline parallelism are used together.</p>
<p>Choosing that combination is a difficult systems problem because every benefit carries a cost. Fully replicated data parallelism is comparatively simple but stores duplicate model states on every device. Fully sharded approaches such as ZeRO-3 and Fully Sharded Data Parallelism distribute parameters, gradients and optimizer states across devices, sharply reducing memory use while increasing communication. Hybrid sharding provides a middle ground by replicating states within smaller groups and sharding them across those groups. Pipeline schedules attempt to reduce idle time by interleaving forward and backward computations, but more active micro-batches can create memory imbalance between stages. Automated parallelism systems address the complexity by searching possible partitions, estimating execution and communication costs and selecting a plan for a particular model and cluster. The review highlights approaches based on dynamic programming, simulators, reinforcement learning, Monte Carlo search and constraint-guided planning.</p>
<p>Hardware and networks form the physical foundation of these strategies. GPUs are well matched to the matrix and vector operations that dominate Transformer training and increasingly support mixed numerical formats, including FP16, BF16 and FP8. Their high-bandwidth memory and specialized tensor-processing units allow large matrix operations to run in parallel, while interconnects such as NVLink and NVSwitch provide faster communication than conventional PCI Express within a server. Other accelerator ecosystems, including AMD GPUs, Google TPUs, Intel Gaudi processors, Graphcore IPUs and Cerebras wafer-scale systems, offer different combinations of memory, compute capacity and software support. The survey emphasizes that software portability remains important: a strategy optimized for one accelerator and programming environment may require substantial adaptation to run efficiently on another.</p>
<p>Communication can become the dominant expense in a distributed job. Gradient synchronization creates periodic bursts of what network engineers call elephant flows, while tensor and expert parallelism can generate intensive all-to-all exchanges. Remote Direct Memory Access allows one machine to access memory on another without involving the operating system, and GPU-direct variants can move data between GPUs across nodes while bypassing the CPU. InfiniBand provides a dedicated high-performance fabric, whereas RDMA over Converged Ethernet brings similar capabilities to Ethernet-based data centers. Network layouts are increasingly designed around training traffic rather than general-purpose workloads. Rail-optimized architectures group corresponding GPUs through carefully arranged switches, while other designs use sparse connectivity, optical circuit switching or topology-aware routing. Load-balancing methods split large transfers across multiple paths, and congestion-control schemes regulate traffic to limit stalls and packet loss.</p>
<p>Storage must also be treated as part of the training engine. A 70-billion-parameter model can produce a checkpoint of about 980 gigabytes, according to the survey, and thousands of accelerators may need to save or reload such states together. Distributed file systems and object stores therefore need high write and read bandwidth as well as consistency, availability, protection and security. Training data presents a different challenge: LLaMA 3 was trained on more than 15 trillion tokens, while crawling, filtering and preparing web-scale data can involve volumes far larger than the final dataset. Caches such as Alluxio and JuiceFS can prefetch data from slower storage, keeping accelerators supplied even when each token is normally read only once. Cluster schedulers must coordinate these demands with GPU, CPU, memory and network resources, while balancing fairness, utilization, job packing, adaptive scaling and energy consumption across multiple users and workloads.</p>
<p>Memory optimization is central because training stores far more than the model’s visible parameters. Mixed-precision training requires parameters and gradients, higher-precision copies for stable optimizer updates, momentum and variance states, activations from the forward pass, temporary communication buffers and allocator space. For a model with Phi parameters, the survey estimates that model states alone can require 16Phi bytes under a common mixed-precision arrangement before activations and other buffers are counted. Activation recomputation reduces the peak by discarding selected intermediate tensors during the forward pass and rebuilding them during backpropagation, trading memory for additional computation. Sharding removes duplicate states, while CPU and NVMe offloading supplements limited GPU memory. Defragmentation techniques use improved allocation policies or virtual-memory stitching to turn scattered free regions into usable capacity. The goal is not merely to fit a model, but to prevent memory management from idling the entire cluster.</p>
<p>At the computation level, the review points to a shift from treating training as a sequence of generic operators toward designing complete dataflows for specific hardware. FlashAttention reduces high-bandwidth-memory traffic by processing attention in tiles, retaining intermediate results in faster on-chip memory and fusing matrix multiplication, softmax and related operations into efficient kernels without changing the exact attention result. Later variants improve parallel scheduling, asynchronous data movement and low-precision execution. Compilers such as TVM, Triton and TorchInductor can generate kernels, fuse operators and exploit memory locality across broader graphs, reducing the need for every optimization to be hand-written. Mixed precision extends the same principle: BF16 and FP16 are widely used, while FP8, INT8, INT4 and even one-bit or ternary representations are being investigated. These formats can reduce computation, storage and communication, but they require scaling methods, higher-precision accumulators or latent weights to prevent numerical errors from undermining training.</p>
<p>The reliability problem grows with both cluster size and training duration. A failure may come from a GPU, a network link, storage, software or a less obvious straggler that slows progress without producing an immediate crash. Because distributed training is often synchronous, a single stalled participant can leave thousands of others idle. Checkpointing provides a recovery point, but writing and restoring enormous states can itself interrupt useful work. The survey therefore treats anomaly detection, rapid fault diagnosis, resilient communication and automatic recovery as core components rather than afterthoughts. It also distinguishes the needs of pretraining from those of later alignment and reinforcement learning, where actor, critic, reward and reference models may alternate between inference and optimization. Newer asynchronous systems stream generated trajectories to trainers instead of waiting for the slowest rollout, while agentic training adds tool calls, sandboxes, verifiers and long-horizon interactions. The review’s broad conclusion is that future progress will depend on co-design: accelerators, memory, networks, storage, schedulers, compilers and learning algorithms must be optimized as one evolving system, with optical computing and optical networks among the possible technologies for overcoming the limits of conventional digital hardware.</p>
<p><strong>Subject of Research:</strong> Distributed infrastructure for efficient large language model training</p>
<p><strong>Article Title:</strong> Efficient training of large language models on distributed infrastructures: a survey</p>
<p><strong>Article References:</strong> Duan, J., Zhang, S., Wang, Z., Jiang, L., Qu, W., Hu, Q., Wang, G., Weng, Q., Yan, H., Zhang, X., Qiu, X., Lin, D., Wen, Y., Jin, X., Zhang, T., &amp; Sun, P. (2026). Efficient training of large language models on distributed infrastructures: a survey. <em>Vicinagearth, 3</em>(1), Article 9. <a href="https://doi.org/10.1007/s44336-026-00038-z" rel="noopener noreferrer">https://doi.org/10.1007/s44336-026-00038-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44336-026-00038-z" rel="noopener noreferrer">10.1007/s44336-026-00038-z</a></p>
<p><strong>Keywords:</strong> large language models, distributed training, GPU clusters, parallelism, AI accelerators, high-performance networking, memory optimization, mixed precision, fault tolerance, Efficient, training, large</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">184390</post-id>	</item>
	</channel>
</rss>
