Physicists at the Jožef Stefan Institute in Ljubljana have borrowed the machinery behind ChatGPT to tackle one of the most stubborn computing bottlenecks in particle physics: simulating how charged particles leave their fingerprints in silicon tracking detectors. In a study published in The European Physical Journal C, Tadej Novak and Borut Paul Kerševan demonstrate the first application of generative machine learning to silicon tracker simulation, using a GPT-like transformer architecture that treats the trail of detector hits left by a particle as if it were a sentence in an unfamiliar language. The result is a neural network that can generate fully correlated sequences of detector hits with tracking performance comparable to the gold-standard Geant4 simulation, at least for the well-behaved muons that served as the benchmark.
The motivation is stark. The Large Hadron Collider’s physics programme depends on Monte Carlo simulation, in which the fates of billions of proton collisions are computed particle by particle, interaction by interaction. The workflow has four stages: generating the physics events, propagating particles through a virtual detector, digitising the deposited energy into signals that mimic real detector readout, and reconstructing tracks with the same algorithms used on actual data. The detector response step, traditionally handled by the Geant4 toolkit, is the most computationally punishing, and the problem is about to get dramatically worse. By the end of the High-Luminosity LHC programme, the ATLAS and CMS experiments will have collected up to ten times more data than in their first three runs combined, and many precision analyses will find their sensitivity limited not by the data themselves but by the statistical size of the simulated samples needed to model backgrounds.
Generative machine learning has already made inroads on the calorimeter problem, where the goal is to reproduce the showers of energy deposited when particles slam into dense material. Generative adversarial networks, variational autoencoders, normalising flows, diffusion models and continuous flow models have all been applied, and the ATLAS experiment already uses GANs in production for fast photon shower simulation. But tracking detectors present a fundamentally different challenge. A calorimeter response can be interpreted as an image with correlated entries. A tracking detector, by contrast, records a sparse sequence of discrete interactions: a particle crossing thin layers of silicon widely separated in space. Rendering that as an image would produce a picture that is almost entirely empty, which is precisely why the Slovenian team turned to sequence-based models instead.
The conceptual leap is elegant. In a large language model, text is broken into tokens and the network learns to predict the next token given everything that came before. Novak and Kerševan applied the same logic to particle tracks. Each hit in a silicon sensor is described by seven features: a discrete particle identifier encoding type and charge, a geometry identifier for the detector module, two local coordinates on the sensor surface, and three components of the particle momentum after the hit. Every feature is tokenised into a shared dictionary of integers, with continuous quantities rounded to two decimal places. A track becomes a flat sequence of tokens, bookended by virtual start and end hits representing the particle’s birth at the beamspot and its exit from the tracker. The transformer’s job is then exactly what GPT does with words: predict the next element of the sequence, one feature at a time.
The technical implementation relies on a decoder-only transformer built on Karpathy’s nanoGPT codebase, with eight layers and eight attention heads each, feed-forward dimensions four times the input dimension, and dropout of 0.2. Model sizes range from 9.1 to 35 million parameters depending on the token dictionary. Because self-attention scales quadratically with sequence length, the authors employed a sliding-window training scheme, chopping full tracks into rolling windows of four hits. This is physically justified since correlations between hits weaken with distance and the dominant feature, local track curvature, is captured by nearby hits. The trick also augments the training data by an order of magnitude, letting the graphics accelerators work with much larger batches. Training used the AdamW optimiser with gradient clipping and a carefully scheduled learning rate, running for roughly 5000 epochs on Slovenian supercomputing hardware.
Validation proceeded at two levels. At the hit level, the researchers compared distributions of individual hit properties between Geant4 and the neural network, always against the rounded Geant4 data on which the model was trained. For single muons, which barely interact with the detector beyond multiple scattering, the agreement is impressive. The number of hits per track is well modelled, with the largest discrepancy for any muon model being just six hits, and directly tokenised quantities such as the momentum component px agree at the percent level. Derived global coordinates fluctuate up to about five percent, with deviations growing in the distribution tails. At the higher level, the simulated hits were fed into the ACTS tracking software and reconstructed using the standard Open Data Detector configuration, allowing a direct comparison of seeding and tracking efficiencies.
Those reconstruction results are the heart of the paper. The technical seeding efficiency measures how often the algorithm finds viable track seeds, while the technical tracking efficiency measures how often full tracks are successfully fitted. A benchmark 11.2-million-parameter model reaches 94.9 percent tracking efficiency with only a slight seeding efficiency drop, comparable to the rounded Geant4 reference. Scaling the model’s layer dimensions by a factor of two improves both metrics further, though they remain below the reference sample. One persistent weakness emerged: the azimuthal angle phi is poorly modelled when the full detector coverage is included, and a dedicated test with a wider phase space showed that the model learns probabilities less accurately as the training domain grows. The authors conclude that detector-aware rounding and tokenisation strategies will be essential for production-quality models.
Electrons and pions exposed the limits of the approach. Electrons undergo bremsstrahlung and pair production, processes that can abruptly change their direction and momentum, and the model’s token dictionary grew 1.4 times larger to accommodate the added variety. Although overall track quality remained comparable to muons, the momentum modelling proved insufficient, with a sizeable fraction of generated electrons carrying too much transverse momentum; the network learns the shape of the momentum tail but is biased toward smaller changes. Pions, which can decay inside the tracker, revealed a related failure: rare, low-probability processes are systematically under-learned, so machine-generated pions almost never decay, artificially inflating the tracking efficiency even though the tracks themselves remain good quality. Reweighting such rare events could help, but the authors note it would erode the statistical power that makes fast simulation worthwhile in the first place.
Computing performance adds a further caveat. GPU inference runs at roughly the same speed as Geant4 simulation on same-generation CPUs, scaling linearly with model size, so better physics currently means slower generation. CPU inference is orders of magnitude too slow to replace conventional simulation, and given the much higher cost of high-performance GPUs, the authors judge the present speed insufficient for general production use. Still, benchmarks across Nvidia hardware generations show clear gains, and half-precision bf16 computation speeds up training by about a third with no loss of physics performance. The team has released its SiliconAI training and validation code and datasets openly, and sees paths forward in extending the sequence representation to secondary particles and in exploring alternative sequence-based architectures. The message is provocative: the same technology that writes essays can, with careful tokenisation, learn to write the passage of a muon through a silicon tracker, and with further work it may help the LHC’s next generation of experiments simulate their way to new discoveries.
Subject of Research: Generative transformer models for fast simulation of silicon tracking detectors in high energy physics
Article Title: GPT-like transformer model for silicon tracking detector simulation
Article References: Novak, T., & Kerševan, B. P. (2026). GPT-like transformer model for silicon tracking detector simulation. The European Physical Journal C, 86(9), Article 1090. https://doi.org/10.1140/epjc/s10052-026-16362-z
Image Credits: AI Generated
DOI: 10.1140/epjc/s10052-026-16362-z
Keywords: transformer, GPT, silicon tracker, detector simulation, Geant4, machine learning, LHC, High-Luminosity LHC, particle physics, generative models, Open Data Detector, computing
Cite Scienmag News
Katie Riggs. (October 6, 2026). GPT-Style AI Learns to Simulate Particle Tracks in Silicon Detectors. Scienmag. https://scienmag.com/gpt-style-ai-learns-to-simulate-particle-tracks-in-silicon-detectors/
Katie Riggs. "GPT-Style AI Learns to Simulate Particle Tracks in Silicon Detectors." Scienmag, 6 October 2026, https://scienmag.com/gpt-style-ai-learns-to-simulate-particle-tracks-in-silicon-detectors/. Accessed 6 October 2026.
Katie Riggs. "GPT-Style AI Learns to Simulate Particle Tracks in Silicon Detectors." Scienmag. October 6, 2026. https://scienmag.com/gpt-style-ai-learns-to-simulate-particle-tracks-in-silicon-detectors/

