Sometime in the mid-2030s, three spacecraft flying in formation around the Sun will start eavesdropping on the gravitational ripples of spacetime itself. The Laser Interferometer Space Antenna, or LISA, will tune in to millihertz gravitational waves — a slice of the cosmic spectrum entirely beyond the reach of ground-based detectors — and among its most anticipated prizes are pairs of black holes caught in the earliest, slowest act of their merger. These binaries can chirp gently for years or even decades before they coalesce, and a single decade-long signal, sampled at its Nyquist rate, may span hundreds of millions of data points, approaching a billion in the most extreme cases. That torrent is a gift for astrophysics and a catastrophe for computation, because extracting a source’s properties means evaluating a likelihood function millions of times across parameter space. A new study published in the journal General Relativity and Gravitation now shows that, for simulated data, nearly the entire torrent can be thrown away — and the correct answer still comes back, up to a million times faster.
LISA, a mission led by the European Space Agency with major NASA participation, will fly a triangular constellation of spacecraft separated by 2.5 million kilometres, timing laser beams between them to sense spacetime strains far smaller than an atomic nucleus. Its millihertz band, roughly 1 to 100 millihertz, hosts some of the most extreme physics in the universe: mergers of supermassive black holes, stellar remnants spiralling into giant ones, and the long, slow prelude of ordinary binary black holes, which may spend millions of years emitting a gradually strengthening chirp before coalescence. The spacecraft are expected to operate for five to ten years, and for the most sluggishly evolving sources — compact binaries whose signal hovers near the upper edge of the sensitivity band throughout the mission — the strain record, sampled at its Nyquist rate, could contain between 100 million and a billion data points. Extracting masses, spins and distances from such a signal is a job for Bayesian parameter estimation, in which a stochastic sampler wanders through parameter space, evaluating the likelihood over and over. With datasets of that size, even one evaluation is a burden, and millions of them become prohibitive.
The bottleneck has a precise mathematical shape. In the Gaussian-noise framework that underpins gravitational-wave data analysis, the likelihood of a proposed set of parameters is a noise-weighted inner product between the data and the template waveform, evaluated against the inverse of the noise covariance matrix. In the frequency domain, where much of gravitational-wave analysis lives, this is cheap, because stationary noise makes frequency samples statistically independent and fast Fourier transforms do the heavy lifting. Yet a significant share of the physics LISA is expected to probe is most naturally written in the time domain: environmental effects that gradually reshape an inspiral, small departures from general relativity carried by additional dynamical fields, and the light-travel-time modulations that arise when a binary orbits a more massive third object. In the time domain, coloured detector noise correlates neighbouring samples, and a naive likelihood evaluation must grapple with the full covariance structure across the entire enormous record. Multiply that by the millions of evaluations a sampler demands, and the calculation collapses under its own weight.
The remedy proposed by Jethro Linley of the University of Glasgow is disarmingly simple to state: throw away almost everything. The scheme, described in the journal General Relativity and Gravitation, retains only a tiny subset of the data — typically between one thousand and ten thousand samples out of potentially hundreds of millions — and defines a modified noise-weighted inner product on that subset which closely reproduces the inner product of the full dataset across the manifold of plausible waveforms. The subtlety is that time-domain samples from LISA-like noise are statistically entangled with their neighbours, and simply discarding them distorts the resulting posterior. The method therefore whitens the residual first, transforming it with the inverse square root of the noise covariance so that retained samples become statistically independent, and computes only the small neighbourhood of raw samples each retained point requires. Once the noise power spectrum is trimmed to the signal’s frequency band, that neighbourhood shrinks to single digits, so the cost of each likelihood evaluation scales linearly with the number of retained samples.
Discarding data on its own would breed dangerous overconfidence: fewer points mean less information, and the posterior would shrink into an implausibly tight ball around the true parameters. The scheme compensates by rescaling the noise weighting of the retained samples with a single correction factor, chosen so that the curvature of the log-likelihood around its peak — encoded in the Fisher information matrix, which measures how sharply each parameter is constrained — matches that of the full-data analysis. Linley derives this factor by minimising the Jeffreys divergence, a symmetric measure of the distance between the Gaussian approximations of the downsampled and full posteriors. For the most stubborn cases, the paper goes further still and constructs an exact scheme that preserves the entire Fisher information matrix by assigning an individually tailored weight to every retained sample, a construction the paper proves can encode the data without any loss of Fisher information at all.
The method does demand one property of its targets: the signal must be slowly evolving, in a precise sense, meaning that its Fisher information content, averaged over windows of roughly a hundred sampling intervals, stays nearly constant across the record. A waveform with an abrupt step defeats the approach outright, because the samples flanking the step carry information found nowhere else. Fortunately, the gently chirping inspirals that dominate LISA’s slow-moving catalogue fall squarely within the safe class, and within that class the choice of which samples to keep turns out to matter less than one might fear. Uniform decimation courts aliasing yet performs surprisingly well, because correlated detector noise encodes high-frequency information in the surviving residuals; random selection suppresses aliasing altogether; and a hybrid of the two, blending roughly equal parts of both, emerged as the best all-round performer in head-to-head comparisons, though cluster sampling lagged until the Fisher-preserving variant gave it a second life.
Proving the approximation trustworthy demanded its own statistical machinery, because the true full-data posterior is precisely the thing too expensive to compute. The study therefore bootstraps its own benchmark. For each test system, 21 independent posteriors were generated with a nested sampler running through Bilby, the standard Bayesian inference library of gravitational-wave astronomy, using simulated inspiral signals with eight free parameters — chirp mass, mass ratio, effective spin, distance, inclination, polarisation and coalescence time, with the coalescence phase marginalised out numerically — injected at a signal-to-noise ratio of eight, the conventional threshold for a detection. Agreement between posteriors was quantified with a combined marginal version of the Jensen–Shannon divergence, weighted by the entropy of each parameter’s marginal distribution, across hundreds of pairwise comparisons per configuration. A downsampled posterior earned acceptance only when its divergence from a reference fell to the level set by a completely different source of variation: the natural spread in posterior geometry caused purely by swapping one realisation of detector noise for another, the irreducible variation that no analysis of real data could ever escape.
The empirical results carry a twist. In most systems, posteriors built from a few hundred retained samples proved statistically indistinguishable from those computed on the full record; the worst-behaved system required only 362 samples, and some configurations converged with as few as 16. A handful of rebellious systems initially refused to cooperate, and the diagnosis proved revealing: their Fisher information about the effective spin parameter was disproportionately concentrated, leaving a single correction factor unable to balance the books. Switching those cases to the exact Fisher-preserving scheme restored order, cutting the required sample count to at most 128. Convergence times fall roughly linearly with the number of retained samples — but only down to a point, since below about 256 samples the posterior fragments into spurious modes and the sampler, navigating artificially jagged terrain, slows again despite the cheaper likelihood. Extrapolated to the billion-sample records a ten-year mission could deliver, the most favourable cases approach speed-up factors of seven million, and even after budgeting for the handful of extra runs needed to verify convergence, gains of order ten-thousandfold or more remain plausible for LISA-like datasets.
The insistence on the time domain is strategy rather than nostalgia. Many of the effects scientists most want to model — environmental influences on an inspiralling pair, small departures from general relativity carried by extra dynamical fields, the Roemer time delays accumulated as a binary orbits a massive third companion — are natural to express as modifications of time evolution and awkward to formulate in the frequency domain. Informal tests on seventeen-parameter models, complete with the Keplerian orbit of a black hole binary around a distant massive companion, produced highly consistent recoveries, hinting that accuracy survives the climb in dimensionality. Downsampling also occupies a distinctive niche among likelihood accelerators: reduced-order quadrature requires detector-specific offline training, heterodyning and relative binning depend on a pre-selected reference waveform, and adaptive frequency resolution lives in the frequency domain. Downsampling, by contrast, needs no retraining or recalibration when the waveform family, the priors or the injected physics change — precisely the flexibility demanded by large-scale waveform-bias surveys, in which signals carrying extra physics are repeatedly recovered with deliberately incomplete models to chart where analyses go wrong.
There is one boundary the method cannot cross: because downsampling silently rewrites the effective noise model, it is strictly a tool for simulated data and cannot be applied to real LISA measurements. Within that simulated kingdom, however, it arrives ready for action. The paper provides the theoretical foundation for Dolfen, a pip-installable Python package that plugs directly into the standard Bilby inference pipeline, implements every variant of the procedure, and prescribes a disciplined workflow: coarse exploratory runs to map the likelihood’s rough landscape, high-accuracy runs at default settings, and convergence checks at increasing sample counts until the posteriors stabilise. Future versions will automate the handling of data gaps and telemetry drop-outs, and the author points to combinations with complementary accelerators as the next frontier. LISA remains roughly a decade away, but the millihertz sky it will open is already being rehearsed on ordinary computers — and if this study is any guide, the dress rehearsal may cost a millionth of what anyone dared to fear.
Cite Scienmag News
Grant Pearson. (August 30, 2026). Downsampling speeds up likelihood approximations for simulated LISA data. Scienmag. https://scienmag.com/downsampling-speeds-up-likelihood-approximations-for-simulated-lisa-data/
Grant Pearson. "Downsampling speeds up likelihood approximations for simulated LISA data." Scienmag, 30 August 2026, https://scienmag.com/downsampling-speeds-up-likelihood-approximations-for-simulated-lisa-data/. Accessed 30 August 2026.
Grant Pearson. "Downsampling speeds up likelihood approximations for simulated LISA data." Scienmag. August 30, 2026. https://scienmag.com/downsampling-speeds-up-likelihood-approximations-for-simulated-lisa-data/

