Fine-tuning a large language model usually demands something deceptively simple: gradients. Every step of conventional training relies on backpropagation, the algorithmic machinery that propagates error signals backward through billions of parameters to tell each weight how it should change. But there are increasingly common scenarios in which that machinery is unavailable. A model may be accessed only through an application programming interface that returns outputs but hides its internals. It may be so large, or stored so aggressively quantized, that keeping activations in memory for a backward pass is prohibitively expensive. It may sit on hardware where automatic differentiation simply is not supported. In all of these black-box and resource-constrained settings, researchers have turned to zeroth-order optimization, a family of methods that updates parameters using nothing more than the values the model returns when queried. A new study published in the International Journal of Machine Learning and Cybernetics by Ximin Zhang and Shuhua Yuan of the Henan Engineering Research Center of Fault-Tolerant Server in Zhengzhou, China, argues that the standard versions of these methods carry a hidden cost that grows catastrophically with model size, and proposes an elegant way to pay far less.
The core idea of zeroth-order optimization is old and intuitively appealing. Rather than computing a gradient analytically, one estimates it by perturbing the parameters slightly in random directions and observing how the loss changes. If nudging a weight upward makes the model worse and nudging it downward makes it better, the estimator knows, roughly, which way to push. Two function evaluations per step, in the so-called two-point estimator popularized in modern language model work such as the MeZO algorithm, are enough to produce a usable descent direction. The catch is statistical. A random direction in a space of millions or billions of dimensions is almost certainly nearly orthogonal to the true gradient, so the difference in loss between two nearby points along that direction is dominated by noise. The variance of the resulting gradient estimate scales with the dimension of the perturbation space, and in parameter-efficient fine-tuning, where the trainable set may still contain millions of parameters, that variance can be severe enough to make training crawl or stall.
Parameter-efficient fine-tuning itself was supposed to solve the scale problem. Techniques such as adapters, prefix-tuning, prompt tuning, BitFit and the now-ubiquitous low-rank adaptation known as LoRA freeze the bulk of a pretrained model and train only small inserted modules or subsets of weights. This slashes memory for optimizer states and makes adaptation feasible on modest hardware when gradients are available. But as Zhang and Yuan point out, even a reduced trainable set frequently reaches into the millions of parameters, and a zeroth-order estimator perturbing all of them simultaneously inherits variance proportional to that count. The reduction that makes PEFT attractive under backpropagation does relatively little for the statistical burden of gradient-free methods. Previous lines of work have attacked the problem from several directions, including sparse perturbations, transferable static sparsity, low-rank structures, curvature-aware estimates, preconditioned and accelerated variants, and random subspace methods that restrict perturbations to a randomly chosen subset of coordinates. The new paper builds on this momentum but changes where the compression happens.
The proposed framework, called Projected Latent Subspace Zeroth-Order Optimization, or PLS-ZO, restricts perturbations not to a subset of coordinates but to a deliberately constructed low-dimensional latent subspace, and it does so at the level of individual weight matrices rather than the flattened parameter vector. The authors construct hierarchical low-rank random projections that map each trainable matrix into a controllable latent space of much lower dimension. Random projection is a classical tool from linear algebra and randomized numerical computation: a high-dimensional vector multiplied by a random matrix retains, with high probability, the geometric information that matters, compressed into far fewer coordinates. By applying this idea hierarchically across the matrices of the model, PLS-ZO creates a latent coordinate system in which the optimization actually takes place. Perturbations are drawn in this latent space, and the difference of loss values between perturbed and unperturbed forward passes yields an estimated gradient with variance governed by the latent dimension rather than by the original parameter count.
A crucial and somewhat distinctive design choice follows from this. Many subspace methods treat the projection as a way to generate cheaper perturbations but still run the optimizer in the original high-dimensional coordinates. PLS-ZO instead performs the main optimizer operations, including momentum accumulation, preconditioning and the bookkeeping of optimizer states, directly in the low-dimensional latent coordinates. Momentum in particular benefits: an exponential moving average of past gradient estimates is a powerful variance reducer, but maintaining and applying it over millions of coordinates is exactly the kind of overhead zeroth-order methods are supposed to avoid. Working in latent space keeps those state tensors small and makes preconditioning, which rescales update directions to account for uneven curvature, computationally practical. The result is an optimizer whose memory footprint and per-step arithmetic scale with the chosen latent dimension, a hyperparameter the practitioner controls, rather than with the full size of the trainable parameter set.
One subtlety of any random-subspace approach is that the subspace must eventually change. A single fixed random projection risks missing important directions of variation in the loss landscape for the entire run, so PLS-ZO periodically refreshes its projections, drawing new random bases in which optimization continues. Naively, each refresh would force the optimizer to discard its accumulated momentum and preconditioner statistics, throwing away the very variance reduction that makes the method work, or would require expensive re-projection of large state tensors. The authors introduce what they call a lazy subspace refresh strategy, which defers and amortizes the cost of regenerating projections, together with a principled state migration rule that projects the latent optimizer states from the old basis into the new one. In effect, the optimizer’s memory of its trajectory survives the change of coordinates, so long-horizon training remains stable instead of repeatedly restarting from a cold state.
Beyond the algorithmic engineering, the paper contributes theory. The authors prove that the bias and variance of the two-point zeroth-order estimator under their scheme are governed by the latent subspace dimension rather than the original parameter dimension, formalizing the intuition that compressing the perturbation space compresses the noise. They further establish a nonconvex convergence guarantee for the setting in which the random subspaces are periodically refreshed, showing that the algorithm converges to a stationary point at a rate consistent with classical zeroth-order analysis despite the moving coordinate system. Nonconvex guarantees of this kind matter because the loss surfaces of fine-tuned neural networks are emphatically nonconvex, and a method whose theoretical behavior degrades under basis changes would be difficult to trust in practice. The migration rule and refresh schedule are precisely the components that make the analysis go through.
On the empirical side, the authors report extensive experiments demonstrating that PLS-ZO consistently achieves highly effective fine-tuning performance under the parameter-efficient setting. The evaluation draws on widely used natural language understanding benchmarks of the kind standard in this literature, including datasets such as SST-2, BoolQ, WiC, WinoGrande and MultiRC that have anchored comparisons among gradient-free fine-tuning methods, applied to robustly pretrained encoder models in the RoBERTa family. The paper’s data availability statement notes that no new datasets were generated or analyzed, meaning the contribution lies in optimization methodology rather than data collection. The experiments position PLS-ZO against the backdrop of a rapidly crowding field: recent work has introduced low-rank zeroth-order fine-tuning, temporally low-rank perturbations, gradient-aligned projections, minimum-variance two-point estimators, Hessian-informed optimizers, quantized zeroth-order training, and learned optimizers, all racing to make forward-pass-only adaptation competitive with full backpropagation.
The significance of this line of research extends well beyond academic benchmarks. As large models are deployed at the edge, embedded in privacy-sensitive pipelines, or offered as closed services, the assumption that a practitioner can differentiate through the model is increasingly fragile. Zeroth-order fine-tuning requires only the ability to run the model forward and read off a loss, which makes it compatible with API-based adaptation, on-device training under tight memory budgets, and models whose weights are frozen by policy as much as by architecture. Methods like PLS-ZO attack the central statistical obstacle, dimension-induced variance, with a control knob, the latent dimension, that lets practitioners trade a small amount of expressive freedom for large gains in estimate quality and compute. If the theoretical promise holds up under wider independent testing, the practical upshot is that adapting a frontier-scale model could become feasible for organizations with neither the hardware nor the internal access that today’s fine-tuning workflows assume, a shift that could meaningfully democratize who gets to specialize the most powerful models in existence.
Subject of Research: Zeroth-order optimization for parameter-efficient fine-tuning of large language models
Article Title: PLS-ZO: projected latent subspace zeroth-order optimization for parameter-efficient fine-tuning
Article References: Zhang, X., & Yuan, S. (2026). PLS-ZO: projected latent subspace zeroth-order optimization for parameter-efficient fine-tuning. International Journal of Machine Learning and Cybernetics, 17(10), Article 488. https://doi.org/10.1007/s13042-026-03317-9
Image Credits: AI Generated
DOI: 10.1007/s13042-026-03317-9
Keywords: zeroth-order optimization, parameter-efficient fine-tuning, large language models, random subspace, low-rank projection, black-box optimization, gradient-free methods, LoRA, latent subspace, convergence guarantees, machine learning, optimization
Cite Scienmag News
Blake Davidson. (September 30, 2026). New Zeroth-Order Method Tames Gradient-Free Fine-Tuning of Large Language Models. Scienmag. https://scienmag.com/new-zeroth-order-method-tames-gradient-free-fine-tuning-of-large-language-models/
Blake Davidson. "New Zeroth-Order Method Tames Gradient-Free Fine-Tuning of Large Language Models." Scienmag, 30 September 2026, https://scienmag.com/new-zeroth-order-method-tames-gradient-free-fine-tuning-of-large-language-models/. Accessed 30 September 2026.
Blake Davidson. "New Zeroth-Order Method Tames Gradient-Free Fine-Tuning of Large Language Models." Scienmag. September 30, 2026. https://scienmag.com/new-zeroth-order-method-tames-gradient-free-fine-tuning-of-large-language-models/

