Back to articles
Reinforcement Learning

LoGRA Cuts the Memory Barrier for LLM Reinforcement Learning

3 min read

Introduction

Reinforcement learning has become an important way to improve the reasoning behavior of large language models, but RL post-training is memory-intensive. In addition to model weights, a training system must handle gradients, optimizer states, and information needed to keep policy updates synchronized. These costs make larger models difficult to train on modest GPU clusters. LoGRA, introduced by researchers affiliated with NVIDIA, addresses the problem by compressing the learning signals retained during RL updates.

How LoGRA works

The central idea is to represent useful gradients as low-rank sketches instead of keeping fully dense gradient information throughout training. A low-rank representation is smaller and can therefore reduce the memory footprint of the update process. The sketches are also used beyond the optimizer step: they support model updates and more efficient policy synchronization.

Compression alone, however, can make an update unsafe. If an approximate gradient produces a step that is too large, the policy may move too far from its previous state and destabilize learning. LoGRA addresses this issue with predicted-KL step control. Before an update is applied, the method estimates the resulting policy change and adjusts the step magnitude accordingly. In this design, gradient compression is paired with an explicit mechanism for limiting policy drift.

The abstract reports several results:

  • LoGRA reduces average training memory by up to 45.7% on reasoning tasks without sacrificing performance;
  • it enables stable training of a 27B-parameter model for more than 1,100 steps on one eight-GPU node;
  • this is achieved in a setting where dense Adam runs out of memory;
  • implementation is available in the LoGRA example scripts of the Molt library.

Why it matters

The practical significance of LoGRA is that it may turn memory-infeasible RL experiments into manageable workloads. Lower memory use can reduce the hardware barrier for teams studying post-training, and it may allow larger models to be explored on a fixed GPU allocation. The method also highlights an important systems point: gradient compression cannot be evaluated independently of training stability. Preserving a compact learning signal is useful only if the resulting policy updates remain controlled.

The available material also leaves several questions open. The abstract does not provide the full sweep of low-rank settings, task-by-task comparisons, communication costs, or behavior across alternative RL algorithms. The reported 45.7% figure should therefore be read as the maximum reduction stated by the paper, not as a guaranteed gain for every workload. Broader validation across model architectures, training horizons, and deployment environments will be important for assessing how widely LoGRA can be used.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles