From Art to Science - A Sigmoidal Scaling Framework for Predictable Reinforcement Learning Compute in Large Language Models
Pre-training LLMs has become a science of predictable returns on compute. But what if the subsequent, crucial phase of reinforcement learning remains an expensive gamble? How can we reliably forecast the success of an RL strategy before committing tens of thousands of GPU hours?
Decoding AI’s Black Box: Reinforcement Learning’s Transition from Alchemy to Predictable Science
In the world of Large Language Models (LLM), we have grown accustomed to a nearly deterministic path of progress. Thanks to established power-law scaling, the relationship between input and output during the model's pre-training phase is clear: invest more compute resources, and you are rewarded with predictable performance gains. This has transformed multi-million dollar investments from sheer gambling into calculated engineering.
However, once a model graduates from pre-training and enters the crucial Reinforcement Learning (RL) fine-tuning phase, that clear picture rapidly dissolves into fog. RL is essential for unlocking an LLM’s truly powerful capabilities, such as complex reasoning and its potential as an Agentic system. Even models like Deepseek-R1-Zero committed an astonishing 100,000 H800 GPU hours to RL training. Yet, unlike pre-training, this domain remains an “art” rather than a science, even with ten times the computational budget. Researchers have long engaged in an expensive lottery, unable to reliably forecast the performance lift a massive investment will yield, making the development and selection of scalable RL methods a costly and time-consuming gamble.
The core insight of this analysis is the solution to this challenge: by employing a systematic theoretical framework, we can peer into the final outcome of large-scale investment based solely on early training dynamics. This allows developers to scientifically screen for RL strategies with genuine scaling potential during low-cost experimentation, fundamentally shifting this “art” toward a predictable “science.”
I. The New Rosetta Stone: The Sigmoidal Scaling Law for RL
Unlike the power-law that governs LLM pre-training, the theoretical cornerstone proposed here is the more familiar Sigmoidal Curve—or S-curve—commonly encountered in business school and economics to describe the adoption rate of new technologies or product market penetration. The S-curve perfectly captures the non-linear relationship between resource input and return, characterized by slow initial growth, a rapid mid-phase explosion, and eventual saturation.
This shape is highly suitable for describing RL training because, regardless of the compute invested, a model’s performance will always have a theoretical ceiling.
Explaining the Sigmoidal Curve Parameters

Source: https://arxiv.org/html/2510.13786v1
By analogizing the RL training process to a capital investment project, the S-curve’s three core parameters gain immediate financial relevance, becoming the critical metrics for evaluating any RL method's scaling potential:
- Asymptotic Pass Rate (A): This represents the ultimate performance ceiling the model can theoretically achieve with infinite compute resources. In investment terms, this is the highest potential Return on Investment or the Total Addressable Market. When designing an RL method, all efforts should prioritize raising parameter A, as it dictates the maximum possible performance.
- Scaling Exponent (B): This determines the steepness of the curve and represents Computational Efficiency. Financially, it’s akin to Capital Efficiency or Growth Velocity. A larger B value means the model approaches its performance ceiling A faster, yielding a higher return per unit of compute invested.
- Midpoint Compute (): This is the amount of compute required to reach half of the total potential gain. This can be viewed as the capital required to reach the break-even point or half the market target. A smaller means the project enters its explosive return phase sooner.
The framework’s greatest strength lies in its predictive power. Researchers no longer need to run the entire, costly training sequence. By collecting data points during early training (e.g., using just 1.5k GPU hours) and fitting these parameters, they can scientifically extrapolate the method's performance at 100,000 GPU hours and beyond.
Predicting ScaleRL Performance at 100,000 GPU Hours via Sigmoidal Curve

Source: https://arxiv.org/html/2510.13786v1
The figure above perfectly illustrates this capability, accurately predicting the scaling trajectory of the ScaleRL method across different model sizes.
II. The Best Practice Recipe: Predictable ScaleRL
Guided by the robust S-curve theory, the research team conducted over 400,000 GPU hours of systematic empirical studies, ultimately distilling a "best practice recipe" named ScaleRL.
The development of ScaleRL adhered to several key principles, including a necessary reevaluation of domain common sense:
- RL Performance Ceilings are Malleable: Unlike pre-training, the RL performance ceiling (parameter A) is not fixed; it can be altered by design choices in the training process (such as loss type or batch size). Therefore, finding designs that elevate A is the primary objective.
- Accepting the "Bitter Lesson": A method that performs best in small-scale experiments may fail to scale. Only methods that the S-curve predicts will have a higher asymptotic rate A are truly worth the long-term investment.
- Revisiting Conventional Wisdom: Many traditional RL tricks (like advantage normalization) are often just shortcuts; they might boost short-term efficiency (parameter B) but have negligible or even negative impact on the final performance ceiling (A).
Through systematic optimization, ScaleRL not only achieves predictable scaling but also comprehensively outperforms existing mainstream methods like MiniMax and DeepSeek in both final performance and computational efficiency.
Comparison of ScaleRL with Existing Mainstream Methods

Source: https://arxiv.org/html/2510.13786v1
As shown above, ScaleRL achieves the highest Pass Rate for the same compute input. Crucially, its Asymptotic Performance (A) reaches 0.610, significantly surpassing other comparative methods, while its Scaling Exponent (B) is also high at 1.97, demonstrating its superior combined effectiveness.
III. Inside the Black Box: Which Design Choices Are Critical?
The success of ScaleRL is not due to a single silver bullet, but rather the synergy of a series of critical design choices. Interestingly, the impact of these choices can be clearly categorized: some are designed to raise the ceiling (Parameter A), while others are aimed at boosting efficiency (Parameter B).
1. Asynchronous Pipeline-RL: The Leap in Efficiency
In RL training, the model must alternate between "generating data" (inference) and "updating parameters" (training). Traditional PPO methods are synchronous, akin to an inefficient assembly line where production must halt until the first part is fully processed. ScaleRL adopts an asynchronous Pipeline-RL setup, allowing different GPUs to work concurrently in a true pipeline: one set of GPUs generates data while another immediately begins training on the previous batch.
This setup significantly reduces hardware idle time, resulting in a substantial increase in Computational Efficiency (Parameter B). Experiments showed that PipelineRL boosted efficiency B from 2.68 to 4.55. Although the ultimate performance ceiling A remained similar, this configuration enables the model to reach that ceiling much faster.
2. FP32 Precision: Ensuring "Account" Accuracy
In general LLM training, using lower floating-point precision (like FP16 or BF16) is standard practice to save memory and compute. However, this research found that maintaining full FP32 precision in the critical LM head (the model’s final output layer) is vital for RL training.
The rationale is that RL training is extremely sensitive to numerical precision. The LM head calculates the policy gradient, the "steering wheel" that guides the model's behavioral adjustments. If low precision is used here, minute numerical differences, much like "rounding errors" in complex financial calculations, can be rapidly magnified in RL’s sensitive gradient computation, leading to numerical instability and capping the model's learning potential.
The results are striking: simply using FP32 precision at the LM head drastically raised the Asymptotic Performance Ceiling (Parameter A) from 0.520 to 0.610. This confirms that in the "actuarial science" of RL, ensuring the precision of critical calculations is the key to elevating the ultimate potential return.
Using FP32 Precision in the Final Layer (LM head) Significantly Increased Asymptotic Reward

Source: https://arxiv.org/html/2510.13786v1
3. Other Key Components: Loss Function and Data Filtering
- Loss Function: The loss function dictates how the model learns from reward signals. The CISPO loss function adopted by ScaleRL proved capable of achieving a higher performance ceiling A than traditional methods like DAPO, while also demonstrating greater robustness across hyperparameter choices.
- Zero-Variance Filtering: In the training data, some user prompts yield the same reward regardless of the model’s response. These "zero-variance" data points contribute nothing to the policy gradient. Filtering them out not only boosts training efficiency but also subtly improves the asymptotic performance A.
IV. Macro-Scale Validation: Is the Predictability Universal?
A truly robust framework must maintain its efficacy under various conditions. ScaleRL demonstrated the universal applicability of its framework through large-scale experiments across multiple axes: model size, generation length, and batch size.
- Model Size (MoE): Training a larger 17Bx16 MoE (Mixture-of-Experts) model with ScaleRL showed the same predictable S-curve scaling behavior as the 8B dense model, but with a higher asymptotic RL performance. Notably, this stronger MoE model surpassed the peak performance of the 8B model using only 1/6 of the RL training compute.
- Generation Length (Context Budget): Increasing the model’s required generation length from 14k tokens to 32k tokens resulted in a slower initial training phase (lower efficiency B), but ultimately raised the asymptotic ceiling (Parameter A). This proves that long-context RL is a necessary investment to improve the performance limit, rather than a simple efficiency trade-off.
- Global Batch Size: Larger batch sizes (e.g., 2048 prompts) were reliably shown to increase the asymptotic ceiling (A). This further confirms the "bitter lesson": small-batch runs, which might appear more efficient early on, are eventually surpassed by large batches as compute scales up.
Conclusion and Strategic Implications
The introduction of the ScaleRL framework successfully elevates LLM RL training to a level of predictability comparable to pre-training. This shift has profound implications for LLM development and deployment strategy.
- Increased Capital Efficiency: The S-curve framework allows organizations to scientifically predict and screen the most promising RL strategies at a small computational cost (days, not months), avoiding the waste of potentially millions of dollars in GPU hours on strategies that are fundamentally unscalable.
- Clear Optimization Roadmap: This research provides a crucial roadmap. When designing next-generation RL methods, priority must be given to factors that raise the Asymptotic Performance Ceiling (A) (e.g., correct loss functions, FP32 precision, larger models/context), rather than solely focusing on marginal gains in Efficiency (B).
- A New AI Valuation Methodology: From a broader strategic perspective, this work provides a powerful analytical tool. It enables AI developers and investors to systematically assess the potential return ceiling (A) and capital deployment efficiency (B) of different technological paths, transforming the RL process from an opaque gamble into a measurable, predictable engineering science.