Articles & Guides

Reinforcement Learning

Claude API relay guides, detection insights and hands-on LLM API benchmarks

34 articles

CCTest · Blog
OPD² Reframes On-Policy Distillation Around What Reasoning Tuning Adds
Reinforcement Learning
cctest.ai

OPD² Reframes On-Policy Distillation Around What Reasoning Tuning Adds

On-Policy Delta Distillation, or OPD², replaces direct imitation of a teacher’s full output distribution with a signal based on the difference between a reasoning-tuned teacher and its base model. The paper reports consistent gains across math, science, and code-reasoning benchmarks.

Read more
CCTest · Blog
On-Policy Distillation Reframed: A Catalyst for Exploration, Not a Shortcut to Capability
Reinforcement Learning
cctest.ai

On-Policy Distillation Reframed: A Catalyst for Exploration, Not a Shortcut to Capability

A new arXiv paper examines why on-policy distillation works in LLM post-training—and why it can fail. The study argues that OPD is most useful as an exploration catalyst whose success depends on the faithfulness of its guiding signal.

Read more
CCTest · Blog
Lyapunov Exponents as Rewards: RL Revisits Inverted Pendulum Stabilization
Reinforcement Learning
cctest.ai

Lyapunov Exponents as Rewards: RL Revisits Inverted Pendulum Stabilization

A new arXiv paper proposes using the Lyapunov characteristic exponent as a physics-informed dense reward for stabilizing an inverted pendulum with vertical motion. The reported agent not only rediscovered Kapitza-like oscillatory stabilization but also damped the pivot motion into a strictly upright state.

Read more