Low-Resource Reasoning: What SFT Changes and RL Repairs
A Greek reasoning study finds that nearly unchanged accuracy can hide major behavioral gains. SFT teaches models to reason in the user's language, while RLVR repairs formatting and channel-control failures.
Read more