Back to articles
Large Language Models

A Model Can Transfer Coding Ability Without Showing Any Code

3 min read

Introduction

Can a model learn to code without ever seeing code? A study from Peking University suggests that it can, at least under a carefully designed experimental setup. The central observation is that post-training may leave traces in decisions that appear unrelated to the capability being trained.

The paper introduces Active Taskless Distillation, or ATD. Rather than asking a coding-trained teacher to solve programming problems, the method examines how that teacher chooses between ordinary words in unrelated contexts. The researchers first identify prompts for which the shared base model is almost indifferent between two candidate words. They then measure which word the post-trained teacher selects and train a student, initialized from the same base model, on those prompt–word pairs.

Key findings

  • Only one word is used as the target. The student does not receive code, target-task examples, teacher parameters, or teacher logits.
  • Near-ties make small updates visible. When the base model has almost no preference between two words, a slight shift caused by post-training can reveal the direction of that update.
  • Coding performance improves. In the primary Qwen2.5-1.5B experiment, 5,664 single-word examples yielded a 5.34 percentage-point improvement on HumanEval+ compared with an exactly nuisance-matched control.
  • The effect is broader than programming. The study reports transfer involving scientific knowledge, commonsense reasoning, and reading comprehension across additional model generations, sizes, and families.
  • The transferred behavior appears composable. Functional analyses suggest that the learned change can combine with other capabilities, while its strength tracks the magnitude of the teacher’s update.

Why it matters

The result extends earlier work on subliminal learning, which largely focused on preferences or traits transmitted through unrelated generations. ATD suggests that capability-relevant information can also be distributed across many low-information behavioral choices. A single ordinary word carries little signal by itself, but a deliberately selected collection of such choices may encode a detectable direction of model change.

For model developers, this raises the possibility that post-training evaluation should include unrelated prompts, not just the benchmark targeted by the training process. For safety and model-governance research, it also creates a new question: can capabilities be indirectly extracted or transmitted through outputs that do not look task-relevant?

The evidence should still be interpreted within the limits of the supplied material. The reported experiments do not establish that the approach will work equally well for every model scale, training recipe, or adversarial setting. ATD is best viewed as a tool for probing the behavioral shadow of post-training, rather than as a general replacement for task-specific data.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles

CCTest · Blog
Instead of Imitating Answers, LLMs Could Learn by Avoiding Bad Reasoning
Large Language Models
cctest.ai

Instead of Imitating Answers, LLMs Could Learn by Avoiding Bad Reasoning

Negative Self-Distillation introduces a reverse form of self-distillation: rather than copying a privileged solution trace, a model is trained to move away from its own flawed reasoning. The approach aims to preserve exploration and self-correction without requiring external labels.

Read more