NEEDLE Removes LLM Backdoors with Weight Orthogonalisation
Introduction
Backdoors in large language models are difficult to detect because they can remain dormant during ordinary use. A model may answer benign prompts normally, yet switch to an attacker-chosen behaviour when a particular word, pattern, or input condition appears. In a deployed system, such a trigger could cause unsafe generation, unexpected actions, or code injection behaviour without being visible in standard capability tests.
Removing a backdoor creates a second problem. An aggressive intervention may reduce the attack success rate, but it can also shift the model’s responses to benign prompts. That shift may damage general capabilities or weaken safety refusals. A paper from Locai Labs introduces NEEDLE as an attempt to make backdoor removal more targeted and less disruptive.
How NEEDLE works
NEEDLE is training-free, although it assumes that the relevant trigger has already been identified. Its procedure has four central elements:
- Estimating the backdoor direction: Activation vectors from triggered and clean inputs are compared to identify the main representational shift associated with the trigger.
- Building a refusal subspace: Activations related to refusal behaviour are analysed to define representations that the intervention should preserve.
- Orthogonalising weights layer by layer: The output weights of attention and MLP components are adjusted so that relevant activations are less aligned with the backdoor direction.
- Correcting downstream effects: Editing earlier layers changes the inputs received by later layers. A closed-form correction is used to keep the edited activations’ projections onto the refusal subspace as stable as possible.
The distinction between these two spaces is the core idea. NEEDLE does not simply erase a broad region of the parameter space. It seeks to remove the direction associated with the trigger while protecting representations linked to safety refusals. It also requires neither a clean reference model nor the original poisoned training data.
Results and implications
The authors evaluate NEEDLE across multiple model families and attack types. Their report says that the method achieves the lowest mean attack success rate among the evaluated defences, including a reported 0% result on challenging code injection attacks. It also produces the lowest KL divergence in the comparison and causes minimal changes to measured capability and safety. These findings should be read within the paper’s specific triggers, models, and evaluation protocols; they do not establish that every unknown backdoor can be removed.
The broader contribution is a surgical view of model repair. Once a trigger is known, a defender may be able to intervene directly in internal representations and weights instead of relying on full retraining. The work also highlights why backdoor defences need more than one metric. Attack success rate matters, but so do benign-prompt behaviour, output-distribution drift, general capability, and refusal safety. Further testing will be needed for unknown triggers, multiple coexisting backdoors, and larger production models.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...