Why LLM Agents Keep Following Instructions Users Have Withdrawn
Introduction
Users rarely express a complete, stable request in one message. They may ask for a draft, change its format, remove a constraint, or explicitly withdraw an earlier decision. A human collaborator would normally treat the latest valid instruction as authoritative. An LLM agent, however, does not automatically erase information that has become obsolete. Earlier requirements can still influence the final response or appear in a tool call.
The paper When Users Change Their Minds: Measuring and Repairing Intent Drift in LLM Agents studies this failure mode under the name “intent drift.” Its contribution is not simply to report that agents make mistakes in long conversations. It proposes a way to test the problem systematically and evaluates a concrete mitigation strategy.
Key findings
- IntentFlux creates controlled dialogue tests. The researchers transform verifiable tasks into conversations containing changes of intent while preserving the original task graders. This makes it possible to compare a model given the final task directly with a model that must recover the same task from an evolving dialogue.
- More obsolete information means worse results. In a 627-case calibration, mean task score declined from 0.476 to 0.384 as the dialogue contained more superseded and withdrawn information. Across eight models, fully correct solutions were significantly less common when the final task had to be reconstructed from a multi-turn exchange rather than stated in one turn.
- Explicit state maintenance helps. StateForge maintains the requirements that are currently active before generation or execution. On General-Test, it increased the mean task score from 0.367 to 0.467.
- State tracking is only part of the problem. Supplying the ground-truth final state improved performance further, but still did not recover single-turn performance. The remaining gap suggests that errors also arise in how the model uses the state, produces the answer, or carries out the action.
Why it matters
The study challenges the idea that better context management simply means retaining more conversation history. For tool-using agents, the important operation is to distinguish active constraints from modified requirements and explicit withdrawals. If obsolete instructions remain equally influential in the prompt, an agent can execute an outdated plan despite having sufficient knowledge and access to the right tools.
IntentFlux is useful because it turns a vague conversational failure into a reproducible evaluation setting. The task, the dialogue changes, and the grader can be examined separately. StateForge, meanwhile, suggests that inserting a structured state-maintenance step before execution is a practical direction, although not a complete repair.
For deployed systems, this points to the value of confirmation and state visibility. Before taking an irreversible tool action, an agent could expose the goal and constraints it believes are currently active, then ask for confirmation when it detects a conflict. The paper does not claim to have solved intent drift, but it offers a clearer basis for measuring it and comparing future agent designs.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...