AI Agents Learn to Avoid and Repair Their Worst Mistakes

Agents fail at the turn that matters. Two new approaches from NVIDIA and Google target that weakness from different angles: one teaches agents to prevent and repair damaging decisions, while the other gives them a durable system for evolving skills across tasks.
NVIDIA researchers, working with Princeton University and the University of Maryland, introduced PivotOPD, an on-policy distillation method for multi-turn LLM agents. Google’s WikiSkill builds a structured wiki from past experiences, creating a persistent knowledge layer that helps agents refine how they perform recurring workflows.
Both developments focus on a problem that benchmark charts tend to hide: an agent can produce fluent actions for several turns before one bad decision sends the entire task off a cliff. The systems differ in design, but their message is similar—successful agents need memory about failure, not just more chances to generate text.
PivotOPD Targets the Decision That Breaks the Task
PivotOPD defines a pivotal mistake as an action that lengthens the shortest remaining path to completing a task or makes the task unsolvable. Its training process adds three components to group-based reinforcement learning: pivot detection, preventive distillation, and recovery distillation.
The method trains an agent to avoid its most damaging early mistake, then teaches it how to recover when prevention fails. Against 13 baselines, PivotOPD delivered the best average results across ALFWorld, WebShop, and Search-based QA for Qwen3-1.7B and Qwen3-8B students.
With the 1.7B student, PivotOPD averaged 73.7% on ALFWorld, 5.5 points above SDAR, and 44.5% on Search-based QA, 5.9 points above RLSD. With the 8B student, it reached 93.0% on ALFWorld, 47.4% on Search-based QA, and 81.9% success on WebShop.
The recovery results are the more revealing measure. PivotOPD recovered from 72.7% of 72 replayed pivotal mistakes, compared with 8.3% for the base model, 20.3% for standard OPD, and 45.8% for preventive-only training.
Standard OPD lowered the overall failure rate from 79% to 56%, but failures after a pivotal turn fell only from 51% to 49%. PivotOPD corrected the pivotal turn in Qwen3-8B failure replays and raised success from 8% to 59%—a reminder that reducing failure is not the same as understanding it.
Recovery still took work. PivotOPD averaged 9.7 turns to recover, against an optimal 6.2 turns. The method also held its lead when Qwen3-8B served as its own teacher, winning all three benchmarks by at least 1.5 points.
WikiSkill Turns Agent History Into Working Knowledge
Google’s WikiSkill addresses the same multi-step problem through persistent skill evolution. It divides the agent workspace into three layers: the Raw Layer, the Wiki Layer, and the Skill Layer.
Each evolution cycle runs training tasks, analyzes execution trajectories, proposes new skills, and validates improvements before adoption. The architecture preserves execution traces, extracts recurring patterns, and tests changes rather than allowing every new lesson to become permanent by administrative accident.
In one case study, WikiSkill refined a broad goal-directed-action skill into a more concrete “break-repetition-loop” skill. The system achieved the highest average score for every model tested across five benchmarks, and its gains grew with model size: 12.3 points at 4B, 17.5 points at 9B, and 23.9 points at 27B models.
Skills also helped smaller models close part of the gap with larger ones. Qwen-3.5-9B with WikiSkill reached 47.4% accuracy, compared with 39.4% for Qwen-3.6-27B without skills. In another result, Qwen-3.5-9B scored 70.2% on ALFWorld using a skill evolved by Qwen-3.6-27B, compared with 63.4% using its own evolved skill.
WikiSkill’s system uses a ReAct Skill Proposer that takes roughly 10 to 20 turns per iteration, plus a Wiki Maintainer call. Its experiments also found that keeping the wiki out of the inference agent’s context improved performance, even though active skills are injected into the model prompt.
Liyan Tang described the design this way: “We kept Karpathy’s shape: immutable sources, an LLM-maintained wiki, an index and log, but the input is the agent’s own execution trajectory, and the output is an executable SKILL.md.” The reference is to Andrej Karpathy’s shape, but the important change is practical: the stored knowledge comes from what the agent actually did.
WikiSkill does not test skill retrieval or triggering as skill libraries grow. That leaves a clear pressure point, especially for systems managing many skills across many workflows. For now, the strongest fit is agents performing multi-step tasks with recurring failure modes—the precise territory where PivotOPD also finds its advantage.
Together, the approaches point toward agents that do more than generate the next action. They must identify the turn that causes failure, recover from it, preserve the lesson, and apply that lesson later. The impressive part is not that models can make mistakes. That feature remains well-supported.
Based on




