Research and evaluation

Can an AI agent learn on the job? Measuring improvement per dollar

A game-based study measures how frozen-model agents improve through saved tools and notes. What its results and learning costs can tell a business buyer.

By Clairevue · · 5 min read

Chess pieces, saved strategy cards, a reusable tool block and blank cost tokens sit beside a separate test tile.
AI-generated illustration of game-based learning and evaluation costs; not the paper’s actual experiment or results.

Suppose a company’s internal helpdesk agent saves lessons after routing requests to the wrong team. In this fictional example, the company pays for extra model calls to review mistakes and update the agent’s instructions. Before expanding that budget, it needs evidence that the next version routes unfamiliar requests more accurately.

A new paper calls that acquisition efficiency agent plasticity: improvement on held-out tasks relative to the cost of learning. More saved notes alone don’t establish an improvement.

Harman Singh and co-authors from Meta Superintelligence Labs and universities published the first version on 6 October 2026. Their experiments use games, rather than helpdesk work, to study agents that revise their own tools and notes while their underlying model stays fixed.

What survives when the conversation ends

In the paper’s board-game protocol, each game starts with a fresh context. The model can inherit persistent Python tools and written strategies, but it doesn’t carry the previous conversation forward. Between rounds, that same model reviews training-game evidence and proposes changes to those files.

The model’s weights remain frozen. The agent can use a better move-picking program and revised instructions from its saved files. An update has to pass structural and executable checks before it becomes persistent; passing those checks doesn’t guarantee better play.

For chess, Go and Hex, the researchers withhold evaluation-game trajectories and scores from the revision process. They can then compare each saved version’s results on games at the training difficulty and against stronger opponents. The stronger opponents test transfer across difficulty, not whether a chess lesson helps with an unrelated business task.

In Hard chess, the authors report that Claude Fable 5’s held-out score rises from 37.5% initially to an average of 73.3% over checkpoints 16–20. That score gives half credit for draws, so it isn’t a pure win rate. GPT-5.6 Sol reaches a late average of 36.9%, having started near zero.

Held-out games can also contain positions encountered during training, so these results don’t establish success on wholly unfamiliar situations.

Count the learning bill, not only the final score

The paper counts model-call costs for playing training games and revising persistent files, including rejected revision attempts. Held-out board-game evaluation calls don’t enter that learning-cost denominator. These dollar amounts are a pricing-based proxy for computation, not a complete deployment budget.

Across the combined chess, Go and Hex curves, Claude Fable 5 reaches the highest fitted final performance, while GPT-5.6 Sol has the highest estimated improvement per dollar up to saturation. The authors fit curves to estimate where learning approaches a plateau; some estimates require extrapolation. Provider prices and those fitting assumptions affect the ranking.

This efficiency measure doesn’t tell a buyer which agent already performs well enough for a job. A low starting score can leave plenty of room to improve, while the improved version still trails a stronger agent. Nor does the experiment isolate the benefit of experience from the extra computation spent developing tools.

For the fictional helpdesk, suppose the original agent correctly routes 70 of 100 held-back test cases and an updated version routes 80. If the learning and reflection calls, including rejected attempts, cost $50, the gain is 10 percentage points, or 0.2 points per dollar. Those are illustrative numbers, not a measured helpdesk result.

Staff review, evaluation runs and maintenance would add to the company’s bill. The gain also needs repeated checks because outputs can vary, and an aggregate score can hide mistakes in a costly request category. Ten more correct classifications don’t establish ten requests resolved or a return on investment.

A saved rule can still fail to change behavior

The authors describe a GPT-5.5 chess agent that repeatedly stores checks against a known early pawn-capture mistake. Its opening book returns a move before those checks run, and the mistake persists. The agent has written down a correction without making it govern that decision.

The paper’s broader analysis uses Gemini 3.1 Pro judgments alongside execution logs to distinguish artifacts that are missing, available but unused, or used while a failure remains. These are observational diagnoses. They don’t prove that forcing an agent to read a note would fix its mistakes, or separate poor advice from poor application.

Honcho’s Dreaming documentation describes consolidating conclusions and replacing outdated ones; it labels the feature experimental. Before paying for such maintenance, ask for before-and-after task results and the cost of producing them.

Test the next version before paying for more revisions

For a helpdesk trial, I’d agree on routing labels with staff and use synthetic examples or permissioned requests with private details removed, representing actual request categories. Keep difficult cases visible instead of letting easy requests dominate the average. OpenAI’s evaluation guidance similarly recommends task-specific tests and human feedback to calibrate scoring.

Give the agent one set of cases to learn from and keep the evaluation cases out of its revision process. Pin the model and runtime settings, record which artifact version each run uses, and start each case in a fresh context. Once you select a candidate, confirm it on a separate final test rather than choosing the best-looking version from repeated peeks at one test set.

Review proposed code and instruction changes before allowing customer-facing actions, and keep the previous version available for rollback. Log whether the agent uses its saved material, but make routing accuracy and harmful mistakes the decision criteria. More tool use is not a successful outcome by itself.

Set the next learning budget against an agreed improvement target, with evaluation and staff time accounted for separately. If repeated held-out checks show no worthwhile gain, retain the better verified version and stop funding further revisions until there’s a specific failure to investigate.