The question
A local agent's work is worth something only if the cloud agent can use it without doing the task again. Our research on delegation kept returning to that condition. A capable local agent whose result cannot be trusted is, for the cloud agent, the same as no result.
So this note asks a narrower question than "can a local model do the task?": what has to come back from a local agent for the cloud agent to accept the work? The question does not depend on which model runs locally or on which machine.
It applies to any agent that delegates, including subagents in the cloud. The local case makes it sharper in three ways. The cloud agent never sees the machine where the work happened, so everything it knows arrives in what comes back. The local agent usually runs a smaller model than the one reviewing it, so review is the norm. And if checking means reading the sources again, the files end up in the cloud context anyway: the saving and the benefit of keeping data on the machine disappear together.
Prior work shows that small local models can do useful work for a cloud model. In Stanford's Minions protocol, a cloud model splits a long-document task into subtasks that a local model runs over short chunks.1 Those results come from benchmarks with a fixed protocol. We look at the handoff inside the agents people already use, where the cloud agent decides what to delegate and whether to accept the result.
We studied the question inside real delegations: Claude Code and Codex handing tasks to a local agent through Forjal, on open-source codebases and document sets, from August to September 2026. What follows is what we learned about the handoff itself. The setup is at the end.
A handoff is a contract, not a call
A delegation can break at six points. At each one there is a signal that looks like success and is not.
| Stage | What has to happen | What does not prove it |
|---|---|---|
| Adoption | The cloud agent chooses to delegate | The local agent being installed and described to it |
| Admission | The local runtime accepts the task and starts it | A catalog lookup or a rejected request |
| Execution | The local agent reads, writes and runs what the task needs | A job marked running, a busy GPU, many tool calls |
| Delivery | The agent that asked receives a usable result | A job that finishes after that agent has moved on |
| Quality | The result does the task, with enough evidence | Valid JSON, literal quotes, a completed status |
| Useful economy | The accepted work costs less, review and repair included | Shorter answers, more errors, or cost moved to a step nobody measured |
Adoption is its own problem. A cloud agent that can delegate will often solve the task itself, even with a local agent installed and well described. When and why an agent decides to hand work off is a research question in its own right, and one of the open questions below.
Delivery is engineering, not intelligence. A delivered result has to be whole, not an answer cut off at an output limit and passed along as complete. It has to reach the cloud agent while that agent is still waiting, not after it has ended its turn. And when the local agent's context fills up, the work has to continue in the same job instead of being lost. None of this depends on the model.
The stages are judged in order. If nothing was delivered, the quality of the local work cannot be judged, and the failure counts against the system, not the model. Scoring an answer that never reached the cloud agent measures the wrong thing.
The result has to carry its proof
A correct answer that arrives without proof costs the cloud agent the same review as a wrong one. To accept it, the cloud agent has to check it, and checking without evidence means reading the sources again. That is the work delegation was supposed to save.
We found that evidence has three separate properties, and each one failed on its own:
| Property | The question it answers | What we saw when it was missing |
|---|---|---|
| Literal | Does the passage exist in the source, exactly as quoted? | Lines joined together or markers dropped, although the original lines had been read |
| Located | Is it where the answer says it is? | Real passages cited at the wrong line numbers, in one answer most of them |
| Sufficient | Does the passage support the conclusion? | A correct fact backed only by the fragment trace.events.filter( |
Research on attribution in text generation asks a related question: is a statement supported by an identified source?2 Our setting adds a twist. The reader of the evidence is another agent, and what it decides is whether to redo the work.
A machine can check literal and located. Sufficient takes judgment. The executor can compare a quote with the file and refuse it when it does not match. In the runs we audited, that check refused the inexact quotes and passed none that were wrong. A model can estimate whether a passage supports a claim, but that estimate is one more judgment the cloud agent has to trust. That is where review cost lives.
Literal quotes can decorate a wrong answer. One result carried four literal quotes while claiming two file edits that never happened and misstating what a setting does. Every quote was true. The claims around them were not. Evidence has to be tied to each claim, so that a claim without support reads as unsupported.
What the agent saw, read and cites are three different sets. In one case the agent saw a file name in a directory listing and reported a version number it had never read. In another, it read the helper functions that answered the question and left them out of its evidence. Only a record of what was actually read lets the cloud agent tell these apart.
A claim of absence is bounded by what was read. "There is no reference to this" and "I found none in the files I read" are different answers. The first went beyond the pages the agent had opened. The second, with the files named, is something the cloud agent can act on.
"I could not prove it" is a useful result. Our stricter acceptance mode refused a correct answer whose evidence was insufficient. That looks like a loss, but a refusal tells the cloud agent exactly what is still open. An unsupported answer that happens to be right gives it nothing it can rely on.
Receipts come from the executor, not the model
What a local agent says it did and what it did are two different records. Only the executor, the part of the system that actually performs reads, writes and commands, can keep the second one. Guidance on agents already stresses getting ground truth from the environment at each step;3 a receipt carries that ground truth back to the cloud agent.
"I edited the file" needs a write behind it. In the result with the four literal quotes, the text reported two edited documents. The workspace was identical before and after, and no write had been attempted. Without the executor's record, the cloud agent would have had to inspect the files to find out.
A test run by someone else is not the agent's test. In another task the local agent wrote a correct patch and meaningful tests. We confirmed both by running them afterwards, outside the job. Its own attempts to run them had not succeeded. Our check does not count as the agent's verification, and the receipt has to say which commands ran, how they exited and which were refused.
Asked for is not delivered. The agent asked for lines 121 to 289 of a log; a page limit returned 121 to 284, with an explicit note to continue at 285. A flag saying the output was not truncated would have been true and misleading. Completeness is judged from what was delivered, not from what was requested.
Unknown is not zero. A multi-step job has to account for every step, not only the last one. When a measurement is missing, the receipt says unknown instead of filling in a number.
The receipt we are building toward carries, for every delegated task:
- what the local agent was granted;
- what it actually read, as delivered ranges;
- what it wrote, and what it ran, with exit codes and refusals;
- evidence tied to each claim, checked by the executor;
- what is unfinished, and where to continue;
- what the local work cost, with gaps marked as unknown.
The review is part of the cost
The unit we measure is work the cloud agent accepts, with its review, its repairs and any lost attempts included. A delegation pays off only when the cloud work it avoids is larger than everything it adds, and the accepted result is just as good.
Cloud cost with delegation = cloud work that remains + handoff + follow-up and collection + review + repair or redo
The equation counts cloud work. Time and local compute are separate costs, recorded apart: a local run can take longer than the cloud agent doing the work itself, which matters when the next step waits for it.
Delegation has a price everywhere, not only between cloud and local. Anthropic reports that its multi-agent research system uses about 15 times the tokens of a chat, and notes that this pays off only when the task is valuable enough.4 The question is always whether the work avoided is worth more than the coordination added.
Every cloud turn rereads the conversation. Each time the cloud agent decides to delegate, writes the brief, collects the result, reviews it or repairs it, it processes much of its history again, even when part of it comes from cache. A delegation that saves one read but adds three turns can cost more than the read.
Less context per turn is not less cost per task. Smaller tool results make each turn lighter. If they also make the cloud agent take more turns to finish, the total can grow. We track both.
Tokens, API price and subscription quota are three measures. Cached input is priced differently, so a session can use more raw tokens and still cost less at API prices; we saw exactly that. Quota and energy we have not measured yet. Any number we publish says which of these it is.
The upside is real when the evidence is right. In a small pilot, cloud conversations given a short set of exact passages made the same three decisions as conversations given the full material, using about a quarter of the tokens per call. A researcher chose those passages. The pilot shows what the right evidence is worth, not that a local agent can already find it on its own.
We do not publish a savings figure for local delegation yet. The equation above is how we will measure one: on equivalent accepted work, with the attempts that failed counted in.
Where this points: delegate what is cheaper to check than to do
Our evidence points to a way of splitting the work that does not depend on how hard a task looks. Hand off the work whose result the cloud agent can check more cheaply than it could do the work itself.
| Zone | To do | To check | Tasks from our runs |
|---|---|---|---|
| Delegate | Slow | Quick | Answer a question from named sources; review documents, finding by finding. Still under test: diagnose a log, citing the lines; change and test one module |
| Do it directly | Quick | Quick | Look up a constant or a version |
| Supervise closely | Slow | Slow | Diagnose a subtle cause across modules; open exploration with no clear end |
| Keep in the cloud | Quick | Slow | Judgment calls on evidence already gathered |
The table places tasks in zones, not at exact positions. The asymmetry itself is well known: some tasks are much easier to verify than to solve.5 Between two agents it becomes a rule for what to hand off, applied by the cloud agent that owns the task.
The failure rate belongs in the rule. A task that is cheap to check still costs a redo every time the local agent gets it wrong. In full: delegate when the handoff, the check and the expected redo (chance of failure × cost of redoing) add up to less than doing the work directly.
Checking is cheap when the sources are bounded and named, the evidence is literal and located, and the receipt shows what was read, written and run. A test the agent ran itself, with its exit code, is cheap to check, although whether it tests the right thing still takes review. Checking is expensive when a conclusion crosses many files, when a claim is about absence, or when the task has no clear end to check against.
Some work is not worth handing off at all. When the cloud agent can find a constant with one search, a delegation adds a brief and a collection turn that cost more than the search.
Judgment stays with the agent that owns the task. Asking a local agent to decide, from evidence the cloud agent already holds, saves little and still has to be checked. In our runs, delegated judgment calls often had to be redone. Reading and gathering evidence travel better than the final call.
Better receipts move work into the delegate zone. Every property from the sections above, from literal evidence to executor records, lowers the cost of checking. That is the part a system can improve without waiting for a better model.
This is a direction, not a certified list. The zones come from a limited number of runs. Our next studies test whether the rule holds across real workflows and other models.
For builders: a checklist
If you are wiring one agent to hand work to another, local or not, these are the checks we now apply.
- Judge the six stages separately, in order. Never score the quality of work that was not delivered.
- Bind every job to what it was granted. The same prompt over different files is a different job.
- Record effects from the executor. Reads as delivered ranges, writes, commands with exit codes, refusals. Never from the model's summary.
- Tie evidence to each claim. Check literal and located by machine before anyone spends time on sufficient.
- Bound claims of absence. "Not found" names the files that were read.
- Accept "I could not prove it" as a result, with what is missing.
- Tell the agent its real limits. The limits and commands described to the model should come from the same code that enforces them. When the two drift apart, the agent plans actions the executor will refuse.
- Mark unknown as unknown in every count and receipt.
- Measure the whole task until acceptance: every cloud turn, the review, the repair and the attempts that failed.
- Separate the system's failures from the model's. Preparation, transport and harness errors get fixed in the system. When the model had enough evidence and still got it wrong, the result limits what you delegate; it does not call for a special fix that only passes that one case.
Open questions
These are the questions the lab is working on next.
- Who selects the evidence? Can a local agent find the passages a cloud agent needs, with the coverage our pilot got from a researcher?
- Does the rule hold? Is "cheaper to check than to do" a reliable guide across real workflows, other local models and work beyond code, such as documents and personal agents?
- How much can a receipt remove? Which parts of the cloud agent's review disappear when the receipt is complete, and which always need judgment?
- When does an agent decide to delegate? Availability did not produce adoption. What makes a cloud agent's decision to hand off well calibrated?
- When does background work pay? A local agent working while the cloud agent moves on helps only if the result arrives in time to be used.
Setup and method
The findings come from delegations run between August 8 and September 16, 2026.
| Item | Detail |
|---|---|
| Cloud agents | Claude Code 2.1 with Claude Sonnet 5; Codex 0.150 |
| Local models | Qwen3.5 9B, Ornith 1.5 9B, Qwen3.6 35B-A3B and Qwen3.8 27B, in several quantizations, on MLX and llama.cpp |
| Hardware | MacBook Air, Apple M4, 16 GB; Mac mini, Apple M2, 24 GB |
| Forjal | Releases 0.13.0 to 0.18.6 and experimental builds of the harness |
| Tasks | Questions, reviews, patches and tests on HTTPX, Werkzeug, Jinja, Click, Pluggy and Graphify; document sets of ADRs, RFCs and reports; a real 390-line log |
| Accounting | Cloud tokens from provider receipts; local work from executor receipts, kept separately |
Where we compared, the same task ran with and without delegation, from the same frozen sources. Results were reviewed against those sources for delivery, material errors and pertinent evidence. In the later protocols, reviewers sealed their verdict before seeing any cost. A result counted as accepted only when all three held.
The cases in this note illustrate mechanisms. They are individual observations, not rates, and the note reports no success rate or average saving.
The one figure is the evidence pilot: four new Codex conversations, run short, long, long, short, on the same three decisions. A researcher selected the short set of passages, and an AI model, not a blind reviewer, checked the decisions. Average raw tokens per call were 18,240 with the short set and 70,766 with the full material. That excludes the cost of finding and checking the passages.
Figures from the sources cited describe their own studies, not Forjal. Sources checked on September 27, 2026.
Notes
-
A. Narayan et al., "Minions: Cost-efficient Collaboration Between On-device and Cloud Language Models," arXiv:2502.15964. ↩
-
H. Rashkin et al., "Measuring Attribution in Natural Language Generation Models," arXiv:2112.12870. ↩
-
Anthropic, "Building effective agents," December 2024, anthropic.com. ↩
-
Anthropic, "How we built our multi-agent research system," June 2025, anthropic.com. ↩
-
J. Wei, "Asymmetry of verification and verifier's law," July 2025, jasonwei.net. ↩
