Batch mode worked. You’d hand localagent-box a repo and a prompt, go make a coffee, and come back to a branch. For small, well-scoped tasks that was enough.
Then I started giving it bigger jobs, and every fix I made uncovered the next problem. None of them were in my original plan. Each one only showed up once the previous one was solved and the system was doing more work than before.
I promised last time that this post would be about interactive mode. It got built, it works, and it turned out to be the least interesting part of the whole thing. The problems worth writing about came after, and I think the order I hit them in is the fascinating part.
Getting agents to work for longer
One opencode run with one prompt has a ceiling. A local model with a modest context window can hold one focused change in its head. Ask it to “add validation to the login form and write tests and update the API docs” and it either loses the thread halfway through or declares victory early.
So I built loop mode. Instead of one shot, the worker drives the agent through a repeating cycle until the goal is met:
- INITIAL_PLAN: survey the codebase and write a checklist of milestones, each small enough to finish in one iteration
- ORIENT: look at the next milestone and the relevant files, and decide on the smallest next change
- ACT: make that change
- REFLECT: evaluate progress and say what’s next, or emit a completion marker if the goal is done
Steps 2–4 repeat up to maxIterations. The whole cycle is config-driven. There’s a default loop.json on the server and a repo can replace it with its own in .localagent-box/loop.json.
A few guard rails turned out to matter more than I expected. The host checks whether the working tree changed between iterations, and if it’s identical twice in a row the run stops as stalled instead of burning iterations going round in circles. Hitting maxIterations commits the partial work rather than throwing it away. And each step can use a different model. Locally I ran a variant of qwen3.8-27b for the coding work and sometimes dropped REFLECT down to gemma4-12b, since judging progress doesn’t need the big model.
It worked. Agents could now grind through multi-step tasks unattended. It also used a frightening number of tokens.
Loop mode was a token furnace
“Tokens are free when you run locally” is true for the bill. It isn’t true for time. On local hardware, every extra token of input is extra seconds of prompt processing, and a small model with a stuffed context window gets noticeably dumber on top of being slower. And once I started using Ollama Cloud to run GLM5.3 for the heavier jobs, it was about money again too.
When I looked at what each iteration was sending, the waste was obvious:
- Too many steps. The original cycle was OBSERVE → PLAN → ACT → REFLECT. Every step in an iteration shares a session, so each later step replays everything before it, including every file the agent read. Four steps meant a lot of replay.
- The same facts, twice. Progress lived in a markdown checklist inside the repo clone, and the full REFLECT output from the last iteration was injected into the next prompt. Both said the same thing.
- Tool calls for bookkeeping. ORIENT spent a tool call reading the plan file. REFLECT spent one or two more editing it to tick boxes, which it did unreliably.
The fix came down to one principle: the host should own anything deterministic. The model is expensive, slow and forgetful. Node.js reading a JSON file is none of those things.
So OBSERVE and PLAN merged into a single ORIENT step. REFLECT stopped touching files altogether and now has to answer in a fixed shape:
DONE: <what changed this iteration>
REMAINING: <what is left>
NEXT: <the single next change>
FILES TOUCHED: <comma-separated paths>
The host parses that, ticks the matching milestones itself, and stores the state in the agent’s data directory rather than in the repo, so there’s no chance of it being committed by accident. When the next iteration starts on a fresh session, the host injects only what the model needs: the next unfinished milestone, the one-line NEXT from the last REFLECT, and the output of git status and a diffstat so it doesn’t have to go rediscovering what changed.
Because ORIENT and REFLECT no longer write anything, they run as OpenCode’s read-only plan agent, which also means a smaller tool schema in every request.
The other big win was letting a repo define a checkCommand. After ACT, the host runs it (usually the test suite) and feeds the exit code and the tail of the output into REFLECT. If the check fails, the host ignores any completion marker. The model doesn’t get to decide it’s done while the tests are red.
Now I had to review all of it
With loops running longer and more reliably, I had a new problem: a pile of branches to review. Reading every agent diff by hand undoes most of the point of fire-and-forget.
I added a fourth agent mode, review, built on Open Code Review (OCR). It’s an open source CLI that diffs two branches, reviews the changed files with your LLM of choice, and returns structured findings. A review agent reuses the same clone-and-checkout pipeline as a coding agent, but runs ocr review instead of OpenCode:
ocr review --repo <workspace> --from main --to agent/login-validation \
--format json --audience agent
If there’s a PR for the branch, the findings get posted as a proper GitHub PR review: a summary with a severity breakdown, plus line comments on the diff. Turn on auto-review and every PR a coding agent opens gets a review queued automatically. The review also gets context from the coding session that produced the branch, so it knows what the change was meant to do rather than guessing.
That naturally led to the next question: if the review found problems, why am I the one fixing them? So findings above a configurable severity now kick off autofix agents. They work through the findings in small batches, one agent at a time, resolve the GitHub review threads as each batch lands, and finish with a single verification review over the result. The review also shows up as a GitHub check run, so you can make it a required status check.
A good example is the PR that added direct Ollama Cloud support and rebuilt the settings page. It touched 40 files. OCR, running glm-5.3-flash on Ollama Cloud, reviewed 26 of them in about 21 minutes and came back with 36 findings: 10 medium and 26 low.
The most useful one was a startup bug. The server only regenerated the OpenCode config when a local ollamaBaseUrl was set, so a cloud-only install running in a container would lose its OpenCode config on every restart until someone re-saved Settings. It also pointed out that the new Ollama Cloud API key was being written to opencode.json with default file permissions. Neither of those would have jumped out at me in a 40-file diff.
The low findings were mostly maintainability nitpicks: duplicated provider labels, an unused export, two refresh handlers that did exactly the same thing. Fair enough, but not why I built it.
It isn’t cheap, either. That one review used 3.7 million tokens, and three files didn’t get reviewed at all because the session hit its context compression threshold. Two of those three were the review runner and the review flow, so the code reviewer couldn’t review its own code. I found that funnier than I probably should have.
None of this replaces a human reading the diff. It does catch the things that slip past me in a big PR, and it means the PR I end up reading is a better one. There’s still plenty of refinement to do, starting with making sure it gets through every file and cutting down how many tokens it takes to get there.
Chaining work into one PR
With loops and reviews in place, individual agent runs were in good shape. The limit was now the size of a single run. Loop mode stretches what one agent can do, but there’s still a point where a piece of work is better split into chunks, each with its own prompt, all landing on the same branch and the same PR.
The original design didn’t allow that. A second session on a branch that was already in use got a 409 BRANCH_IN_USE, and checkout always created the branch fresh from main.
The shared-branch queue fixes both. You can now create several sessions against the same branch, and they run strictly one at a time on it:
- Only one worker per repo and branch at a time, while other branches still run in parallel up to the global concurrency limit
- Session N+1 starts only once session N has completed and pushed
- N+1 fetches the branch from the remote instead of recreating it from base, so it builds on N’s work
- If N fails, everything after it waits, with a visible reason. You can retry N in place or tell the queue to skip it and start the next one
I deliberately didn’t build a pipeline system. No multi-prompt composer, no drag-to-reorder, no new job type. It’s the existing FIFO queue with a per-branch mutex and a rule about predecessors. The git handoff is push-then-fetch, which sounds wasteful until you remember that every workspace is isolated and disposable anyway.
The piece I’m happiest with is how it plays with review. The first chunk to push opens the PR and later chunks just update it. Auto-review holds off while there’s still coding work queued on the branch, then runs once over the finished whole rather than reviewing every half-done intermediate state.
This is now how most of localagent-box itself gets built. I write a plan doc that breaks a feature into PR-sized tasks, then queue a batch session per task on one branch. Reporting reviews as GitHub check runs landed as 15 queued batch commits on a single branch before I merged it as one PR. The workspace bootstrap in the next section was built the same way, one branch per phase, with 7 to 11 chunks each.
Plain batch mode, not loop mode, does most of this work now. Once the plan doc has split the work into small tasks, each chunk is the kind of focused, well-scoped job batch mode was always good at.
Agents spending half their time on npm install
Queuing ten small chunks instead of one big run meant ten times the setup, and watching them go through the logs one after another made a pattern hard to miss. The first few minutes of nearly every run went on the agent working out how to set up the repo. Find the lockfile, run npm ci, wait, run the tests, discover it needed a build step first, run that, try again.
Every chunk starts from a fresh clone, so all of that was repeated every single time. And it was the worst possible work to hand a model: slow, token-hungry, and fully deterministic.
Same principle as before: move it to the host. Before the OpenCode session starts, the worker now runs a workspace bootstrap. A repo can tell it what to do, in order of precedence:
- A committed
.localagent-box/setup.shscript, for anything multi-step - An explicit
setup.commandin.localagent-box/environment.json - Named runtime profiles like
nodejs-pnpmorpython - Lockfile auto-detect, when the operator turns it on
// .localagent-box/environment.json
{
"version": 1,
"profiles": ["nodejs-pnpm"],
"cacheKey": "acme-monorepo-pnpm9",
"verifyCommand": "npm test -- --passWithNoTests"
}
The verifyCommand runs after setup as a smoke test, and a failure there always stops the agent before it starts. I’d rather an agent fail fast with “your environment is broken” than spend ten minutes trying to fix a broken environment and then fix my code on top of it.
Once setup succeeds, the host prepends a short workspace ready block to the agent’s first prompt saying what ran, so it doesn’t go looking for how to install dependencies all over again.
Fresh clones also meant a fresh npm ci every time, which was often slower than the agent’s actual work. There’s now an optional dependency cache on a persistent volume, keyed by a hash of the lockfile (or an explicit cacheKey). On a hit, the host restores node_modules into the fresh clone before the setup command runs. It’s off by default and only covers the Node profiles for now, because caching is exactly the kind of thing that works fine until it quietly doesn’t.
The same lesson, five times
Looking back, every one of these fixes follows the same pattern. I started out treating the model as the thing that does the work, with the harness as a thin wrapper around it. I’ve ended up with the opposite: the harness does everything it can deterministically (tracking progress, running checks, managing branches, installing dependencies) and the model only gets the parts that need judgement.
That matters more with local models than it does with frontier ones. A frontier model can shrug off a messy context and a few wasted tool calls. A 27B model on my own GPU can’t, and every token you save is one it doesn’t have to wade through.
What I have now is a set of parts that work well on their own. The missing piece is the workflow that connects them. Right now I’m still the glue: I write the plan doc, break it into tasks, queue the sessions, and keep an eye on the reviews.
Next I’m wiring it all up. Linear holds the tasks, Slack is where I talk about the work and hear back about it, and localagent-box does the work in between. A feature gets broken down into Linear issues, the issues become queued sessions on a branch, and the reviews and PRs come back to me in Slack. The pieces are all built. Now they need to talk to each other.