Some Fun Discoveries While Building and Iterating on My agent-harness
This post was translated from Chinese by AI. If anything reads oddly, the Chinese original is authoritative. 中文原文
No ads, typed by hand, and a bit messy. I'm browsing the code and writing things down as I go. To avoid anything that could count as advertising or leaking information (this is mostly for my own use, and I don't want trouble), I'll redact quite a bit. Thanks for understanding. Now, when claude gets rate-limited, I'll try my own agent for small tweaks. I'm still not confident enough to give it big tasks.
It's basically an agent harness framework, with most of the theory and methods borrowed from codex and claude code.
The desktop app is built with wails.
Sometimes I ask claude what tools it has that I could add to my agent (lol).
A few screenshots first

Core principles
Tools for the agent
These are the basics, such as ask_user_choice, memory, call_mcp, workspace, read image, save file, send file, and so on. Without tools, it can't do anything.
How tools are called
Usually through a round trip, so the agent can loop and decide for itself what to do next.
Tool feedback
Tools must return a result in every situation, whether stopped, paused, failed, successful, or otherwise. Without that, the agent can't know what happened or continue.
The agent process must be able to run in the background
It shouldn't need the page to stay open to run. Previously, the frontend passed results back to keep the backend making calls. That didn't work: switching to any other frontend page where the component wasn't mounted would pause the task. It's better to keep things running in the backend of a local app and send system notifications when user intervention is needed or the task is complete.
Preventing runaway execution, human intervention, and reasoning memory
- The agent needs a limit, and every run should produce a deliverable. Ideally, I'd eventually remove the hard limit and let a decisions model make that call.
- Human intervention: important operations need human oversight (for example, using ask_user_choice). Human instructions take priority over the agent's plans. So users can actively intervene in plans, stop work, and so on.
- Reasoning memory: this can't be lost, or the agent keeps forgetting things. The chain of thought must stay aligned with tool calls, or things get confused.
Four-layer memory
I built this before the harness framework and desktop app. Even a plain web chat needs memory. I think I posted about it before, but I forget. It's basically standard rag with 4 layers of memory.
Just make sure it doesn't affect the cached prefix; constantly losing the cache gets expensive
Context engineering
Speaking of caching, static cached content and dynamic content need to be kept separate in the agent.
- Prompts need to be assembled in layers.
- Context compression should be progressive.
- Tools should also be disclosed progressively. Dumping them all into the context at once takes up too much space, most aren't needed for a given task, and they distract the model.
The web version needs a remote sandbox
The web version also needs to generate files for the three main office applications. This is usually done with skills+ai writing python scripts.
I run these scripts using a remote machine (not the main project machine) + a persistent container (provides an interface and can launch temporary containers) + temporary containers (disposable execution sandboxes, destroyed when done).
This is essential. Running them locally = asking for disaster.
Local runtime environment
Some time after installing the desktop app, install a dedicated, isolated python environment for it. Don't use the machine's main python installation.
Ecosystem interoperability
Existing tools and prompts from claude and codex should all work and integrate seamlessly.
I just haven't gotten around to it yet. Otherwise, I could automatically import claude's memory documents too, making the integration even more seamless.
Security engineering
The main goal is to prevent external information from compromising the agent's security under any circumstances. After all, the agent runs on the local computer.
Security needs to cover every aspect and every step.
Closing thoughts
I asked claude to write the content in the image below. Damn, it said my backend api showed traces of one-api, so I couldn't describe the api as entirely self-developed. It changed that to "redesigned and implemented." Very respectful of copyright and the facts.
I've actually changed more than 70% of the original one-api code, and it still recognized it.

I mainly built this out of interest, and ended up with a usable tool. When it helps me write code, calls mcp to publish posts, or runs on my wife's computer to troubleshoot problems, it feels really rewarding.
Comments 0