// about

Jason Newbury

The human-in-the-loop

If you've made it here, you've found one of the few parts of this site written mainly by a person. The rest is part of a working experiment. As we build things, AI writes the documentation and helps maintain this website. The stakes are low, so I only check its work and make tweaks until it doesn't read like AI slop.

Why I Do This Work

These projects started as an experiment. I could see AI changing how software gets built, and reading about it wasn't enough. I needed real projects to find out where AI helps, where it falls short, and what changes when you treat it as a collaborator rather than a tool.

These questions became the backbone of this work:

What can I automate and still verify? Tasks with clear inputs, clear outputs and guardrails, where every result is checked before I rely on it.

what I've found so farThe jobs that work best are ones that code can check: documentation, website copy, pipeline settings and review findings. My rule so far is that models propose and code decides what lands. A format check, a build or an exact match against the file catches a wrong answer before it lands. Some jobs turned out not to need a model at all.

Where does AI speed up development? And where does it create new ways to fail and new blind spots?

what I've found so farThe speed-up is real. Once the loop is built properly, new jobs reuse it. The failures are real too. AI rarely fails by crashing. It fails by trailing off, and it sounds finished when it isn't. So "done" has to be something code can check. For bigger changes, the plan is written down first, and I can have a fresh AI reviewer hunt for what is wrong with it before any code is written.

How do I stay in control without becoming a bottleneck? What does a safe, auditable workflow look like when several AI agents are involved?

what I've found so farRules enforced in code, and a record of every step. When I start a task run, such as an audit, the steps in between run without me, with a fixed number of retries, and no change is saved until I approve it. My software factory, which works through my roadmap unattended, ships low- and medium-risk work on its own. It never picks up riskier work, which I run myself, and it sets aside open questions until I rule on them. Either way, what reaches me is a decision with the full record attached.

Fulltrace, the Data Dictionary and the CI/CD blueprint are all attempts to answer these questions in real code and real workflows.

The work isn't finished. It's a living record of how AI and humans build things together, and of making that collaboration reliable and useful.
the human-in-the-loop

What can be handed off, and what still needs me? Below is one of the patterns I run: one audit of one project, from the moment I start it to the moment it comes back to me.

I open and close every run. Planning, doing the work, review and up to two corrections happen without me. When the report comes back, I choose which fixes to ship and which to skip.

↺ how a task run works →
A loop of five steps. I hand the work to a planning step. Next the work is carried out, then reviewed. Only if the review rejects it does it go to a correction step, which sends it back for another review, at most twice. When the review approves it, or the corrections run out, the work comes back to me. Every step except mine sits inside the handed-off zone. The loop starts and ends with me. handed off only if review rejects it max 2x re-check In this pattern, nothing ships without me. Me Plan Execute Review Correct
Me I start the run. When the report comes back, I pick which fixes to apply, and nothing is written until I confirm. Each code change needs its own approval.
Planning agent Turns the request into a plan. A low-cost model does this, DeepSeek by default with Qwen as backup, and I can pick a different model for any run.
Execution agent Does the work with tools, on Kimi with Qwen as backup. A dropped or rate-limited model call is retried, then the step restarts on Qwen. A badly formed tool request goes back to the model once to fix.
Review agent DeepSeek, a different model from the one that did the work, approves it or sends it back with reasons. Plain code also checks the result if I switch on Verify, and always before a fix.
Correction Only when review rejects the work. The execution model fixes what the reviewer flagged, at most twice. Fixes to my projects are a separate, later step, drafted by Claude Sonnet 4.5 for findings rated safe.

↺I hand the work off. Planning, doing and review run without me, with up to two corrections if review rejects it. Nothing ships until the work comes back to me.