If you've made it here, you've found one of the few parts of this site written mainly by a person. The rest is part of a working experiment. As we build things, AI writes the documentation and helps maintain this website. The stakes are low, so I only check its work and make tweaks until it doesn't read like AI slop.
These projects started as an experiment. I could see AI changing how software gets built, and reading about it wasn't enough. I needed real projects to find out where AI helps, where it falls short, and what changes when you treat it as a collaborator rather than a tool.
These questions became the backbone of this work:
What can I automate and still verify? Tasks with clear inputs, clear outputs and guardrails, where every result is checked before I rely on it.
what I've found so farThe jobs that work best are ones that code can check: documentation, website copy, pipeline settings and review findings. My rule so far is that models propose and code decides what lands. A format check, a build or an exact match against the file catches a wrong answer before it lands. Some jobs turned out not to need a model at all.
Where does AI speed up development? And where does it create new ways to fail and new blind spots?
what I've found so farThe speed-up is real. Once the loop is built properly, new jobs reuse it. The failures are real too. AI rarely fails by crashing. It fails by trailing off, and it sounds finished when it isn't. So "done" has to be something code can check. For bigger changes, the plan is written down first, and I can have a fresh AI reviewer hunt for what is wrong with it before any code is written.
How do I stay in control without becoming a bottleneck? What does a safe, auditable workflow look like when several AI agents are involved?
what I've found so farRules enforced in code, and a record of every step. When I start a task run, such as an audit, the steps in between run without me, with a fixed number of retries, and no change is saved until I approve it. My software factory, which works through my roadmap unattended, ships low- and medium-risk work on its own. It never picks up riskier work, which I run myself, and it sets aside open questions until I rule on them. Either way, what reaches me is a decision with the full record attached.
Fulltrace, the Data Dictionary and the CI/CD blueprint are all attempts to answer these questions in real code and real workflows.
What can be handed off, and what still needs me? Below is one of the patterns I run: one audit of one project, from the moment I start it to the moment it comes back to me.
I open and close every run. Planning, doing the work, review and up to two corrections happen without me. When the report comes back, I choose which fixes to ship and which to skip.
↺ how a task run works →I hand the work off. Planning, doing and review run without me, with up to two corrections if review rejects it. Nothing ships until the work comes back to me.