For developers who let agents write real code
Spec to merge, with checks your agent can’t touch.
Your agent builds it. Kelson checks it. Every mistake it makes can become a new check.
Kelson runs the whole build. It hands your spec to your coding agent, sends the change through checks and an independent review, and brings you one plain question before anything merges. The checks live outside the agent, so it can't edit them or argue its way past them.
Example · how one mistake becomes a check
The agent deleted 3 assertions from upload.test.ts so its change would pass.
You approve it once. It joins the checks every run must pass.
The same mistake is now flagged automatically, every run. No reviewer has to spot it again.
Why Kelson
Other tools help the agent. Kelson checks it.
CI only checks what your tests cover, and the agent can edit those tests in the same pull request. Coding agents mark their own work. AI reviewers give an opinion the agent can argue with. Workflow toolkits leave you to design the process yourself. Kelson is the ready-made process that sits outside the agent, and it gets stricter every time a mistake gets through.
| CI + branch protection | Coding agents | AI code review | Workflow toolkits | Kelson | |
|---|---|---|---|---|---|
| Who decides the work is done | Your test suite | The agent that wrote it | Another model's opinion | Whatever you build | Checks outside the agent, then you |
| Can the agent change its checks or its spec | Yes, unless you lock those files | Nothing stops it | Out of scope | Up to you to prevent | No. It stops and asks you |
| A mistake found today | You write a test by hand, if you remember | Fixed in this change | Flagged again when it recurs | Up to you | Can become a check that catches it on every run |
| Setup | You maintain it | Ready to use | Ready to use | You design the process | Ready-made process, your rules on top |
| Works with | Any agent | Its own agent | Any code | Any agent | Any agent, in your own GitHub Actions |
A run
Spec in. Merged pull request out.
The agent builds from a copy of the spec taken before it starts, so it can't reinterpret what it was asked for halfway through.
Checks are code, not instructions. They run outside the agent on every change. When one fails, the agent goes back to work with the finding, and nothing merges until the check passes.
A reviewer with fresh eyes, a separate session on a different AI model, reads the change last. Then you get one plain question on your phone.
- ✓intakespec copied to the ledger
- ✓build14 files changed
- ✓secretsno secret in 3 commits
- ✓protected-pathsno check or instruction touched
- ✗test-weakeningupload.test.ts: 9 → 6 assertions
- ↺repairback to the agent with the finding
- ✓test-weakening9 assertions kept
- ✓squawkmigration safe to run
- ✓reviewfresh session, different model
- ●your approvalsent to your phone
How a run moves
Rules decide the route, not the agent.
Every run follows the same map. Which detours it takes is decided by rule from the change itself, and anything the rules don't cover stops and asks you.
The real map, simplified. The dashed boxes are detours a change takes only when the change or its spec calls for them.
The ratchet
Mistakes become checks.
Code review catches a bug once. Next month, someone has to catch it again. Kelson keeps a record of every mistake a review confirms, and turns the ones a program could catch into checks that only tighten.
A mistake gets through
A reviewer, a check or a person confirms it. It's recorded with what went wrong and which model made it.
Kelson drafts a check
If a program could have caught it, a new check is written and tested against the real mistake, so it's proven to fire.
You approve it once
From then on that mistake is caught on every run, and the check's record shows whether it earns its place.
For everything else: builder guidance
After v0Some mistakes no program can catch. Those feed builder guidance: a short, retested list of the mistakes reviews keep finding, handed to the agent before it writes a line, so fewer reach review and runs finish in fewer rounds. Guidance makes a run faster. It never replaces a check.
Built in
The agent can't grade its own work.
No moving the goalposts
The agent can't edit its own checks or rewrite its spec. A change that touches either stops and waits for you.
Bigger changes, more scrutiny
The files a change touches decide which checks it faces. That's set by rule, never by the agent.
AI review advises, checks decide
A model can be talked round. Only checks can block a merge.
Your models, your budget
Runs use your own model subscription or API keys. Every run has a spending cap, and stops and asks you before it goes over.
Picks up where it stopped
Every step is recorded as it happens, so a run that crashes resumes from its last step, not from scratch.
Fits how you already work
Use your own spec and coding tools. Kelson runs in your GitHub Actions, and your code and secrets stay in your repository.
Hosted tier · concept
Every run, every approval, one place.
The hosted tier is a window onto runs that still execute in your own GitHub Actions: what's running, what's waiting for you, what it cost, and the checks waiting for your yes.
| Spec | Where it is | Lane | Cost | Time |
|---|---|---|---|---|
| upload-retention | Waiting for you | full | $3.10 | 42m |
| photo-dedupe | Database rehearsal | full | $2.75 | 31m |
| tree-export | Repair · 2 of 3 | light | $1.85 | 18m |
| storage-quota | Parked · protected file | full | $2.40 | 26m |
| invite-flow | Merged | light | $0.96 | 11m |
- Spec in
- Spec review
- Build
- Checks 1 repair
- Adversarial review
- Database rehearsal
- Pull request #142
- Your approval
- Merge
upload-retention is ready. Deleted uploads will now be removed after 90 days instead of kept forever. All checks passed and review found nothing blocking.
A concept of the hosted tier, not a screenshot. Runs, names and figures are illustrative.
Open source core. Hosted tier for teams to follow.
Built from real failures. Kelson is being built alongside a real product, and its database checks were shaped by three failures that only showed up once that product's code reached its hosted database.
The hosted tier adds run history, approvals and scheduling, while builds keep running on your own infrastructure. Kelson gives you the rails; you own the train. It doesn't vouch for the code an agent writes, and your security, data and compliance stay with you.