Answer · trust and control
How do I review code that an AI wrote if I cannot read code?
Updated 2 October 2026
You review the evidence instead of the syntax. Four checks give you most of what you need. If the project has a verified test command, ask for a test that failed before the change and passes after it. That shows the test covers the case you care about. Ask for a review from a model that did not write the code. The configured review model can vary. Require a green continuous integration run on a clean checkout. That reports whether the configured checks passed outside the machine that made it. Then use the feature yourself and confirm it does what you asked. Keep your own approval on anything that touches money, customer data, or access control. A wrong answer there is costly, and a passing test can still miss it.
This does not make you a reviewer of the code itself, and it is not meant to. It makes you the person who decides if the evidence is good enough. That is a job you can do.
What to do, in order
Test the behaviour yourself, first
Open the change and use it the way a customer would. This is the one check that needs no technical skill. It also catches the most embarrassing kind of failure: code that runs fine and does the wrong thing.
Demand a test that failed before the change
A test written with a feature can pass without checking anything. When the project has a verified test command, ask for red first proof. The test goes in without the change, so it must fail. Then the change lands, and the same test must pass. If that proof is not available, report the test evidence you do have. Do not claim this proof.
Get a second opinion that shares no context
Ask a different model to review the change and argue against it. Use a fresh session. A model that reviews its own work brings its own assumptions to the review. So the useful review comes from a session with no memory of writing the code.
Require a green build on a clean checkout
Continuous integration runs the whole test suite on a fresh copy of the repository. Reading code cannot answer one question. Does it still work when the author's machine is gone? A clean run answers that.
Keep the risky changes on your own desk
Payments, login, permissions, personal data, and anything that deletes or migrates records. Those need your approval. The other four layers may have done well. That does not cover them. When the stakes are high, ask a qualified human to look too.
What each layer proves, and what it does not
A test that fails first proves the test measures the change. It does not prove the change is a good idea. An independent review catches reasoning errors and missing cases. It can still miss what nobody thought to look for. A green build reports that its configured checks passed in a clean environment. It does not report that the suite covers the thing you care about. Your own use of the feature proves the behaviour is right. It tells you nothing about the code paths you did not touch. These layers help because each one fails in a different way.
Why the second opinion has to be independent
A model that reviews its own output tends to agree with it. The same thing happens when an author proofreads their own writing. They read what they meant. The page says something else. A reviewer that has never seen the reasoning works from the diff alone. A human reviewer is in the same spot, and that is why the review is worth something.
How Keelen runs these layers for you
A separate adversarial reviewer reads the diff. It does not see the conversation that wrote the code. The review model comes from your project settings. Red first proof and clean checkout tests need a verified project test command. Gated auto-merge needs the repository checks and a review window you set up. Plan review, manual merge, and branch-only mode each pick where a person must decide. Check the real evidence before you accept the result.
What the test checks can catch
Fail first and clean checkout tests need a verified project test command. Gated auto-merge needs the repository's required checks to be set up. The red first proof gate applies new tests without the code change. They must fail first. Tests that pass at that point are rejected. When the assertion guards are on and can check a change, they block detected net removal of assertions. They also block deletion of existing test files that contain detected assertions. They can detect smaller expected sets or lowered numeric floors. They can also detect contentless matchers and newly disabled tests. Each guard has a supported scope. Supported exceptions can let those changes through.
What still needs a person
Declared exceptions use an Assertion-removal: or Assertion-relaxation: commit trailer with a reason. The agent can add these trailers itself. The gates do not check human approval. The guards can be turned off. They can skip a check if the comparison base is missing or the Git diff fails. Some changes and patterns fall outside detection. A passing or skipped guard does not prove that no checks were weakened. The tests and code still need human review. With manual merge, Keelen opens the PR and you click merge.
What happens when the checks cannot be satisfied
Keelen handles infrastructure faults. It holds those apart from real dead ends. Provider quota failures are handled the same way. They are not real dead ends. A real dead end can become a decision card. The card uses plain English. That card can pause the project. Recovery does not promise uninterrupted progress. It does not promise zero provider charges.
Ask for the tool list
This fictional workshop equipment tracker shows how to follow a change through the work record. It is not a customer result or a record of a completed project. You write the Request in plain English. Ask for a list of checked out tools, with the borrower and the due date. State what counts as acceptable behaviour. Choose manual merge when you want the merge decision to stay with you. Manual merge is one of the project's merge policies. Keelen opens the pull request, and you click merge.
Check the planned task
Keelen classifies the Request into the project's standing roadmap. Then it expands the top item into a planned task with acceptance criteria. Read the task before any code is written. Check that its criteria match your Request. Criteria that drift from what you asked for start the mismatch. The task should carry the details you named. It should keep the acceptable behaviour you asked for.
Read the pull request
Keelen implements the task as a pull request on your repository. Open the pull request and compare its diff with the planned task. A pushed branch or an open pull request is not a merge. Branch only mode pushes a branch with no pull request. Manual merge waits for your click. Until a merge is recorded, the change has not landed.
Inspect review and checks
Read the independent review and the actual checks on the pull request. When the project has a verified test command, Keelen checks that applicable new tests fail without the change. On a clean checkout, those same tests must pass. Without that command, this failing test proof is not available. A separate review is more evidence. So are the required CI checks you set and the review window you picked. The reviewer does not have to be a different model. Inspect the actual result. That includes failures and skipped checks. Do not read those as a pass.
Confirm the outcome
Confirm the actual merge state of the pull request before you rely on the change. Then use the project's deployment process to make the tool available. Check the requested behaviour in the running app. Keep three states apart. A pull request exists. A merge is recorded. A deployment happened. Only a recorded merge and a completed deployment put the tool in front of the coordinator who needs it.
Follow a later change
Suppose weeks later the coordinator wants one more view: only the overdue tools. Submit a second Request and follow it the same way. Read the planned task. Open the pull request. Inspect the review and the checks. Then confirm the merge and the deployment. This continuation is hypothetical. It works because the due dates already exist in the tool. It is the same loop applied to a smaller change instead of a new project.
When Keelen is not the answer
- You can read the code. Ordinary review practice is better than this checklist, and you should use it.
- Your software may be regulated. It may be safety critical or life affecting. That work needs a qualified human audit. No set of automated checks can take its place.
- You want a guarantee. This raises the floor a long way and removes the most common failures. It is not a proof of correctness, and anything that claims to be one is overselling.
Hand the work to a loop
Connect a repository and write what you want in plain language. Review the tested pull requests that come back.
FAQ
How do I stop my AI agent from changing or deleting tests?
Fail first and clean checkout tests need a verified project test command. Gated auto-merge needs the repository's required checks to be set up. The red first proof gate checks that new tests fail without the code change. Tests that pass at that point are rejected. When the assertion guards are on and can check the change, they block detected assertion removal or weakening within their supported scope. Supported exceptions can let changes through. The agent can declare an exception itself. That does not prove human approval. Guards can be turned off or skip checks when inputs are missing. They do not catch every way a test can lose meaning. The tests and code still need human review. Manual merge keeps the merge click with you.
How do I know if AI-written code is safe to merge?
Check the evidence around it instead of the code itself. Look for a test that failed before the change and passes after it. Look for a review from a model that did not write the code, and for a green build on a clean checkout. Use the feature yourself. Keep your own approval on anything that touches money, customer data, or access control.
Can AI reliably review code that another AI wrote?
It catches a useful share of real problems. That includes missing cases and reasoning errors. This only works when the reviewer is a separate session that did not write the code. It is not equal to a qualified human on high risk changes. Treating it as equal is the mistake to avoid.
What should I never let an AI merge without a human looking?
Payments and personal data top the list. Authentication and permissions matter too. The same goes for a database migration that deletes or rewrites records. Those changes are hard to undo. A passing test suite can still stay silent. No automated check settles those.
Do I need to learn to code to run a product this way?
No. You do need to say exactly what you want and judge whether the result matches. That is a product skill instead of a programming one. Learning to read a diff does help later. It is not the entry requirement.