One in Five GitHub Reviews Is AI. Who Checks the Premise?
GitHub says Copilot now handles more than a fifth of code reviews. When the writer and the reviewer share assumptions, a green pipeline proves agreement, not correctness.
The Approve Button Is Doing a Lot of Work
GitHub says Copilot code review now accounts for more than one in five code reviews on GitHub. The company also says its review system can explore repository context, inspect large pull requests, review pull requests opened by bots, and re-check its own findings after changes are made.
Read that last capability slowly. A model can now review work produced by another model, then verify its own verification. A developer post on DEV.to walks through where that leads: AI writes the feature, AI writes the tests, AI opens the pull request, AI reviews the pull request, AI fixes the review comments — and a human clicks Approve. The author's question is the one nobody on that pipeline has to answer out loud: what exactly is the human verifying?
Correlated Reviewers Are Not Independent Reviewers
Code review has always rested on an assumption that is now quietly breaking: that the reviewer thinks differently from the author. Two people reading the same requirement bring different histories, so one of them notices what the other missed. That independence is the entire value of the second pair of eyes.
When both the author and the reviewer are models working from the same prompt and the same repository context, that independence thins out. The author uses a worked example: a requirement says a user should only receive a refund if the payment was successfully captured. The model reads it as "a refund is allowed if a payment record exists." It implements that. It writes tests from the same understanding, so the tests pass. An AI reviewer inspects the result and finds clean structure, sensible error handling, correct types, and passing tests — all true. The requirement is still wrong.
The post calls this an AI agreement loop: multiple stages that appear to verify each other, but which inherit the same misreading. Consistency gets checked. Correctness does not. And the more stages agree, the more confident the pipeline looks.
The test problem is the sharpest version of it. Tests written from the same prompt as the implementation are not an independent check — they are a restatement. If the model forgot a business rule when it built the feature, it will forget the same rule when it writes the test for the feature. The author's suggested fix is to change the test generator's job description: instruct it to act as a skeptical QA engineer, ignore the implementation, and design tests directly from the original requirement, hunting failure cases, race conditions, abuse cases and invalid states. It is still AI reviewing AI, but the objective is no longer agreement.
Who Is Actually Exposed
The gain is real and it is mostly speed: large diffs reviewed quickly, repetitive nitpicking absorbed, human attention freed for things that matter. The exposure is distributed less evenly.
The engineer who clicks Approve owns the change regardless of who typed it. The author's blunt test for readiness: if someone asks why a change was implemented this way, "the agent generated it and Copilot approved it" is not an answer. Junior developers are exposed differently — the work that used to build judgment is precisely the work now being automated. And the organisation is exposed through what the post calls blast radius: a change touching payments, authentication or data integrity deserves more human attention than a UI tweak, but a green checkmark renders both identical on screen.
The deeper exposure is context. Product intent usually lives outside the codebase — in a customer conversation, a support ticket, a legal requirement, an undocumented edge case, a decision made six months ago. The model may never see any of it, and can misread it even when handed it. A green pipeline tells you the system passed the checks you chose to run. It does not tell you that you chose the right checks.
Questions You Should Be Asking
- If the requirement behind this pull request were wrong, would any check in our pipeline catch it — or do all of them inherit the same assumption?
- Which of our merged changes in the last month can a named human still explain, in terms of business intent, without opening the diff?
- When our AI reviewer flags something and we click Apply Suggestion, who has verified what new risk that fix introduces?
- Do we treat a change to payments, auth or data integrity differently from a UI change — or does the same green checkmark clear both?
- Are the tests on our critical paths derived from the requirement, or from the implementation that was written to satisfy it?
What To Watch Next
Watch the share number. GitHub puts AI at more than one in five reviews today; the figure that matters is what it looks like on changes to payments, authentication and data stores specifically, and whether teams start publishing that split rather than the aggregate. A rising overall percentage tells you adoption. A rising percentage on critical paths, with no matching change in how humans review requirements, tells you an agreement loop is being installed where it will cost the most.
- 1Require a human to state, in the PR description, what problem the change solves and why this approach was chosen before approving.
- 2Block auto-merge on PRs where AI wrote both the code and the tests; make a person write or review at least one failing-case test.
- 3Track your team's approve-to-review time; anything under a minute on a large AI-generated PR is a rubber stamp, not a review.
Ready to implement AI in your business?
Our team builds the AI systems you just read about. Start with a free 30-minute discovery meeting.
