Writing code got cheap. Reading it didn’t. And the instincts good reviewers rely on were trained on code that humans struggled to write.
Much of what experienced reviewers know about how to read a diff was learned on code a human struggled to write. When most code was written manually, the effort of producing it acted as a natural throttle on how much reached a reviewer. But more importantly, human-written code carries information about the process behind it: If an engineer is confused, rushed, or struggling with the design, some evidence of that usually makes it into the diff.
Good reviewers learned to treat those things as signals. They didn’t just review logic; they also used traces of the author’s reasoning to decide where the logic deserved more attention.
AI-generated code has changed both sides of the equation: It increases the amount of material that can reach review, and it can remove many of the irregularities reviewers learned to watch out for. That means that today, code can be wrong without any warning signs.
One clear example of the first problem comes from curl, the open-source project behind the curl command-line tool and libcurl library. Its security team reviews vulnerability reports submitted by outside researchers, a process that has the same basic bottleneck as code review: Anyone can generate a report cheaply, but an experienced maintainer still has to determine whether the report is real.
In 2025, curl maintainer Daniel Stenberg described an “explosion in AI slop reports“ as people used AI to generate plausible vulnerability submissions. He called it “death by a thousand slops.” Historically, more than 15% of reports to curl’s bug bounty had resulted in confirmed vulnerabilities, but in 2025, that rate fell below 5%. Reviewing the false reports still consumed the security team’s time, so in January 2026 curl ended the bug bounty it had run since 2019.
Then something interesting happened.
The models (or at least the way researchers used them) got better. In March, curl returned to HackerOne without reinstating its monetary bounty. By April, Stenberg reported that the low-quality “AI slop” problem had largely disappeared, and that the share of submissions that became confirmed vulnerabilities had returned to roughly 15-16%.
But volume never returned to normal; reports arrived at about twice the 2025 rate, which was already more than twice the rate of previous years. The quality problem had improved, but the capacity problem had not.
Here’s what that distinction means today: AI doesn’t need to produce bad code for review to become a bottleneck. If it produces more reasonable code faster, the amount of code that has to be understood can still grow faster than the number of people capable of understanding it.
In other words: Generation scales, but comprehension doesn’t.
The bottleneck moved into review
There’s now enough telemetry to see the same pattern inside engineering organizations.
A 2026 longitudinal study followed 802 developers and 196,212 pull requests at an AI-forward company whose leadership had explicitly set a goal of doubling engineering output. By April 2026, merged pull requests per developer reached 2.09 times the pre-mandate baseline.
Per-reviewer load roughly doubled as well, and the shape of review changed. Automated review overtook human review, while the share of pull requests receiving substantive human feedback fell significantly.
PRs 2.5× bigger
This study separates two claims that are often blurred together:
AI can increase how much code developers produce
That doesn’t imply that an organization has increased its capacity to evaluate that code at the same rate
Other datasets show the same imbalance from different angles.
Faros AI’s 2026 data across 22,000 developers found tasks completed per developer up 33.7%, while median time in PR review increased 441%.
LinearB’s benchmarks, based on 8.1 million pull requests, found that AI-assisted PRs were much larger than unassisted PRs at the 75th percentile: 408 lines versus 157. They also waited about 2.4 times longer to be picked up for review. Fully agentic PRs waited more than five times longer.
The obvious response is to automate more review, which helps with syntax, style, obvious mistakes, duplicated code, and a growing set of static checks. But it doesn’t remove the harder problem: deciding whether a change preserves the properties of the system that matter. That requires a different way of reading.
A different reading order: Blast radius first
Traditional review often starts with a question: What was the author trying to do here?
You read the ticket, read the code in roughly the order it was changed, reconstruct the author’s reasoning, compare it with your own mental model of the system, and look for places where the two diverge.
But that strategy assumes there’s a reasoning path to reconstruct. And with generated code, there may not be.
Instead, start with what must still be true after the change. Once that becomes the question, reading from top to bottom is no longer obviously the right order, because the most expensive failures may not be evenly distributed through a diff. A destructive database operation deserves attention before a formatting helper. A new payment retry deserves attention before an internal transformation. A changed authorization boundary deserves attention before a renamed variable.
From there, read generated code in six passes, ordered by consequence rather than line number.
The order is deliberate: Passes 0 and 1 take only a few minutes combined, and both come before the point where most reviewers would normally start reading the implementation.
In regulated systems, make the invariant executable
At Solstice, we have an even more complex version of this problem because we build software for regulated work.
Consider a promotional email template containing approved safety information. If we ask a model to clean up the template, its changes may be completely reasonable from a software perspective: normalize whitespace, remove duplicated markup, standardize the punctuation, and combine repeated blocks.
But suppose one of the blocks is safety language that has gone through medical, legal, and regulatory review and must remain exactly as approved. The model doesn’t need to make the email look broken to cause a serious problem. It can render perfectly, pass every test, and look like an improvement in the diff, but still alter text that was not permitted to change.
No amount of careful reading catches that reliably. An invariant does: This block must be byte-identical to its approved source. Then, assert that invariant in CI, and if the text changes, the build fails.
The 6-pass protocol as an agent skill
The same visual grammar can be reused beyond code review. Here is the generalized agent skill behind it:
--- name: mermaid-colors description: Apply a consistent semantic color palette when creating or updating Mermaid diagrams. Use for flows, states, and interactions where color helps distinguish inputs, processing, effects, failures, outcomes, and external systems. --- # Mermaid colors Use color to identify a node's role in the story, not to decorate the diagram. Keep the mapping stable across diagrams. Labels and arrows must still make sense without color. ## Palette | Role | Fill | Stroke | Text | Use for | |---|---|---|---|---| | `userNode` | `#dbeafe` | `#2563eb` (1.5px) | `#1e3a8a` | Actor, user action, or incoming input | | `parseNode` | `#e0e7ff` | `#6366f1` (1.5px) | `#312e81` | Processing, derivation, or decision | | `effectNode` | `#fef3c7` | `#d97706` (1.5px) | `#78350f` | Side effect, state change, or escalation | | `bugNode` | `#fee2e2` | `#dc2626` (2px) | `#7f1d1d` | Identified focal fault or failed invariant | | `resultNode` | `#fee2e2` | `#dc2626` (1.5px) | `#7f1d1d` | Negative outcome or visible symptom | | `legacyNode` | `#fee2e2` | `#dc2626` (1.5px, dashed) | `#7f1d1d` | Deprecated or retired path, when relevant | | `okNode` | `#dcfce7` | `#16a34a` (1.5px) | `#14532d` | Confirmed success or desired outcome | | `extNode` | `#f1f5f9` | `#64748b` (1.5px) | `#0f172a` | External system, service, or data store | ## Apply the palette - Assign roles by meaning, not shape or position. An error-handling *step* may be `parseNode`; use red only for an actual fault or negative outcome. Do not style every ordinary output as `resultNode`. - Define only the classes the diagram uses. For flowcharts, use Mermaid's `classDef` once per role and apply it with `class`. Use the same values for state diagrams where supported. For sequence diagrams, use supported `box` or `rect` coloring with these fills converted to `rgb(...)`; `classDef` does not style participant lifelines there. - Reserve the 2px `bugNode` outline for the focal fault when one has been identified. Use `legacyNode` only when the path is actually deprecated; a dashed red border is a status cue, not an instruction to delete it. - For call paths, bug traces, or architecture diagrams of an existing codebase, inspect the source first. Every node that maps to code should show its verified function name and repo-relative file on separate lines; do not replace it with a paraphrase. Use plain-language labels for people and external systems. If a symbol cannot be verified, label it as unknown or omit the node; never invent a name. - Show the real direction of calls or data, with one idea per arrow. If information flows both ways, draw two directed arrows rather than one ambiguous bidirectional arrow. Keep edge labels short, split a crowded flow into focused diagrams when needed, and preserve branches or cycles that matter. Add a legend only when the color meaning would otherwise be unclear.
A default reading order for all code
This isn’t a checklist for “AI code.” It’s a default reading order for all code. But that distinction matters less and less all the time, because authorship is becoming harder to observe from the diff itself. You won’t reliably know which lines were generated, rewritten, completed or merely reformatted by a model.
If you keep only three habits, keep these:
Read the blast radius before the logic. Spend two minutes asking what the change can touch that you can’t undo.
Verify existence, don’t recognize it. The dependency that looks familiar is exactly the one you’ll be tempted not to check.
Write the invariant before you read the diff. Otherwise, you’ll compare the change with the old implementation, which is the wrong reference.
And say what you didn’t read. A review that implies total coverage is a review nobody should trust.
If you're interviewing with us, this is the kind of judgment we look for.
// solstice.eng
All posts →

