Abstract interwoven threads separating — connection diagram losing its links, knowledge transmission breaking apart

Code Review Used to Be How Craft Propagated

The transmission mechanism that moved senior engineers' judgment through an org was the artifact itself — and AI agents moved the judgment out of the artifact.

For a generation, code review was how senior engineers' judgment propagated through an org — and AI agents moved the judgment out of any artifact code review can see. A senior engineer's reasoning used to become visible in the artifacts they touched: PR review comments, commit messages explaining architectural decisions, code that demonstrated patterns worth following. Other engineers read those artifacts and absorbed the reasoning. The transmission was slow but real.

Then AI agents took over implementation. The reasoning that determines quality moved out of the artifact. The instrument that transmitted knowledge between engineers is still in place. The knowledge is no longer passing through it.


The reasoning and the artifact used to move together

According to Baltes et al. (2018) ("Towards a Theory of Software Development Expertise"), software development expertise is context-specific and cannot be inferred from experience or output volume alone. A survey of 335 developers found that expertise lives in the specific decisions an engineer makes: how they scope a problem, when they push back on a design, which edge cases they anticipate. Transferring that judgment between engineers requires active observation. It does not travel through tenure.

For a long time, code review was that observation instrument. An engineer reviewing another engineer's PR saw the reasoning made visible: why a function was structured this way, what edge case the test was protecting against, which refactor would have been simpler and why it was rejected. The artifact carried enough of the thinking for a reader to absorb the judgment behind it.

The engineering intelligence category built on this assumption. DORA metrics, PR templates, code review requirements, documentation standards — all of it assumed that the work was visible in the output. If the output carried the thinking, measuring and propagating the thinking meant measuring and propagating the output. The assumption was reasonable. For a decade, it was correct.


Then the work moved upstream of the artifact

A session — the back-and-forth interaction between an engineer and an AI agent, from initial prompt through verification and acceptance — is now where the quality-determining decisions happen. The engineer frames the problem, the agent proposes an approach, the engineer either challenges it or accepts it, and after several turns a PR gets opened. The PR looks the same whether the engineer directed carefully or accepted uncritically. The reasoning does not.

Both produce a PR. Only one produces judgment worth learning from.

According to Ulfsnes et al. (2024) ("Transforming Software Development with Generative AI: Empirical Insights on Collaboration and Workflow"), as developers shift from consulting colleagues to consulting AI tools, the peer-learning dynamics that propagated knowledge through agile teams are disrupted. The study of 13 professionals found that efficiency gains at the individual level come at the cost of collective knowledge accumulation. The instrument that moved reasoning between engineers is being replaced by a tool that does not move it.

A senior engineer spends 12 turns scoping the threat model, challenging the agent's first implementation, and verifying edge cases the AI had missed. A junior engineer accepts the first output after two turns. Both produce the same PR. Code review sees both artifacts identically. The senior's judgment — the specific decisions that made the work good — lived in the session, not the output.


What the org cannot see, it cannot keep

This is not a measurement problem. It is a retention problem.

When a senior engineer leaves an organization, the code they wrote stays. In the artifact era, the reasoning behind that code was also partially retained — documented in PR comments, encoded in the patterns their junior colleagues had absorbed through review. The institution kept something even when the individual left.

In the agentic era, that transfer is no longer happening. The reasoning that determines quality lives in sessions no other engineer sees. The patterns a senior engineer has developed for working with AI agents — scope first, challenge first answers, verify edge cases — exist only in their individual practice. When they leave, those patterns leave with them. The remaining team is left with the code and none of the judgment that produced it.

The org is accumulating AI-generated output while losing the institutional capacity that used to make that output good.


Every senior engineer's workflow compounds across the org

Session observability does not replace code review. It replaces what code review used to provide — the surface where an org could see how its best engineers actually worked, and keep some of that even as people came and went.

When sessions are observable, the decisions an experienced engineer makes while working with an AI agent stop being private practice and become institutional record. Which edge cases they run before accepting an output. When they push back on a proposed design. How they scope a security-critical change versus a routine fix. The specific interaction patterns that separate engineers producing reliable AI-assisted work from those producing debt become something the org can study, not just something individual engineers eventually pick up or never do.

The result is not a training program. It is compounding infrastructure. Every session an experienced engineer runs adds to what the org can see, study, and spread. How a senior engineer scopes an ambiguous requirement. Where they push back on an agent's first design. Which verification they never skip. Each of these stops being one person's habit and becomes a resource every other engineer can learn from. Code review used to drip institutional knowledge into the org slowly. Sessions accumulate it continuously.

An AI-native engineer's workflow should not be a personal trade secret that walks out the door when they take another offer. For the first time, it does not have to be. The org gets stronger every time that engineer works.


Maestro measures the session — the observable interaction between an engineer and an AI agent — where the reasoning that determines software quality now lives. See how it works.

Ready to transform your engineering organization?

Start making data-driven decisions about your engineering processes with AI-powered insights.