Abstract visualization of data fragments dissolving — representing the collapse of artifact-based measurement

The Output Was Never the Work

Every engineering intelligence product is built on the same assumption — and AI has made it impossible to maintain.

The premise was reasonable. For a decade, it was also correct.

An engineer who merged ten pull requests in a week had done more work than one who merged two. Cycle time correlated with delivery capability. Deployment frequency predicted reliability. The entire engineering intelligence industry was built on this: measure the artifacts, understand the team. Count what moves through the pipeline and you understand the pipeline.

The premise was not a mistake. It was a model that fit its era.


The model that fit its era

Every product in the engineering intelligence category is built on the same foundation. The feedstock is developer tool data — commits, pull requests, cycle time, deployment frequency. The output is productivity insights. The theory of improvement is: measure more, understand more, improve more.

This made sense when artifact production required human effort. When a developer merged ten pull requests, that represented real thinking, real scoping, real decisions about what to build and how to build it. The work happened before the PR — in conversations, in design, in twenty minutes spent understanding why a bug existed rather than just patching the symptom. But the PR was a reasonable proxy for all of that. You couldn't see the thinking. You could count the output.

The artifact and the work were correlated. Not equivalent — but close enough that the correlation was useful.

According to Chen et al. (2026) ("Beyond the Commit: Developer Perspectives on Productivity with AI Coding Assistants"), measuring developer productivity with AI coding assistants requires six factors beyond commit and PR counts — including technical expertise growth and work ownership — because short-term output metrics no longer capture whether engineers are developing skill or just accepting suggestions.

Industry benchmarks were built on millions of pull requests. Performance frameworks codified deployment frequency and cycle time as the authoritative measures of engineering health. Teams were evaluated against these numbers. Leaders reported them upward.

The model worked. Until it didn't.


Then AI arrived

A developer using an AI coding assistant can generate ten pull requests in an afternoon. The artifact count is identical to what a highly productive human produced in a week. The dashboard sees the same number. The PR velocity metric ticks up. Every target gets hit.

Both produce a PR. Only one produces good software.

The artifact was always a proxy for effort. AI broke the proxy. The correlation between output count and actual work is gone — not weakened, not degraded, but structurally severed. The effort is no longer visible in the artifact. It moved somewhere else.

The engineering intelligence industry's response has been to add AI metrics. Track code attribution percentages. Measure token costs. Correlate agent traces to commits. Count how much of your code is AI-generated and call it an AI impact score.

Still artifacts. Still not the work. The category responded to the collapse of the proxy by measuring more of the same thing.

One product in this space has a customer case study that illustrates this precisely: "80% of our code is now generated by AI with up to 30% faster issue cycle times." It's presented as evidence that the platform is working. What it actually shows is that the measurement framework has been left behind. Faster cycle times on AI-generated code tells you nothing about whether the engineers are directing the AI well, verifying its outputs, or accumulating debt that will surface six months from now.

The category is using AI-generated velocity as proof that the category still works. It doesn't.


Where the work actually went

The work is a session now. An engineer in 2026 doesn't live in their IDE. They live inside the session — framing problems, evaluating proposals, verifying outputs, deciding what to accept and what to push back on. The quality decisions happen there. The failures originate there. The difference between engineers who are using AI well and those who aren't lives entirely in the session.

Sessions have structure. A 12-turn session where the engineer scoped the problem, challenged the agent's first answer, and ran the edge case is different from a 90-turn sprawl where nothing was verified and the engineer accepted outputs they couldn't explain.

Both produce a pull request. The pull request looks the same. The session does not. Among 3,380 developers studied by Zakharov et al. (2025) ("From Teacher to Colleague: How Coding Experience Shapes Developer Perceptions of AI Tools"), experienced engineers perceive AI as a junior colleague they direct and verify — while less experienced developers treat it as instructional. The same artifact; entirely different sessions producing it.

The artifact count never captured the work. It captured a proxy that, for a decade, moved with the work closely enough to be useful. That proxy is gone. What remains is the session — observable, structured, and measurable, if you're looking at the right thing.


What becomes possible now

Session data makes the invisible visible. The craft your best engineers have developed — scope before you prompt, challenge the first answer, run the edge case the agent didn't think of — is currently locked in their heads, invisible to the org, lost when they leave.

That craft can be made visible. It can be measured against the question that actually matters: whether your engineers are using AI well, and whether they are getting better at it over time.

Standards can exist. Not a vibe — a standard. A quality session for a security-critical change looks different from one for a routine bug fix. For the first time, you can define what that difference looks like, measure it, and build toward it. The right workflow for the right task stops being a matter of individual judgment and becomes something the org knows.

The gap between your AI-native engineers and everyone else doesn't have to widen quietly. It can be coached.

The output was never the work. AI just made that fact impossible to ignore.


Maestro measures the session — the observable interaction between an engineer and an AI agent — where the actual quality decisions happen. See how it works.

Ready to transform your engineering organization?

Start making data-driven decisions about your engineering processes with AI-powered insights.