Engineering is the only function you can see
Eng has the metrics, so eng gets the credit. The judgment about what's worth building stays invisible.
Eng has the metrics, so eng gets the credit. The judgment about what's worth building stays invisible.
Engineering is the most measured function in most companies, and it is not close. Commits, deploys, cycle time, incident counts, review latency, and DORA, the four standard delivery metrics. There is a dashboard for almost everything an engineer does. You can see the work, so you can manage it, reward it, and settle disagreements about it with numbers.
The trouble is that visible gets mistaken for important. The function you can see becomes the function that gets credit, and the functions you cannot see drop out of the story of how the product got good. Something ships and works. The metrics point at engineering, because engineering is where the metrics are. The decision about what to build left no dashboard behind, so it disappears from the account.
Engineering is the most measured function in most companies, and it is not close.
This was a distortion before agents. It is a serious one now. Building got cheap, so the engineering numbers inflate. More ships, and it ships faster, and the charts look excellent while telling you almost nothing about whether the right things got built. The judgment that set the product's value, deciding what was worth building and reading whether it landed, leaves no data behind. You are measuring the part that got easy and ignoring the part that got decisive, and the dashboard cannot tell the difference.
So teams optimise what they can see. They push throughput because throughput is measurable. They under-invest in judgment because it is not. They end up with a well-instrumented machine building the wrong things efficiently.
More engineering metrics will not fix this. Measure the other half instead. It can be done, and it is less convenient. Start with the read. Shipping is the baseline now, and whether it landed is the part nobody measured. Every launch should carry a verdict drawn from real numbers: landed, watch, or stalled. Not a launch-day feeling. Not a slide. A standing verdict that updates as the numbers come in, and that one person is willing to be wrong about in public. That single move puts a scoreboard next to the decision, where there was none before.
Then attach a name. Every launch traces back to a call, and every call has one owner. One decision, one owner. When the verdict arrives, it belongs to the person who made the call, not to the team that executed it. That is what makes judgment reviewable: a record of calls and how they read out, rather than a score for taste in the abstract.
The read is only as strong as the numbers under it. If those numbers live in four systems with four definitions of the same word, the verdict is a matter of opinion again and the loudest reading wins. One owned source of truth, audited, with every number traceable back to where it came from, turns "I think it landed" into something a room can act on. Agents make this newly practical. An agent can pull the numbers, assemble the evidence, and check a claim against its source. It does not deliver the verdict. A person does that, after looking at what the agent assembled.
None of this makes judgment as clean to measure as deploy frequency, and it should not. The point is to stop letting what is easy to measure decide what gets valued. What made the product good rarely leaves a trace in the engineering dashboard. Manage as though only the traceable things happened, and you will starve the part that mattered most while your charts go up and to the right.