The main problem is that there exists a type of management, that does not want to understand the meat of the process they are managing, but just wants to compare numbers.
The problem is that certain aspects of human culture, certain aspects of engineering quality are hard to put into numbers.
When we talk about writing a book, a simple number would be pages-written-per-day. Not that this isn't a useless metric. But the actual metric you want would be more something like good-pages-written-per-day, because throwing garbage out quickly is easy, the hard bit is doing good writing. But then the question is, what makes a good page? And many managers don't have the understanding to make that judgement. Quite frankly, if you're a manager and you cannot judge the quality of the product you're working on everything becomes hit and miss like throwing pudding onto the wall and seeing what sticks.
What’s the alternative to this? Without having a fallback to human driven judgement that itself can be gamed.
Maybe you’d suggest that the VP must ask all their direct reports to verify it manually. Then you rely on each person below you to have good judgement and also act in good faith. The director asks the managers who asks the leads who may or may not give accurate reports.
This isn't "don't measure things" it's, "understand that when you measure something, it often turns into a goal. This can have unexpected consequences".
Even with human judgement this happens. Yhrtr are jokes about payment by line of code that have existed for decades.
The main thing is we need to measure secondary things. User satisfaction, defect rate, bugfix rate, etc (and this isn't to say that those are good measures that work everywhere. They may or may not.)
At the end of the day, the challenge is to think, and not assume they a number means what you think it means, or more or fewer of something will always be good.
No yes, I agree with you but frankly I'm tired of the cliche that you can't measure PRs or LOC when in reality it is very much correlated with whatever you want to optimise. I do think it can be gamed, but relying on vibes and human judgement can also be gamed (which is what I tried to point out).
You are completely right that the VP can measure user satisfaction but here's the thing: that feedback loop has a much longer time period. You could also measure your company by revenue or its stock valuation. But the point is to have metrics that have a shorter latency. How can you achieve it? Its a hard problem to solve.
Yes, whatever you want to optimise is a good way to phrase this. However, it'd not just what you want to optimise, you will get side on effects. Even when things seem reasonable, longer term things come out. I don't have a good solution for this. But training people to target one (or even a few) things is incredibly difficult.
I also question the trust aspect. Part of this (which can also be gamed of course) is the level of trust you have that a team is doing their best to work towards some goal. More metrics indicate lower trust in some ways.
I think specifically this (metrics vs judgement), gaming happens in both, but metrics lead to much more wild problems if gamed, because they have a much harder "you said do this more so I did" backstop than "well this seemed to be what you wanted".
Thinking about it a bit more, if you make the argument that human judgement here is itself a compound metric of many different inputs, this semi resolves itself. But we still have the problem that the metric has no backing that can be handed around other than "seems good to me", which changes for many people.
The problem is that certain aspects of human culture, certain aspects of engineering quality are hard to put into numbers.
When we talk about writing a book, a simple number would be pages-written-per-day. Not that this isn't a useless metric. But the actual metric you want would be more something like good-pages-written-per-day, because throwing garbage out quickly is easy, the hard bit is doing good writing. But then the question is, what makes a good page? And many managers don't have the understanding to make that judgement. Quite frankly, if you're a manager and you cannot judge the quality of the product you're working on everything becomes hit and miss like throwing pudding onto the wall and seeing what sticks.
Maybe you’d suggest that the VP must ask all their direct reports to verify it manually. Then you rely on each person below you to have good judgement and also act in good faith. The director asks the managers who asks the leads who may or may not give accurate reports.
It’s not clear that’s any better?
Even with human judgement this happens. Yhrtr are jokes about payment by line of code that have existed for decades.
The main thing is we need to measure secondary things. User satisfaction, defect rate, bugfix rate, etc (and this isn't to say that those are good measures that work everywhere. They may or may not.)
At the end of the day, the challenge is to think, and not assume they a number means what you think it means, or more or fewer of something will always be good.
You are completely right that the VP can measure user satisfaction but here's the thing: that feedback loop has a much longer time period. You could also measure your company by revenue or its stock valuation. But the point is to have metrics that have a shorter latency. How can you achieve it? Its a hard problem to solve.
I also question the trust aspect. Part of this (which can also be gamed of course) is the level of trust you have that a team is doing their best to work towards some goal. More metrics indicate lower trust in some ways.
I think specifically this (metrics vs judgement), gaming happens in both, but metrics lead to much more wild problems if gamed, because they have a much harder "you said do this more so I did" backstop than "well this seemed to be what you wanted".
Thinking about it a bit more, if you make the argument that human judgement here is itself a compound metric of many different inputs, this semi resolves itself. But we still have the problem that the metric has no backing that can be handed around other than "seems good to me", which changes for many people.