The evals-as-OKRs framing is the sharpest point here. Most orgs hand everyone an agent harness before anyone defines what good output looks like which is exactly how you get the loops problem.
Ironic that the fix is the oldest management skill there is: writing down expectations clearly enough that someone (or something) else can execute them. Managers who can do that just became the most leveraged people in the building.
I still find though that evals can be quite disconnected from the reality. And if you use LLM-as-a-judge, you basically doing what the articles describes as spending tokens on spending tokens.
Agreed on the parallel, mismanagement is mismanagement whether it's people or tokens, and evals-as-OKRs is a sharp way to put it.
One thing worth adding to the context hoarding point: it's not just tribal knowledge people are protecting. It's the last visible proof of what they know. Once that gets encoded into an eval, the person doesn't just lose leverage, they lose the thing that made their value legible to anyone else. That's a harder ask than "share your secret sauce."
You’ve hit on so much of what’s been on my mind. The middle management parallels are uncanny and companies are entirely missing it. If the next big challenge is change management, what’s your thesis on how to crack that? I’d think human led consulting but the need is at scale…
Really enjoyed this piece. I think you're identifying the next management problem almost perfectly.
One thought I'd add is that AI may force us to rethink not just management, but economics itself.
Today we still measure organizations through labor, software, headcount, and now tokens. But those are all inputs. The scarce resource isn't labor anymore—it's an organization's ability to continuously resolve uncertainty.
I've been exploring this as Resolution Economics: the idea that the true productive capacity of an organization is its resolution capacity—how efficiently it converts uncertainty into coordinated action and completed outcomes.
From that perspective, many of the "loops" you describe aren't prompting failures. They're manifestations of unresolved organizational constraints. Better prompts help, but they don't eliminate the underlying bottlenecks. What scales isn't token spend; it's the organization's ability to identify, prioritize, and remove constraints across humans, software, and autonomous agents.
That also changes the role of management. Management becomes less about supervising workers—human or AI—and more about continuously improving the organization's constraint architecture.
I suspect the next generation of enterprise operating systems won't primarily manage AI. They'll manage organizational resolution.
The big question is whether AI will ever be cheaper than humans only for it to follow the typical Disrupter blue print that raises prices after the competition (humans) is irrelevant.
Great writing. However, I feel Ramp’s data would accurately reflect only a very specific subsection of corporate America. Would I be wrong to presume this? Does Ramp have a diverse enough customer base to be able to make AI impact conclusions for the average American? Thanks!
Appreciated the eval-writing frame. What I've been thinking about is what the specific forcing function you name (10X-per-quarter token cost) selects for. The McCallum management fix that produced railroad coordination infrastructure required regulatory and commercial pressure alongside safety concerns; cost pressure alone would have produced cheaper trains and worse safety. The token-cost forcing function looks like it produces a similar narrowing (cheaper models, more efficient prompts, LLM-as-a-judge evaluating LLM outputs), which produces a different management fix than the eval-writing infrastructure you describe.
Thank you for this piece, it sparked this thought for me:
The deeper shift may not be that humans are becoming cheaper than software. It’s that judgment is becoming more valuable than execution.
For decades, repetitive work wasn’t valuable because of the output. It was valuable because it developed judgment, built trust, and prepared people for greater responsibility.
If AI increasingly performs that execution, organizations can’t assume judgment will simply emerge. Leadership has to intentionally redesign capability formation, mentorship, and graduated accountability into AI-native workflows.
The competitive advantage won’t be who deploys the most AI. It will be who develops the best human judgment alongside it.
Couldn’t agree more—only that it’s not the first time humans are less expensive than software. This has been true in emerging markets for a long time.
I had a similar observation when I was working in Southeast Asia and saw US tech companies pitching SaaS products to SEA clients. It’s crazy how quickly AI pushed us to this point.
Focusing on the skill with which humans prompt AI models is so 2025. Multi-agent task breakdowns powered by models with high-level reasoning ability is what will drive productivity this year and in the years ahead.
Transformation companies are more appealing as an investment now because they should be able to scale beyond billable hours. But if optimal use of artificial intelligence is a differentiator worth having, why not buy the laggard with all of its legacy data and capture all the benefit of the new scale-up of productivity and competitiveness?
The evals-as-OKRs framing is the sharpest point here. Most orgs hand everyone an agent harness before anyone defines what good output looks like which is exactly how you get the loops problem.
Ironic that the fix is the oldest management skill there is: writing down expectations clearly enough that someone (or something) else can execute them. Managers who can do that just became the most leveraged people in the building.
Writing is the scarcest resource
I still find though that evals can be quite disconnected from the reality. And if you use LLM-as-a-judge, you basically doing what the articles describes as spending tokens on spending tokens.
Great post, made me think in a completely new way about tokens vs humans, we need a new Fred Brooks’ book “The Mythical Token Month” for these times!
Ha—stealing this
Agreed on the parallel, mismanagement is mismanagement whether it's people or tokens, and evals-as-OKRs is a sharp way to put it.
One thing worth adding to the context hoarding point: it's not just tribal knowledge people are protecting. It's the last visible proof of what they know. Once that gets encoded into an eval, the person doesn't just lose leverage, they lose the thing that made their value legible to anyone else. That's a harder ask than "share your secret sauce."
Yes--no employee wants to self sabotage.
You’ve hit on so much of what’s been on my mind. The middle management parallels are uncanny and companies are entirely missing it. If the next big challenge is change management, what’s your thesis on how to crack that? I’d think human led consulting but the need is at scale…
Really enjoyed this piece. I think you're identifying the next management problem almost perfectly.
One thought I'd add is that AI may force us to rethink not just management, but economics itself.
Today we still measure organizations through labor, software, headcount, and now tokens. But those are all inputs. The scarce resource isn't labor anymore—it's an organization's ability to continuously resolve uncertainty.
I've been exploring this as Resolution Economics: the idea that the true productive capacity of an organization is its resolution capacity—how efficiently it converts uncertainty into coordinated action and completed outcomes.
From that perspective, many of the "loops" you describe aren't prompting failures. They're manifestations of unresolved organizational constraints. Better prompts help, but they don't eliminate the underlying bottlenecks. What scales isn't token spend; it's the organization's ability to identify, prioritize, and remove constraints across humans, software, and autonomous agents.
That also changes the role of management. Management becomes less about supervising workers—human or AI—and more about continuously improving the organization's constraint architecture.
I suspect the next generation of enterprise operating systems won't primarily manage AI. They'll manage organizational resolution.
Good analogies
The big question is whether AI will ever be cheaper than humans only for it to follow the typical Disrupter blue print that raises prices after the competition (humans) is irrelevant.
Interesting shit. Now every startup will claim they’re a NeoTransformFirm.
™️
This is great
Thank you!
Great writing. However, I feel Ramp’s data would accurately reflect only a very specific subsection of corporate America. Would I be wrong to presume this? Does Ramp have a diverse enough customer base to be able to make AI impact conclusions for the average American? Thanks!
Appreciated the eval-writing frame. What I've been thinking about is what the specific forcing function you name (10X-per-quarter token cost) selects for. The McCallum management fix that produced railroad coordination infrastructure required regulatory and commercial pressure alongside safety concerns; cost pressure alone would have produced cheaper trains and worse safety. The token-cost forcing function looks like it produces a similar narrowing (cheaper models, more efficient prompts, LLM-as-a-judge evaluating LLM outputs), which produces a different management fix than the eval-writing infrastructure you describe.
Wrote about that thread here: https://rlsutter.substack.com/p/cost-pressure-produces-the-wrong
Thank you for this piece, it sparked this thought for me:
The deeper shift may not be that humans are becoming cheaper than software. It’s that judgment is becoming more valuable than execution.
For decades, repetitive work wasn’t valuable because of the output. It was valuable because it developed judgment, built trust, and prepared people for greater responsibility.
If AI increasingly performs that execution, organizations can’t assume judgment will simply emerge. Leadership has to intentionally redesign capability formation, mentorship, and graduated accountability into AI-native workflows.
The competitive advantage won’t be who deploys the most AI. It will be who develops the best human judgment alongside it.
Nice one. Great to connect with you.
I write about System design, tech, how to build things in tech at a deep level, LLMs and How to invest money, please checkout some of my posts -
https://howtosystemdesigneverything.substack.com/
[Bookmark: 150000 Reads Top System Design] Weekly Round Up : https://howtosystemdesigneverything.substack.com/p/bookmark-150000-reads-top-system?r=14q3sp&utm_campaign=post-expanded-share&utm_medium=web
Couldn’t agree more—only that it’s not the first time humans are less expensive than software. This has been true in emerging markets for a long time.
I had a similar observation when I was working in Southeast Asia and saw US tech companies pitching SaaS products to SEA clients. It’s crazy how quickly AI pushed us to this point.
Highly likely this is the most interesting, thought provoking and entertaining thing you read today. Trust me.
God has come to Earth; with a podcast, an agenda … and a personality.
The Godcast: A compelling original work of ‘spiritual fiction’ that examines humanity as it faces its biggest challenges.
https://godcast.substack.com/p/the-godcast-1-s1-e1-part-1?r=18ajrt&utm_medium=ios
Focusing on the skill with which humans prompt AI models is so 2025. Multi-agent task breakdowns powered by models with high-level reasoning ability is what will drive productivity this year and in the years ahead.
Transformation companies are more appealing as an investment now because they should be able to scale beyond billable hours. But if optimal use of artificial intelligence is a differentiator worth having, why not buy the laggard with all of its legacy data and capture all the benefit of the new scale-up of productivity and competitiveness?