7 Comments
User's avatar
Mitchell Kosowski's avatar

The "ground truth shows up as a phone call to someone's desk a week later" line is the whole ballgame. Benchmarks measure whether the task ran; production measures whether anyone can vouch for it and those are very different problems.

Same shift coding just went through: producing the output stopped being scarce, standing behind it became the bottleneck.

Hector Stanley's avatar

Such a good article, thank you so much Fabrizio Serafini!

I've been having a lot of fun with CU models on my laptop, seeing what tasks they're good at automating when I'm away from my keyboard.

I think it's also worth calling out the models that are currently topping the OSWorld charts (that don't come from Anthropic or OpenAI), such at Holo3 from H Company, and Qwen 3.7. Holo3 sits above 80% on the bench, which is impressive for a small company!

Paresh Yadav's avatar

Based on the POC we ran, this is where the challenge and the opportunity stands, (how) can the agent "learn" the tribal knowledge --> "The hard part is not whether an agent can navigate an SAP screen, it is whether it understands how a particular company actually gets work done: **the tribal knowledge**, the internal terminology, the preferred formats, who to escalate to and when, how to handle failures, how to verify output reliably. "

Cyrus Azamfar's avatar

The "context is the moat, not the model" point matches what we see, but there's a second market hiding under the BPO framing. Every cost comparison here is agent vs. someone already doing the work — offshore or in-house. For small businesses the baseline is usually nobody. The content doesn't get written, the follow-ups don't get sent, the listing never gets updated. There's no incumbent labor line to compare against.

That changes what "reliable" has to mean. When the alternative is a task that never happened, the bar isn't 85% task completion with a verification harness — it's whether the output is good enough to ship and easy enough to correct. Different failure economics entirely.

It's the thing we've been building at Leapd.ai — agents that build your business and then run the growth side of a small business end-to-end rather than clicking through someone's existing back office. The hard part is exactly what you describe: not the navigation, it's knowing how this business actually talks and who it's talking to.

Leon Wildcard's avatar

there's already tens of (funded) startups for various layers of this. yet you can't just go and ask your ai to vibe its way through the whole process of finding them, making an account, integrating into the process. it seems like the natural next step in interface evolution

Brandon Cebulak's avatar

web3 is not dead... just the internet... agree?

Jerry Xue's avatar

Verify is the most important part in the loop. We are building a independent verification layer.