Discussion about this post

User's avatar
Mitchell Kosowski's avatar

The "ground truth shows up as a phone call to someone's desk a week later" line is the whole ballgame. Benchmarks measure whether the task ran; production measures whether anyone can vouch for it and those are very different problems.

Same shift coding just went through: producing the output stopped being scarce, standing behind it became the bottleneck.

Hector Stanley's avatar

Such a good article, thank you so much Fabrizio Serafini!

I've been having a lot of fun with CU models on my laptop, seeing what tasks they're good at automating when I'm away from my keyboard.

I think it's also worth calling out the models that are currently topping the OSWorld charts (that don't come from Anthropic or OpenAI), such at Holo3 from H Company, and Qwen 3.7. Holo3 sits above 80% on the bench, which is impressive for a small company!

5 more comments...

No posts

Ready for more?