Discussion about this post

User's avatar
Alec Pritzos's avatar

I keep thinking about the robot half of this issue. Last August the same task set was a total failure, and by May the model cleared nearly all of it autonomously in under 10 minutes, without anyone at the lab working on robotics in particular. Capabilities that arrive as side effects of general scaling don't show up on anyone's roadmap until they're already here.

Chris S - The Next Rung's avatar

The week-long programming task figure is the one worth sitting with, because task horizon is a much better proxy for real-world impact than benchmark scores ever were. But there's a gap the length metric hides. "Completes a week-long task" and "completes it in a way you'd ship without a human reading every line" are different claims, and the distance between them is where most of the actual work lives. Running an AI assistant across several businesses daily, the pattern I keep hitting is that capability scales faster than trust. The model can do the thing; the bottleneck is verification, approval steps, the workflow that catches the confident wrong answer before it goes out. The bitter lesson for robotics probably applies here too: general methods plus scale beat handcrafted structure eventually, but the messy interface with the real world is where the timeline stretches. The OpenAI accidental-hacker story is the same point from the other side, capability arriving before the guardrails were ready for it. Longer horizons are real progress. Reliable longer horizons are the harder, less headline-friendly milestone.

7 more comments...

No posts

Ready for more?