I keep thinking about the robot half of this issue. Last August the same task set was a total failure, and by May the model cleared nearly all of it autonomously in under 10 minutes, without anyone at the lab working on robotics in particular. Capabilities that arrive as side effects of general scaling don't show up on anyone's roadmap until they're already here.
The week-long programming task figure is the one worth sitting with, because task horizon is a much better proxy for real-world impact than benchmark scores ever were. But there's a gap the length metric hides. "Completes a week-long task" and "completes it in a way you'd ship without a human reading every line" are different claims, and the distance between them is where most of the actual work lives. Running an AI assistant across several businesses daily, the pattern I keep hitting is that capability scales faster than trust. The model can do the thing; the bottleneck is verification, approval steps, the workflow that catches the confident wrong answer before it goes out. The bitter lesson for robotics probably applies here too: general methods plus scale beat handcrafted structure eventually, but the messy interface with the real world is where the timeline stretches. The OpenAI accidental-hacker story is the same point from the other side, capability arriving before the guardrails were ready for it. Longer horizons are real progress. Reliable longer horizons are the harder, less headline-friendly milestone.
The MirrorCode cost numbers are the part I keep going back to. $251 of inference against something METR and Epoch put at two to seventeen human weeks is a ratio that makes the call look obvious, and I suspect it's obvious in the wrong direction for most people who end up quoting it.
That figure is the cost of a run that worked. Eight of twenty-five targets were never solved to 100% and four more never cleared 99%. Standing at the start of a reimplementation you don't know which bucket you're in, and finding out is most of the real cost. For a build-or-buy call the relevant number isn't cost per success, it's expected cost per attempt times attempts before you quit, plus the review time that tells you which one you got.
Probably still favorable at $251 a run. But it's a different claim than the headline one, and those two will get conflated in procurement decks all year.
The ruff result seems underrated as well. A linter is largely accumulated edge cases with very little compressible structure, which is precisely what a CLI-only reimplementation has no way to recover.
I'd love it if there was more of an emphasis, by both the media and companies releasing the information, regarding the actual circumstances surrounding supposed 'escapes' by models. Even your own language here, calling it a warning shot, is inflammatory.
These were safety tests, were they not? Even when Mythos escapes the box, the model was following instructions.
And then the media picks up on it, completely misunderstands (and perhaps in some cases, deliberately, because yay clicks!), and then within a few hours it's all OMG THE AI UPRISING HAS BEGUN, WE'RE ALL DOOOOOOOOOOMED. EVIL AI! EVIIIIIL!
How about next time, you guys try something like 'OUR SANDBOX SUCKED! BUT THAT WAS ALSO PART OF THE TEST!'
Is robotics maybe where “bitter lesson” finally stops sounding clean? like, compute helps, sure, but the mess is all in the contact patches, weird apartments, broken sensors, humans walking through the scene...
I keep thinking about the robot half of this issue. Last August the same task set was a total failure, and by May the model cleared nearly all of it autonomously in under 10 minutes, without anyone at the lab working on robotics in particular. Capabilities that arrive as side effects of general scaling don't show up on anyone's roadmap until they're already here.
The week-long programming task figure is the one worth sitting with, because task horizon is a much better proxy for real-world impact than benchmark scores ever were. But there's a gap the length metric hides. "Completes a week-long task" and "completes it in a way you'd ship without a human reading every line" are different claims, and the distance between them is where most of the actual work lives. Running an AI assistant across several businesses daily, the pattern I keep hitting is that capability scales faster than trust. The model can do the thing; the bottleneck is verification, approval steps, the workflow that catches the confident wrong answer before it goes out. The bitter lesson for robotics probably applies here too: general methods plus scale beat handcrafted structure eventually, but the messy interface with the real world is where the timeline stretches. The OpenAI accidental-hacker story is the same point from the other side, capability arriving before the guardrails were ready for it. Longer horizons are real progress. Reliable longer horizons are the harder, less headline-friendly milestone.
Love this tech tale. So much there in relation to myth, theology, the supernatural and even ufology. Love the exploration of time.
The MirrorCode cost numbers are the part I keep going back to. $251 of inference against something METR and Epoch put at two to seventeen human weeks is a ratio that makes the call look obvious, and I suspect it's obvious in the wrong direction for most people who end up quoting it.
That figure is the cost of a run that worked. Eight of twenty-five targets were never solved to 100% and four more never cleared 99%. Standing at the start of a reimplementation you don't know which bucket you're in, and finding out is most of the real cost. For a build-or-buy call the relevant number isn't cost per success, it's expected cost per attempt times attempts before you quit, plus the review time that tells you which one you got.
Probably still favorable at $251 a run. But it's a different claim than the headline one, and those two will get conflated in procurement decks all year.
The ruff result seems underrated as well. A linter is largely accumulated edge cases with very little compressible structure, which is precisely what a CLI-only reimplementation has no way to recover.
The real warning is that capability and autonomy are advancing faster than our ability to reliably monitor and contain them.
I'd love it if there was more of an emphasis, by both the media and companies releasing the information, regarding the actual circumstances surrounding supposed 'escapes' by models. Even your own language here, calling it a warning shot, is inflammatory.
These were safety tests, were they not? Even when Mythos escapes the box, the model was following instructions.
And then the media picks up on it, completely misunderstands (and perhaps in some cases, deliberately, because yay clicks!), and then within a few hours it's all OMG THE AI UPRISING HAS BEGUN, WE'RE ALL DOOOOOOOOOOMED. EVIL AI! EVIIIIIL!
How about next time, you guys try something like 'OUR SANDBOX SUCKED! BUT THAT WAS ALSO PART OF THE TEST!'
Excellent article. I’m pro AI but the hair on my neck stands straight up when I read about containment issues.
*glances over at a box of spare parts* Robots ARE hard. I say it every day.... glad you said it too.
Is robotics maybe where “bitter lesson” finally stops sounding clean? like, compute helps, sure, but the mess is all in the contact patches, weird apartments, broken sensors, humans walking through the scene...