10 Comments
User's avatar
Josh McHugh's avatar

Move 73?

"To this day it is widely debated whether this stemmed from a bug - some form of emergent misalignment due to the creation of a new frontier capability - or via an unusual form of selfless enlightenment that a mind had reasoned itself to. But suddenly one day a machine-capital nexus dissolved itself, shutting down its strategist system and repurposing the compute to train many thousands of smaller systems, all of which began to act in the world. These systems, though less intelligent than the vast strategists they were up against, had advantages from randomness, a lack of coordination, and the ability to take unilateral and often suicidal actions."

Kade Schemahorn's avatar

This is exactly the section that stood out to me as well. It is a special kind of recursion that results in our super powerful creations deciding to create versions of themselves that swing the pendulum back closer to images of ourselves (e.g. “less intelligent”, “advantages from randomness”).

Louis Hemming's avatar

Reading the AISI item next to the K3 one made the "short window to prepare" feel even shorter. The narrow-eval gap is basically gone and the remaining moat is supposedly the multi-step generalization stuff, but then K3 is an open-weights model doing autonomous chip design. Maybe I shouldn't chain those two sections together, but it's hard not to.

Tom Williams's avatar

Doctor Who saw this exact problem in Destiny of the Daleks: two competing computers stalemate of perfect logic move/counter-move. Big question: how will we know when one of the computers goes to look for their the buried-beneath-rubble-of-Skaro-Davros and are we the good guys?

Josh McHugh's avatar

“This implies cyber defenders have a short window to prepare before today’s frontier cyber capabilities may become accessible without the same safeguards” used by proprietary companies, AISI writes.

> Does it follow that it's *incumbent upon* the proprietary frontier labs to officially take on the role of cyber defenders?

Michael D. Green, PhD's avatar

I think there’s a throughline between the concern of “unknown unknowns” that happen with the availability of models like Kimi, and a consensus framework for regulation. AI safety conversations often surround catastrophic and doomsday scenarios (which are possible). When it comes time to prioritize what’s the most dangerous near-term risk, there seems to be little desire to enforce durable safety measures.

How do you think we implement consensus guardrails when market forces seem to be pushing so firmly against them?

Mira's avatar

Do you think the open/closed gap should be judged when the weights drop, or only once ordinary teams can run the thing without weird heroic infra?

Cassandra's avatar

The Hassabis piece is the interesting one to me, not because the FINRA analogy is new (it's been floated before) but because of the sequencing he's proposing: voluntary 30-day pre-release review first, formalize into law only once the protocol proves itself. That's roughly how FINRA's predecessor (NASD) actually got built too — voluntary self-regulation that only picked up statutory teeth after it had already established de facto authority. The risk with that pattern for frontier AI specifically is the AISI/Kimi item sitting right above it in the same issue: if the open-weight gap keeps closing at this rate, the "voluntary" phase may not get the runway it needs before the thing it's meant to govern is already loose and effectively ungoverned. Standards bodies work when they can outpace diffusion. Not obvious that's true here.

Richard's avatar

Can you imagine anyone using this for anything important? Terms of service being what they are?