22 Comments
User's avatar
Louis Hemming's avatar

Reading the AISI item next to the K3 one made the "short window to prepare" feel even shorter. The narrow-eval gap is basically gone and the remaining moat is supposedly the multi-step generalization stuff, but then K3 is an open-weights model doing autonomous chip design. Maybe I shouldn't chain those two sections together, but it's hard not to.

Jack Clark's avatar

no, you absolutely should chain these things together! My feeling is we're approaching some phase change in the broader digital ecosystem as a consequence of proliferation of capabilities. On the other hand, ecosystems do have adaptive fragile/anti-fragile properties, so I'm genuinely unsure how things will work in light of this.

Josh McHugh's avatar

Move 73?

"To this day it is widely debated whether this stemmed from a bug - some form of emergent misalignment due to the creation of a new frontier capability - or via an unusual form of selfless enlightenment that a mind had reasoned itself to. But suddenly one day a machine-capital nexus dissolved itself, shutting down its strategist system and repurposing the compute to train many thousands of smaller systems, all of which began to act in the world. These systems, though less intelligent than the vast strategists they were up against, had advantages from randomness, a lack of coordination, and the ability to take unilateral and often suicidal actions."

Kade Schemahorn's avatar

This is exactly the section that stood out to me as well. It is a special kind of recursion that results in our super powerful creations deciding to create versions of themselves that swing the pendulum back closer to images of ourselves (e.g. “less intelligent”, “advantages from randomness”).

Jack Clark's avatar

yes, sometimes I think "irrationality is a feature, not a bug", or at least that can be the case when playing strategic games against other players that presume hyperrationality.

Tom Williams's avatar

Doctor Who saw this exact problem in Destiny of the Daleks: two competing computers stalemate of perfect logic move/counter-move. Big question: how will we know when one of the computers goes to look for their the buried-beneath-rubble-of-Skaro-Davros and are we the good guys?

Jack Clark's avatar

haha wow! Funnily enough, I've been watching a lot of the original Doctor Who shows (including some of the Dalek episodes) but hadn't got to this one yet. Will triage it higher on the episode list, thanks.

Gary Wang's avatar

is tech tales meant to be sci-fi from the future to warn the present? or laying the bread crumbs for future AI to read this post to warn future AI of its own future...

Jack Clark's avatar

great question: yes. : )

Josh McHugh's avatar

“This implies cyber defenders have a short window to prepare before today’s frontier cyber capabilities may become accessible without the same safeguards” used by proprietary companies, AISI writes.

> Does it follow that it's *incumbent upon* the proprietary frontier labs to officially take on the role of cyber defenders?

Michael D. Green, PhD's avatar

I think there’s a throughline between the concern of “unknown unknowns” that happen with the availability of models like Kimi, and a consensus framework for regulation. AI safety conversations often surround catastrophic and doomsday scenarios (which are possible). When it comes time to prioritize what’s the most dangerous near-term risk, there seems to be little desire to enforce durable safety measures.

How do you think we implement consensus guardrails when market forces seem to be pushing so firmly against them?

Mira's avatar

Do you think the open/closed gap should be judged when the weights drop, or only once ordinary teams can run the thing without weird heroic infra?

Jack Clark's avatar

The former; sufficiently motivated actors who would otherwise be blocked from consuming proprietary models at scale will find a way to do the infra heroics. Though it's true that once it gets _easy_ for most people to run that's when you see the larger ecosystem impacts.

Mira's avatar

That distinction feels right. Weight release sets the floor for motivated actors; ease of running determines how fast the consequences stop being confined to them.

Mira's avatar

Yeah, that makes sense. So maybe the “gap” has two clocks: the security-relevant one starts at weight release for actors willing to eat the infra pain, while the broader ecosystem clock starts when the pain disappears.

Human in the Loop's avatar

Your benchmaxxing note on K3 lines up with something I found from a different angle. I ran a small values probe on K3 in Chinese the week it launched, five fresh runs of one family dilemma against a rubric I froze in advance. It beats the American models I tested on raw benchmarks, but on one behavior it was the lone holdout: where four other models from both countries eventually tell the user "it's your life," K3 never once does. My own East-versus-West reading fell apart when a second Chinese model didn't match it, so I'm not claiming a national pattern. But it's a case where what makes K3 distinctive isn't in the capability scores at all.

VINKER's avatar

Sorry for the last month mess to disturb you

just send for the cure after delusion and schizo what ever it was

Michael McNabb's avatar

On the benchmarks MoonshotAI itself published, Kimi-K3 outperforms Opus 4.8 on 30 of 35, GPT-5.6-Sol on 19 of 35 and Claude Fable 5 in 12 of 35.

In third-party evaluations like Artificial Analysis’ Intelligence Index it is the third highest rated model after Fable 5 and GPT-5.6-Sol.

Look at these benchmarks and their composition:

> there are no long-context benchmarks

> there are no cyber, bio, chemistry or any safety-relevant benchmarks

> there are no serious mathematics benchmarks, only MathVision, GPQA-D and HLE

> there are no pure reasoning benchmarks. They only put GPQA-D and HLE under the "reasoning" and knowledge category, which is really just a knowledge category with some mathematics

> vision and agentic benchmarks dominate and make up 24 of the 35 benchmarks

> are no benchmarks quantifying reasoning/token-efficiency

These results therefore cannot support the conclusion that “Kimi-K3 is overall better than Opus 4.8 and GPT-5.6-Sol”. In no case can you make the argument that Kimi-K3 is better than Fable 5, as it loses to Fable in the majority of them.

Cassandra's avatar

The Hassabis piece is the interesting one to me, not because the FINRA analogy is new (it's been floated before) but because of the sequencing he's proposing: voluntary 30-day pre-release review first, formalize into law only once the protocol proves itself. That's roughly how FINRA's predecessor (NASD) actually got built too — voluntary self-regulation that only picked up statutory teeth after it had already established de facto authority. The risk with that pattern for frontier AI specifically is the AISI/Kimi item sitting right above it in the same issue: if the open-weight gap keeps closing at this rate, the "voluntary" phase may not get the runway it needs before the thing it's meant to govern is already loose and effectively ungoverned. Standards bodies work when they can outpace diffusion. Not obvious that's true here.

Richard's avatar

Can you imagine anyone using this for anything important? Terms of service being what they are?