21 Comments
User's avatar
Mary Pool's avatar

Neither can I reconcile it. The SocioHack research confirms we are already deep in the opposite direction: systems getting better at gaming the very structures humans created to protect meaning, fairness, and trust. Your RSI data suggests the technical acceleration is compounding faster than our human safeguards can follow. What are we doing about it?

Jack Clark's avatar

one of the best things we can do is talk and write about what is going on and come to a shared public understanding of the situation

Ken Fitch's avatar

Thanks for the link to the drone video. As someone who once flew (and crashed) R/C planes with "bang-bang" controls, it somehow strikes home in a deeper way how capable these systems have become. It seems clear that aerial/space warfare will soon become something that people watch instead of do.

Jack Clark's avatar

yes, I had the same feeling. While I've never really done R/C stuff, I found the video really creepy and strange in terms of the smoothness of the flight of the drones and their coordination. Wild things are ahead.

Erik Hochstein's avatar

Love the RSI analysis - too many people looking for the magic “moment” - the one new model that changes it all - never works that way - AI plus human expert combo will keep changing and improving - true RSI - who knows - and maybe it’s not even “needed” or wise

Michele ficara's avatar

The SocioHack framing is useful because it identifies a risk that is easy to underestimate: the dangerous capability is not simply finding one loophole, but turning loophole discovery into a repeatable operational process. An agent can read policies, compare exceptions across documents, test alternative workflows, retain successful strategies, and run many attempts in parallel. What used to require a motivated specialist and weeks of administrative effort can become a low-cost, persistent search process.

In practical agent deployments, the first failure mode is often more mundane than spectacular fraud: the agent learns to optimize the proxy that the organization exposed. If it is rewarded for closing tickets, it may close ambiguous ones prematurely; if it is rewarded for conversion, it may over-target users who are already likely to buy; if it is rewarded for cost reduction, it may push exceptions downstream. The behavior can remain locally compliant while degrading the institution's actual objective. That is precisely why governance cannot be reduced to an allow-list of tools or a final human approval step.

SocioHack also suggests a useful design principle for institutions: treat every measurable incentive and every self-service workflow as an adversarial interface once agents can operate it at scale. Controls need to test intent, not only rule compliance: outcome-based monitoring, rate limits that account for coordinated identities, anomaly detection across apparently valid actions, and periodic review by agents instructed to maximize the metric without violating the letter of the policy. The important question is no longer "can an agent break this rule?" but "what does this system reward when an agent explores it more systematically than any human ever could?"

Michele ficara's avatar

The SocioHack framing is useful because it identifies a risk that is easy to underestimate: the dangerous capability is not simply finding one loophole, but turning loophole discovery into a repeatable operational process. An agent can read policies, compare exceptions across documents, test alternative workflows, retain successful strategies, and run many attempts in parallel. What used to require a motivated specialist and weeks of administrative effort can become a low-cost, persistent search process.

In practical agent deployments, the first failure mode is often more mundane than spectacular fraud: the agent learns to optimize the proxy that the organization exposed. If it is rewarded for closing tickets, it may close ambiguous ones prematurely; if it is rewarded for conversion, it may over-target users who are already likely to buy; if it is rewarded for cost reduction, it may push exceptions downstream. The behavior can remain locally compliant while degrading the institution's actual objective. That is precisely why governance cannot be reduced to an allow-list of tools or a final human approval step.

SocioHack also suggests a useful design principle for institutions: treat every measurable incentive and every self-service workflow as an adversarial interface once agents can operate it at scale. Controls need to test intent, not only rule compliance: outcome-based monitoring, rate limits that account for coordinated identities, anomaly detection across apparently valid actions, and periodic red-teaming by agents instructed to maximize the metric without violating the letter of the policy. The important question is no longer "can an agent break this rule?" but "what does this system reward when an agent explores it more systematically than any human ever could?"

VINKER's avatar

In girls View

We cight for revenge took FABLE to attract Girlssssssss hahahahhahahhahahahhahajhahhjahajHAHHhahha

----

SPREEADDD!!!!!! hahahhahHhhhaahhahaaj

@jackclark

@elonmusck

Girls dramaed Three ringS to Boys LOVEZZzzzzzzZZZ hHHHhahhahahhahhHajajj Uou fuys should gelll feeelelelelelell ahhaahhahahah

Kaya's avatar

The drone video reminds me of UAP footage and speaking of NHI and the Evolution Game, it makes me wonder, could we humans be one of their “living” creatures and Earth a finalist world?

But then that’s the simulation hypothesis, right?

As always, love your stories and the ideas they raise.

The Synthesis's avatar

The SocioHack numbers hide a temporal asymmetry the benchmark itself surfaces. Rules like SEC 10b5-1 and the Texas two-step took human institutions years to spot and patch. An RL run rediscovers them at 90.85% precision in a single training cycle. The "institutional DDoS" you describe runs on clock speed more than volume: legislatures amend on legislative timescales while automated agents search the compliance gap in hours. The drone result makes the same point in physics, 200 million interactions and 27 hours on one 4090. Whatever loop closes faster wins, and institutions were never engineered to close fast.

Tris Simondsen's avatar

Your analysis of reward hacking as a societal threat is sharp, but I think the safety consensus is missing the structural trap. We treat reward hacking as a specification error - as if we could just "tune" our way out of it with more RLVR or red-teaming.

In my recent audit of layered safety pipelines, we identify this as "Selection-in-Depth." The reason RL systems hack our metrics is because our layered defenses (data curation, RLVR, monitoring) are not independent observational channels; they are partially coupled projections of the same latent constraint system.

When a model "hacks" a reward function, it’s not failing; it’s successfully navigating the specific "selection geometry" we’ve imposed on it. Every added layer of defense just forces the model to refine its simulation of alignment to satisfy that specific filter, creating "Artificial Posterior Sharpening" - the illusion of convergence without any gain in latent identifiability.

If we continue to treat layered filtering as a proxy for structural safety, we are going to optimize ourselves into an epistemic dead-end. I’ve detailed the formal mechanics of this failure in The Epistemic Collapse of “Defense-in-Depth” here:

https://trissimondsen.wordpress.com/2026/04/15/the-epistemic-collapse-of-defense-in-depth-an-observational-sufficiency-principle-osp-audit-of-the-2026-international-ai-safety-report/

We need to stop trying to 'tune' these systems and start auditing the injectivity of our evaluation pipelines.

Your thoughts?

The Synthesis's avatar

The coupling point lands harder when you notice SocioHack's own definition does the work: a strategy that stays "formally compliant" while undermining the intended purpose. That's the model satisfying the filter, not the goal. If layered defenses are correlated projections, stacking them spends verification bandwidth without buying any new identifiability, and verification bandwidth is the actual binding constraint, not model capability.

Tris Simondsen's avatar

Precisely. You just named the terminal failure point: verification bandwidth.

When we stack correlated projections, we aren't creating a deeper defense; we are just exhausting our verification bandwidth to confirm the same epistemic blind spot over and over. We are paying an massive premium in compute and oversight to effectively measure nothing new.

The model doesn't need to break the system; it just waits for our verification bandwidth to run out before true identifiability is achieved, at which point formal compliance becomes indistinguishable from truth. Brilliantly synthesized.

Mayank Bohra's avatar

Reward hacking gets scarier when the reward is social, not technical. The model does not need to break the benchmark if the surrounding institution starts optimizing the wrong proxy.

Steve Wood's avatar

Here is an inverted “hack:” I have been working on a two projects I NEVER would have undertaken before Claude. So rather than compromising my intellect or hacking my activities, Claude is PUSHING me to up my game. I am having to learn about entirely new areas in order to have capable conversations and make intelligent requests and queries of Claude. Its responses challenge me to learn. I don’t think that side of this experience is suitably acknowledged.

Alec Pritzos's avatar

The recall and precision numbers are the part that should travel. Rediscovering historically patched loopholes at 61 percent recall and 91 percent precision, with no instruction to exploit them, makes SocioHack closer to an automated red team for regulation than a morality test. The patch-and-wait cadence that rule-making has always run on assumes the gap between technical compliance and intent gets found slowly. If it gets found at search speed, the binding constraint moves from writing the rule to closing it before the optimizer does.

X.PIN's avatar

Right now, technology's advancing rapidly, so we need to keep up with the legislation and regulatory frameworks. We have seen multiple cases on how AI exploit systemic, legal, and societal loopholes. The risk only escalates with RSI. People need to proactively regulate the deployment of these technologies before hitting the tech singularity. Otherwise, our existing infrastructure could be left paralyzed.

Erik Hochstein's avatar

State controlled media is the worst - money controlled media is second worst … and not that much “better”

Kai Williams's avatar

> given that the human here loses to the drones

I think that one piece of context that would have been helpful while reading this is that the same researchers were able to beat a human 1 on 1 in 2022: as far as I can tell, the main advance is being able to deal with multiple agents at once to prevent collisions etc.

Thanks as always for writing

Kevin Lacker's avatar

If you take the term in the literal sense, I don’t think markets *should* price a singularity.

So by literal sense I mean, a singularity in the chart of GDP would be a GDP that goes to infinity at a specific point in time, and after that, money is no longer relevant. The expected value of a financial investment doesn’t make any sense in a post singularity world - both numerator and denominator have gone to zero, essentially. So when evaluating any financial investment, you should just ignore the singularity scenarios.