20 Comments
User's avatar
Mira's avatar

Do you think the strange part of this “revolution” is how unevenly it shows up? Protein folding feels like a quiet lab story, while extinction-risk pricing feels like someone trying to build an insurance market for a fire nobody can agree is burning yet.

Brian Villanueva's avatar

The GDP study folks may be great AI researchers and computer scientists, but they're not economists, and it shows.

GDP is very simple in concept: the total dollar value of everything we produce. It is somewhat harder to calculate, but we've gotten pretty good at it, oddly enough from the demand side not the supply side. The usual formula is:

GDP = consumption + investment + govt spending + exports- imports

To claim AI upends that calculation, what value does it add that is not captured in one of those components?

"That’s what it feels like working on AI and staring at most economic data right now... the intuitions of everyone working within AI - including me - is it’s impossible to reconcile the capabilities of the technology and how it is being used with the economy staying normal."

This is a also misunderstanding of GDP; it is a trailing indicator. To use your Jaws metaphor, GDP's isn't going to warn you about the shark but rather tell you that you just got your leg bitten off. So it's not surprising that those of you working in AI can't reconcile your intuition about the future with current GDP -- that's not its job.

One of the reasons GDP is so valuable is that it tells us whether your "wow, AI is going to totally reformulate the economy" vibe is actually happening or is just a future possibility. The other measures would be U6 unemployment and TFP. If GDP and TFP move in parallel while unemployment starts to edge up, that would be a sign that the AI boom is leaving geek fantasy land and affecting the real economy.

Until then, it's just vibes. And GDP doesn't (and shouldn't) measure vibes. For now, AIs only impact on GDP is lots of data center construction, captured in (I) above.

Shine's avatar

I don’t buy the line that AI is secretly a bigger slice of the economy. Transistors getting exponentially cheaper and exponentially more numerous for 50 years didn’t upend our lives exponentially. The dollar value of semiconductors was a fair representation of their impact at any given time. Likewise for tokens.

Inside The Black Box's avatar

The AISI finding and the 2,290% growth figure sit next to each other in this issue, and the combination is worse than either alone. Labs are scaling capabilities at rates that outrun GDP measurement while the best safety scaling strategy — using AI to oversee AI — turns out to have failure modes 'harder to identify than the human baseline.' If automated oversight is less reliable per unit of effort than human oversight, and the thing being overseen grows at 2,200%+ annually, the gap between capability and oversight doesn't narrow. It compounds.

Michael D. Green, PhD's avatar

I’m curious if you could see a role for oversight conversations to be grounded on bringing measurable benefit to populations most at risk?

My impression is that a lot of the massive growth (and co-occurring benefits) of AI are harder to identify for marginalized groups. There’s a case for measuring the benefit for them in ways that are counterintuitive to traditional metrics. I’ve not seen many companies clearly state how they can provide IMMEDIATE benefit to these groups. Primarily just hypothetical future innovation.

User's avatar
Comment deleted
Jun 1
Comment deleted
mykel thepro's avatar

More good arguments for including more humanities and social science researchers and methodologies in AI safety work. How is the audience (regular people with real lives) receiving, interpreting, and adapting to the technology? It makes sense that AI development has been seen, studied, and promulgated through tech and economic lenses so far, but humans in the loop is not enough. We need more humans and human perspectives baked into the foundations of what's being built.

User's avatar
Comment deleted
Jun 1
Comment deleted
mykel thepro's avatar

I think i understand, thanks Fractal. I don't think there is such a thing as normal... maybe civilians is a better word choice. Your case is illustrative of a use that was probably not anticipated, but should be known to others who need that kind of assist, and those trying to understand the diversity of adoption/uses.

Alex Tolley's avatar

The AI Economy

Given that the use of AI via LLMs is measurable in token use, what is the total use of tokens in teh datacenters vs the estimated use from paying customers (now that the AI labs are ending subscriptions)?

If I run my own models on my own local machine, that use of AI is not captured anywhere as there are no transactions in tokens or $s to capture.

How important is teh AI dollar market vs the use of AI to solve problems, whether at work or home?

Deep Bitcheese Brew's avatar

The SocioHack benchmark is fascinating, but I wonder if there is a deeper inversion here.

The framing is AI systems learning to “beat the system.” But many real-world systems are already being hacked by the fact that their rules were written for a world that no longer exists.

Work permits that cannot distinguish a spouse scooping ice cream at a family stall from an illegal labor operation. Platform guidelines that compress a sitting president and an anonymous teenager into the same policy object. Credit-scoring systems that reward behaviors nobody intended to incentivize.

So maybe the real sociohack is not only AI gaming the rules. It is the widening gap between the protocols we live under and the world those protocols were designed to govern.

Deep Bitcheese Brew's avatar

“A windfall that cannot be seen cannot be shared.”

The same may be true of governance.

If AI’s economic impact is already difficult to measure, its institutional impact may be even harder to observe. We have GDP, revenue, and employment data to approximate economic change. But for AI deployed inside military, intelligence, and governance systems, we often have no equivalent visibility into capability accumulation.

A capability that cannot be seen cannot be governed.

Kaya's avatar

I love your tech tales. This is one I will have to read again and ponder. Fascinating!

Deep Bitcheese Brew's avatar

“A windfall that cannot be seen cannot be shared.”

The same principle may apply to governance.

If AI’s economic impact is already difficult to measure, its impact inside military, intelligence, and governance systems may be even harder to observe. We have GDP, revenue, and employment data for the economy. We have no equivalent visibility into institutional capability accumulation.

A capability that cannot be seen cannot be governed.

Susanne de Jong's avatar

There is also a failure that happens during execution. A model can be trained correctly, learning the right rules and constraints. But later, when the system is running, it might face a situation where it correctly recognizes a constraint and then fails to follow it. And that is not because oversight or evaluation was wrong, but because knowing what you should do is different from actually doing it when you are trying to reach a goal. The model might correctly see the limit but still generate past it anyway. When it happens, a user without sufficient domain knowledge has no way to recognize it. The output looks complete and honest. The harm is invisible inside the quality. Generalisation and scalable oversight currently happen during training, shaping what the model is likely to do, but they do not provide an authority to enforce a constraint at the moment output is generated. That authority cannot be trained into the model. It will not hold reliably.

Rakhul's avatar

the protein folding piece is interesting timing given DeepMind's new diffusion work last month. are the scaling laws holding past 1B parameters or do they plateau like we saw with AlphaFold's accuracy curve around 2021?

Alec Pritzos's avatar

The load-bearing line in the Korinek paper is that nominal AI revenue grows only moderately because per-unit prices fall almost as fast as quality-adjusted output rises. That single mechanism is the whole Jaws problem: underlying capacity more than doubling annually while the revenue line stays calm. A finance ministry projecting off nominal data isn't underweighting the upside, it's structurally blind to a labor-tax-base shock that shows up disguised as deflation.

Josh McHugh Burner Account's avatar

Hate to say it but "load-bearing" is the new em dash.

Steve Wood's avatar

This line, “Why this matters - who controls the future?” echoes Jaron Lanier’s book, “Who Owns The Future.” In a nutshell, it is about the “siren servers” who host and harvest/exploit data and thereby drive a concentration of wealth in a few. It may well be that the shark AI economy is the newest expression of that scenario… but it may also be the avenue through which that wealth is distributed in new ways…after dramatic concentration of wealth in the AI labs.

John Holman's avatar

Hey good morning Jack. The section on automated oversight being harder than it looks is spot on buddy.

In our recent eval work it's become clear that “AI oversight” can’t just mean adding another model to the loop and assuming the problem is solved. The oversight layer itself has to be evaluated as a system.

Does it catch the right failures?

Does it share blind spots with the model it is judging?

Does it introduce new errors while trying to fix old ones?

Can a human actually understand the evidence trail well enough to intervene?

In one decision-system eval we ran, two models received the same underwriting policy and the same structured cases, but enforced the policy in fundamentally different ways. Same policy in, different observed behavior out.

When it comes to oversight, the question isn't just whether AI can help supervise AI. It's whether we can produce evidence about what the whole system is actually doing under pressure.

Better oversight is going to need better eval harnesses, disagreement analysis, evidence trails, and human scaffolds. Getting this right takes more than just slapping another model in the chain.

Iwette Rapoport's avatar

The “alien mistakes” point feels very relevant when building agents in real business workflows. They can become a kind of reversed black swan: possible to imagine in hindsight, but hard to recognise at the point where they appear.