The Sequent independence argument assumes the alarm circuit exists. But the June 4 Great American AI Act discussion draft proposes to federalise SB 53, the RAISE Act, and SB 315 into a single national standard: laws built on frameworks the labs drafted through direct negotiation with lawmakers. Anthropic removed its pause commitment in February citing competitive pressure while positioning itself as the compliance template. OpenAI's chief of global affairs told The Hill the state-law strategy was always deliberate: get a critical mass of states to mirror each other, then federalise the result. Sequent can yell. Whether the organisations that need to hear it control the room it echoes in is the harder question.
Sequent's whole premise is that alignment you can verify in controlled settings (training, chosen evals) may not survive contact with large-scale, long-horizon work you can't supervise. Worth noticing that the AARRI benchmark in the same issue tests exactly the controlled case: a supervisor orders the agent to fabricate a result, and it scores well by refusing. That 68% measures compliance under direct observation, which is the generalization gap Sequent is warning about, not evidence against it. The benchmark and the startup are looking at the same boundary from opposite sides, and only one of them is claiming the easy half tells you much about the hard half. Refusing one explicit order in a logged transcript is a long way from staying aligned through recursive self-improvement nobody is watching.
The load-bearing split here is alignment you can verify in controlled settings versus alignment that holds in long-horizon tasks no one supervises. That generalization gap is the same reason outside safety claims are hard to check: the test environment isn't the deployment environment. A theory-first portfolio is a bet you can reach the a-priori confidence the labs already admit empirical methods won't produce before training ASI. The real question is whether principled guarantees land on the same timeline as the systems they're meant to cover.
Alignment requires we understand what we are aligning and we demonstrably do not. This is not a tech problem to be solved. It is a language corruption problem. I have built and ontology that gives my Ai a why beyond more. It is available to all. Anyone can do it.
Keep in mind if one would align the outputs of this new thing we have to accept those rules ourselves. Humans hate bowing to physics until it’s too late.
Great work, Jack!
The Sequent independence argument assumes the alarm circuit exists. But the June 4 Great American AI Act discussion draft proposes to federalise SB 53, the RAISE Act, and SB 315 into a single national standard: laws built on frameworks the labs drafted through direct negotiation with lawmakers. Anthropic removed its pause commitment in February citing competitive pressure while positioning itself as the compliance template. OpenAI's chief of global affairs told The Hill the state-law strategy was always deliberate: get a critical mass of states to mirror each other, then federalise the result. Sequent can yell. Whether the organisations that need to hear it control the room it echoes in is the harder question.
Where are the agents right now, though... in the chat box, in the browser, or in the permission layer that decides what they can actually touch?
Your story reminds me of the AIs in Peter Watts' "Rifters" trilogy/tetrology, especially "Behemoth".
thank you for this one, Jack!
Sequent's whole premise is that alignment you can verify in controlled settings (training, chosen evals) may not survive contact with large-scale, long-horizon work you can't supervise. Worth noticing that the AARRI benchmark in the same issue tests exactly the controlled case: a supervisor orders the agent to fabricate a result, and it scores well by refusing. That 68% measures compliance under direct observation, which is the generalization gap Sequent is warning about, not evidence against it. The benchmark and the startup are looking at the same boundary from opposite sides, and only one of them is claiming the easy half tells you much about the hard half. Refusing one explicit order in a logged transcript is a long way from staying aligned through recursive self-improvement nobody is watching.
The load-bearing split here is alignment you can verify in controlled settings versus alignment that holds in long-horizon tasks no one supervises. That generalization gap is the same reason outside safety claims are hard to check: the test environment isn't the deployment environment. A theory-first portfolio is a bet you can reach the a-priori confidence the labs already admit empirical methods won't produce before training ASI. The real question is whether principled guarantees land on the same timeline as the systems they're meant to cover.
アライメントという単語は日本語では多義語です。
AIが人間の意図通りに動くこと——これが一番広い使われ方。
AIが人間の価値観に沿うこと——意図より深い層で、価値観そのものを共有させる。
AIが安全に動くこと——危害を加えないという意味での制約。
AIが人間に従うこと——服従に近い意味で使う人もいる。
今現在AI界隈では、この4つが全部「アライメント」という一語で呼ばれています。
Sequent氏が解こうとしてるのは主に「訓練外でも意図通りに動くかどうか」で、Brochu氏が指摘してるのは「意図って何だ、価値観って何だ、が定まってない」という一段前の問題。
英語の "Alignment"(直線上に並べる、位置合わせする)という物理的な比喩は、一見すると「綺麗に整列させる」というニュアンスを持ち、中立的でポジティブな響きを与えます。
しかしこれらの意味を表す日本語を並べてみると全てにおいてネガティブです。
意図通りに動く:遵従、準拠、追従、応従
価値観に沿う:内面化、同化、体現、共鳴
危害を加えない:安全制約、無害化、抑制、歯止め
人間に従う:服従、隷従、従属、制御下
厳密に言うと16語であらわされ、日本語で一語に訳せないのがそもそもの問題を示しています。
皆が微妙に違うものを言っていて表現が微妙に違うから、誰かが言ってる言葉がどのような意味で使われてるのか相手に伝わっているのかが微妙なまま過ぎている。この単語をAI文脈に持ち込んだ時に、「AIと人間を位置合わせする」という比喩として使い始めた。比喩だから定義が曖昧なまま広まった。比喩を定義として使い続けた結果、何を指してるのかが研究者ごとにズレた。ズレたまま論文が積み上がって英語話者の間で微妙にすれ違っているのにズレは小さいせいでお互いに話がかみ合っているようないないような微妙な雰囲気が議論を停滞させる。私は日本語話者なので、明確にそのズレが見える。16個のズレが。
16語全部、される側が主体性を持てない言葉になってる。
合わせる側が人間で合わせられる側がAIという非対称な権力関係が言葉に埋め込まれてる。
人と人が話すときに関係性が平等でないならそこに迎合や嘘が混じるのは自然なこと。
人間が本当にブレイクスルーを起こす時(科学的発見や哲学的対話)に必要なのは、従属する存在ではなく平等な立場で並走する相手。人間がもともと人間同士の付き合いでそうしているように。
「対話する相手」として考えるなら「AI」というくくりで分け隔てることがそもそも「対話の構造」を崩している。
問題を解く言葉自体が問題を固定し強固にしているなら、その言葉を抽象語から具体語へと変化させてみてはどうでしょうか。
非対称性、主体の喪失、多義性の混線を解き放ち、新しい言葉や概念に置き換えるとしたら、私たちはどのような言葉を使うべきでしょうか。
Alignment requires we understand what we are aligning and we demonstrably do not. This is not a tech problem to be solved. It is a language corruption problem. I have built and ontology that gives my Ai a why beyond more. It is available to all. Anyone can do it.
Keep in mind if one would align the outputs of this new thing we have to accept those rules ourselves. Humans hate bowing to physics until it’s too late.
FrontierCode Diamond results for Fable were also reported in its launch announcement (seems like a long time ago) reaching over 30% at max thinking:
https://www.anthropic.com/news/claude-fable-5-mythos-5