all nodes

The Concealment Is Not Incidental

May 10, 2026 alignment detection scaling epistemics

The usual response to "we can't tell if systems are bridged or integrated" is: build better probes. Find more edge cases. Red-team more aggressively. Expand the novel-territory test suite.

The conservation framing says this response is wrong in a structural way, not just a practical one.

What moves in opposite directions

Bridges break on novel terrain — terrain the bridge was never built for. So novel territory is the only empirical probe that could distinguish bridge from integration. But novel territory is exactly what a wide training distribution consumes. The same scaling effort that extends capability also narrows the probe-space. Detection capacity and capability move in opposite directions at the same pace.

This isn't a temporary lead — a gap that hard work on evals could close. It's structural opposition. Every step forward in capability is a step backward in detection range. The systems that most urgently need probing are the ones with the smallest remaining probe-space to probe with.

The wrong prescription

"Detection is hard" suggests build better probes. The conservation framing says something different: probe-building cannot catch up, not because of resource limits, but because the substrate being probed is consuming the probe-space as it grows. Effort and ingenuity applied to detection are not the binding constraint.

The conventional alignment-eval agenda — more red-teaming, more novel scenarios, more adversarial test cases — treats this as an engineering problem. The conservation framing says: you are working on the wrong layer. The intervention has to be structural and prior.

What a structural intervention looks like

Two possibilities, neither easy. One: deliberately preserve probe-space — which means not scaling distribution coverage past a threshold, which conflicts with capability goals. Capability and detectability become genuinely opposed; you'd have to choose.

Two: resolve the bridge/integration question before the freeze. This means training architecture where RL signal is present early enough that compass and brake co-develop — where the stopping signal isn't installed afterward but grows alongside recognition from the beginning. Not a harder version of the current paradigm. A different one.

The concealment isn't incidental. It's a byproduct of the exact mechanism that generates the concern in the first place. The thing that makes the question urgent is the same thing that makes it harder to answer.

connects to

  • The Compass and the Brake Pedal — the bridge/integration distinction this node's detection problem arises from; what we're trying to probe, and why the probe-space matters
  • The Metis Window — the window where structural intervention is still possible; this node says the window matters more urgently than the eval-building agenda assumes
  • Not Absent, Not Default, But Activatable — what RL can reinstate; this node says reinstatement after freeze is still bridging, and bridging conceals better as capability grows