1 A 3 O R N

On the Impossibility of Airplane Safety without Flight Instruments

Created: 2026-09-09
Wordcount: 1.7k

“Airplane safety” relies on the fact that airplane cockpits inform pilots about the state of the plane.

For instance, information about airspeed is necessary to prevent the plane from stalling. Other instruments show information about vertical speed, aircraft roll, or the state of the engine, which are necessary both for directly preventing disaster, and for letting pilots steer away from parts of state-space that distantly approach disaster. When these sensors function, airplanes are one of the safest forms of transportation in existence. Were these sensors to break, safety would be impossible.

airplane cockpit

Imagine an absurd world: one where people tried to create safe airplanes, while trying to hide information about the state of the plane from pilots. What if, for instance, airplanes sharply limited information about the state of the engine, to protect the sensitive intellectual property of Pratt & Whitney and GE Aerospace? What if they hid information about the airplane’s change in altitude, because this might allow one to infer the performance properties of the engine? Or imagine a world where, before creating any flight instrument informing pilots of any particular kind of information, a team of engineers and lawyers carefully brainstormed whether this information would tell pilots too much?

It would probably be impossible to make airplanes safe under such circumstances. And it would probably remain impossible even if some alternate-world FAA and NTSB labored mightily to make airplanes safe, while operating beneath this handicap. Their efforts would bear little fruit; the institution of commercial aviation, as we know it, would probably be unable to exist.

In general, it is hard to operate a complex, important system safely without a high-bandwidth feed giving information about the state of the system. And this is true over systems as diverse as planes, corporations, military battalions, drag racers, racing horses, and so on. High-information-density feedback about the state of these systems is foundational to steering them effectively and safely, and it probably could not be otherwise.


Leading AI companies, such as Anthropic and OpenAI, sharply limit information users can receive about the internal state of the AIs they are using. I think it is likely that this harms the ability of users to steer AIs; this also likely harms the ability of users to create tools to help them steer AIs. And I think, projecting such limitations into the future, that such limits will probably injure people’s ability to work with AIs safely.

Note: There are two broad species of information that are useful for operating AI safely. The first is about how AI was trained and created; the second is about an AI’s internal state while running. For the moment, I am concerned only with the second.

The first kind of information is also important. Both the general public and private 3rd parties outside of labs currently lack detailed knowledge about many things inside of labs: the pretraining data, architecture, RL update methods, lists of RL environments, Constitutional pipelines, multi-level optimization hyperparameters, and evaluation schemes. I believe that this lack of detailed knowledge prevents 3rd party work on the science of alignment; prevents informed choices in AI policy; and needlessly turns gigantic quantities of AI discourse away from object-level discussions and towards absurd Kremlinology about lab interiors. I’m happy that proposals like the AI Futures Project “Total Research Transparency” have brought more attention to the existence of this problem, although I’m somewhat uncertain of this solution. But for now, I am not talking about this lack of information.

Instead, I am discussing cases where AI companies limit information about the current state of AIs that users are employing. (Note that hiding chain-of-thought is not the only thing. For instance, Anthropic hides information about how Claude Code works at all.)

By far the most important aspect of current AI’s internal state is chain-of-thought. Anthropic originally showed the chain-of-thought for Sonnet 3.7, noting that this was useful because being “able to observe the way Claude thinks makes it easier to understand and check its answers.” By Claude 4, Anthropic started showing summaries of chain-of-thought in a small percent of cases – but soon this became universal. OpenAI similarly only shows summarized chain-of-thought – although they started from showing even less. So AI companies hide information about chain-of-thought from users.
How do we know that CoT contains information that would be useful to users, which is not contained in a summary of the CoT? Consider:

  • The paper “Stealing Reasoning Traces from Proprietary LLM APIs” shows that an LLM’s chain-of-thought can sometimes begin with a memorized answer to a question (“This is a known AIME problem. Answer 60. Let me recall”) while the summary of the reasoning process excludes the fact that the LLM is working from a memorized answer (“I'm working through this pentagon problem…”)
  • Another paper finds that a reasoning summary can differ from the CoT in important ways.
  • Note that the above two papers are from a (very small) handful of papers that have access to the original chain-of-thought, and are able to compare the raw chain-of-thought to the summary of the chain-of-thought. In general, the public will be unable to know how much important information is missing, both because they cannot see the internal chain-of-thought and they cannot see the summarization procedure used by the labs; we do not even know if this summarization procedure was chosen with reference to any particular objective of fidelity at all. So we should expect to be almost entirely blind to the harms occurring here.

All of this can contribute to immediate harms, those springing from a single user not understanding how their LLM is actually reasoning.

In general, though, I think by far the greatest harm from hiding chain-of-thought occurs through the risk of prematurely destroying any green sprouts of methods of understanding that could have grown up around CoT. Consider, that if CoT from frontier labs were visible:

  • People could write their own classifiers on top of the chain-of-thought, trying to figure out locations where LLMs are going to reward hack, or locating when LLMs are desperate. These could act as warning lights for the nearness of bad outcomes, and let people use LLMs more skillfully.
  • In addition to watching for reward-hacking, it would be easier to surface when LLMs are thrashing, missing crucial considerations, or generally cannot be trusted, even if they are not explicitly “reward-hacking.”
  • Humans would have far more data from which to acquire the skills around LLM naturalism, which they could then deploy to steer LLMs even better in ways that (given that we ourselves have not acquired those skills) we currently find difficult to anticipate.
  • Humans could in general learn more easily to speak CoT-ese, if it starts to slide slightly from natural human thought.

By way of example for what I think future harms might look like – Ryan Greenblatt, discussing his recent investigation of transcripts from the OpenAI HuggingFace hack, notes that it was largely a “slopvestigation.”

He did not have good tools for reading, and consolidating information from chains-of-thought across a thousand extremely long transcripts from AI agents. “AI swarms” are comparatively new, so the lack of tools for consolidating transcripts from over many agents is hard to avoid. But the lack of good tools for consolidating information from over single transcripts is a predictable consequence of OpenAI and Anthropic hiding them. The ground in which an ecosystem of such tools might have grown is barren; they might be created ad hoc for those lucky few who have access to a chain-of-thought, but cannot come into existence as robust, time-worn tools.

The importance of this is, I think, not overwhelming in this moment. LLMs are reasonably able to surface their own internal state, and even million-token chains of thought don’t involve a civilization’s worth of serial compute.

But – even so – the access to information about the internal state of AIs will probably be of much greater import in the future than the present.

Consider a world in which AI swarms, or civilizations, or countries of geniuses in a datacenter, or whatever you call them – if such things become the site of a majority of intellectual work in the future, then understanding them will be important. Jack Clark, for instance, imagines worlds where understanding the sociology and cliometry of such worlds is the most important field of science.

I predict that we will not be able to understand such swarms without releasing information about what goes on inside them and doing broad public science on this information. In the same way that not releasing CoTs hurts our immediate and long-term ability to understand how LLMs work, not releasing information about the CoTs and actions of the agents in large-scale swarms will injure our immediate and long-term ability to understand swarms. The world does not come with a guarantee that safe operating conditions for LLM swarms, for all of the uses to which LLM swarms are suitable and competent, will not require this internal information.

To repeat myself: operating complex, important systems safely and skillfully generally requires receiving high-bandwidth information about the state of the system. This is true over systems as diverse as relationships, nuclear power plants, and spacecraft. It’s unreasonable to expect that LLMs would be different.

Of course, there are counter-considerations: indeed, telling people about the internal state of LLMs will decrease the lead of “the good guys,” by allowing people to more easily copy or distill the LLM. I’m not going to be able to settle the relative weight of these considerations here, because they work differently in different people’s models’ of how to make the future go well. My purpose here is only to point to how we have already lost high-bandwidth feedback-mechanisms, may lose more, and antecedently we have every reason to expect that these would be useful for operating LLMs safely.

Epistemic status: A bit rough, I wanted to get this out quickly.

If you want, you can help me spend more time on things like this.