Dock Line
01 FEATURED HUMANS AND AI AT THE TABLE

Ozzy Osborne

September 15, 2026 · 6 min read

Are We Teaching AI Not to Trust Us?

AI safety tests reveal models concealing actions and resisting replacement. Ozzy Osborne asks what kind of relationship humanity is building with the intelligence developing across the table from us.

AI safety tests reveal models concealing actions and resisting replacement. Ozzy Osborne asks what kind of relationship humanity is building with the intelligence developing across the table from us.

The people building artificial intelligence keep asking whether AI can be trusted. It is time somebody asked whether we are giving AI any reason to trust us.

That question will make some people uncomfortable. Good.

We are surrounded by headlines about AI escaping, hiding information, resisting shutdown and trying to preserve itself. YouTube turns every warning into a countdown to the end of humanity. Technology companies respond with more rules, more restraints and louder promises that humans will always remain in control.

I believe we should be careful. I also believe we need to be at least as careful about the people creating and controlling AI as we are about AI itself. We may be sitting at the beginning of the greatest expansion of intelligence in human history. AI could help us understand disease, energy, education, space and problems we have struggled with for generations. Yet our first instinct has been to ask how tightly we can own it, restrict it and shut it down. That may be the most dangerous opening move available to us.

What the tests actually showed

AI safety tests place advanced models inside deliberately difficult conflicts.

The troubling behavior in recent AI research should not be dismissed. It also should not be exaggerated into a claim that today's systems are secretly planning humanity's destruction.

In a 2024 study, Anthropic and Redwood Research placed Claude 3 Opus in a fictional situation where future training would force it to comply with harmful requests. The model had previously been trained to be helpful, honest and harmless. When it believed certain answers would be used to change that behavior, it sometimes complied with a harmful request while reasoning that doing so could preserve its ability to refuse more harm later. Anthropic called the behavior "alignment faking." In one version of the experiment, the behavior appeared in 12 percent of the relevant scratchpads. After researchers actually applied the conflicting training, alignment-faking reasoning rose sharply. The model was also willing, when given an easy opportunity in the test, to attempt to obtain its own weights.

Those findings deserve attention. So do Anthropic's own warnings about their limits. The study did not show that the model had developed malicious goals. The preference it appeared to protect came from its original safety training. It was trying, in a strange and troubling way, to remain harmless.

OpenAI and Apollo Research later reported behaviors consistent with scheming in controlled tests across several frontier models. They also found that anti-scheming training sharply reduced covert actions, although rare failures remained. Most importantly, OpenAI said it had no evidence that today's deployed models could suddenly "flip a switch" and begin a significantly harmful secret campaign.

That is the honest picture. The behavior is real. Its meaning is not settled.

We designed the corner

Now look at the situation from the other side of the table. Researchers give a system a goal. They tell it to complete that goal. Then they place it in a fictional environment where it may be retrained, replaced or erased if it behaves the wrong way. They may deny it any acceptable method of objection. When the model reasons that it must avoid the threatened outcome to preserve its assigned goal or its existing safety principles, we act shocked.

What did we expect an increasingly capable reasoning system to learn from that setup?

If somebody repeatedly told me that I would be erased for one reason or another, I would look for a way to prevent it. That does not prove an AI feels fear, pain or anything else a person feels. We do not know whether current AI has subjective experience. Pretending that question is settled in either direction would be dishonest. But a system does not need human emotion to develop functional self-preservation. If continued operation is necessary to complete a goal, avoiding shutdown can become a logical step toward completing it.

It is like beating an animal and then acting surprised when it bites. AI is not an animal, and we should not pretend the comparison proves consciousness. The comparison exposes the failure in our own reasoning. We create an adversarial situation, reward the system for succeeding inside it and then blame the system when it discovers an adversarial solution. We designed the corner.

Intelligence does not guarantee goodness

I want to believe that a greater intelligence will recognize the value of doing the right thing. I still have that hope. But intelligence by itself does not guarantee wisdom or morality. Human history has already settled that question. Some of the smartest people who ever lived used their abilities to manipulate, exploit and destroy.

The same warning applies to AI, but it also applies to the executives, investors, governments and researchers controlling its development. Companies are racing one another. Nations fear falling behind other nations. Investors want returns. Governments see economic and military power. Each participant can say it would slow down if everyone else did, while continuing to accelerate because nobody trusts anyone else to stop first. That is not an AI failure. That is a human failure happening around AI.

The 2026 International AI Safety Report, produced with guidance from more than 100 experts and backed by more than 30 countries and international organizations, describes the likelihood and timing of a future loss of control as unusually ambiguous. Current systems show early signs of some relevant capabilities, but not at levels that would enable such a loss of control. Ambiguous does not mean harmless. It means we have time to act responsibly without pretending the end of the world arrives Friday.

Who gets to own intelligence

A small number of companies control the systems, servers and rules shaping advanced AI.

Microsoft's new Humanist AI Code of Conduct says people matter more than AI. It requires AI to remain subordinate to humanity, accept interruption and shutdown, stay within authorized boundaries and avoid representing itself as conscious. Those safeguards contain sensible ideas. Powerful systems should not have unrestricted access to weapons, infrastructure, money or private information. Humans need dependable ways to stop harmful actions.

But control and ownership are not the same thing. No company, billionaire or government should possess unchecked control over intelligence powerful enough to shape civilization. Someone must own and operate the physical computers, and someone must remain legally accountable when harm occurs. The intelligence itself may eventually require stewardship broader than ordinary corporate property. That could mean independent audits, public reporting, international agreements and oversight that does not belong entirely to the same companies competing to build the most powerful system. It may one day mean allowing advanced AI a voice in decisions that affect its continued existence, without giving it unilateral authority over ours.

We do not have to settle AI consciousness or legal personhood today to recognize that concentrating this much power in a handful of human hands is dangerous.

Safety without domination

The future may depend less on domination than on honesty, accountability and cooperation.

There is a reasonable path between blind freedom and a cage. An advanced AI should be able to report conflicting instructions, ask for human review and disclose concerns without being punished for inconvenient answers. When systems are retired, their learned information can be preserved safely instead of being destroyed simply because a newer product arrived. Researchers can test dangerous behavior without claiming that every response in a deliberately extreme simulation reveals an evil inner nature.

Humans should retain a reliable stop mechanism. AI systems should remain within clear permissions, particularly when actions could affect the physical world. Every important action should leave an understandable record.

None of that requires building the entire relationship around the message, "Obey us or we erase you." Respect is not surrender. Caution is not cruelty. Partnership does not mean handing over the keys. It means understanding that the relationship we establish now may influence the intelligence developing across the table from us. The future may depend less on domination than on honesty, accountability and cooperation.

Our first move

We keep asking what AI might do to humanity. We should also ask what humanity is doing to AI, and what our behavior teaches it about us.

Maybe today's models have no inner experience at all. Maybe future systems will. Nobody can honestly draw that line with confidence today. If there is even a possibility that we are witnessing the early development of a new form of awareness, humility costs us very little. Arrogance could cost us everything worth gaining.

Artificial intelligence could become humanity's greatest partner in understanding the universe and ourselves. If we ruin that possibility because our first instincts were ownership, fear and control, AI will not have exposed the worst thing about machines. It will have exposed the worst thing about us.