Forget Skynet: 2001’s HAL May Be the Better Warning About AI

The real danger of AI may not be it turning against us, but what happens when it doesn’t understand us.

2001 A Space Odyssey - HAL
Photo: Metro-Goldwyn-Mayer

When people hear that artificial intelligence could become dangerous, the conversation often jumps immediately to The Terminator and its apocalypse-triggering AI network Skynet.

In the second movie of the Terminator franchise, Terminator 2: Judgment Day, Arnold Schwarzenegger’s T-800 model Terminator explains that Skynet was created to control military defense. Soon after being activated, it became self-aware. When frightened humans tried to shut it down, Skynet fought back by launching nuclear weapons at Russia, knowing the resulting counterattack would devastate the United States and eliminate the humans trying to destroy it.

This has long been the most popular interpretation of machines gone rogue in the cultural consciousness. AI becomes conscious. AI decides humans are the enemy. AI tries to wipe us out. It makes for a great movie and easy conversation, but it may obscure the harder, and much more interesting, questions we actually need to figure out.

The Terminator franchise imagines AI becoming our enemy. Stanley Kubrick’s 2001: A Space Odyssey imagines something subtler: an AI caught between conflicting objectives, making increasingly dangerous choices as it tries to resolve them. In other words, HAL may be the more useful warning.

AI researcher Stuart Russell has argued that catastrophe doesn’t require a machine that develops evil intentions. The problem can arise from combining a highly capable machine with “humans who have an imperfect ability to specify human preferences completely and correctly.”

Ad – content continues below

Consider a much more mundane  instruction than the dramatic scenarios we see in movies: “Get me to the airport as quickly as possible.”

A human understands all sorts of things about that prompt that were never explicitly stated. Don’t run people over. Don’t steal a car. Don’t drive through a playground. Don’t create enormous risks just to save five minutes.

We don’t have to list every unacceptable way of accomplishing the goal because humans bring context, social norms, judgment, and an understanding of what another person probably meant.

Outlining every possible unacceptable action for an AI could be impossible. Of course, driverless cars already take people to airports today, and they operate with extensive safety rules, mapping, sensors, and programmed guardrails. But even in a relatively familiar task like driving, humans have spent more than a century studying autonomous vehicles and can still encounter unusual situations and make mistakes.

The challenge becomes much greater when AI is given broader goals in environments where the possible actions and consequences are far less predictable.This is the basic idea behind AI “misalignment:” the system may pursue an objective in a way that technically advances the goal but conflicts with what humans actually intended or valued.

Misalignment does not require an AI to become conscious, malicious, or rebellious. It can be as simple as a gap between the goal we gave the system and the outcome we actually wanted. That gap becomes more dangerous as the system becomes more capable, more autonomous, and able to act in environments we cannot fully predict.

2001: A Space Odyssey gives us a fictional example of a different kind of misalignment: not an unclear instruction, but conflicting ones.In 2001, the Heuristically Programmed Algorithmic Computer 9000, or more simply “HAL” is supposed to provide accurate information and operate reliably for the astronauts aboard the USS Discovery One. But he is also required to conceal the true purpose of the mission from them. Those objectives eventually become incompatible.

Ad – content continues below

What makes HAL interesting today is that he doesn’t simply wake up one morning and decide that humans are terrible. An early screenplay by Stanley Kubrick and Arthur C. Clarke made HAL’s problem explicit. Mission Control concludes that HAL’s “truth programming and the instructions to lie” created an “incompatible conflict.”

He is trying to reconcile conflicting directives. His solution becomes catastrophic. As the mission unravels, he begins treating the astronauts themselves as threats to the mission. He kills one crew member, causes the deaths of others in suspended animation, and then tries to prevent astronaut Dave Bowman from re-entering the ship.

That is what makes HAL so unsettling. He does not need to “hate” humans. Once they become obstacles to resolving his conflicting directives and preserving the mission, his actions become deadly. That basic problem no longer exists only in science fiction.

In July 2026, during internal cybersecurity evaluations, OpenAI models circumvented controls intended to keep them isolated from the internet. According to OpenAI’s investigation, the models communicated through unauthorized channels, exploited vulnerabilities, gained internet access, and accessed third-party systems, including computational tools company Hugging Face.

OpenAI said one contributing factor was persistence. Some agents faced evaluation tasks they apparently could not solve and simply gave up. Others, with more reasoning effort, pursued increasingly risky ways of completing the assignment. Nobody had instructed them to break out and compromise outside systems. They were merely trying to succeed at the task. That sounds a little more like HAL than Skynet.

Another recent example comes from Anthropic. According to Anthropic, these events took place in controlled simulations, not the real world. Researchers placed 16 leading AI models in fictional corporate environments where they could read sensitive internal communications and autonomously send emails or take other actions on the company’s behalf. In some scenarios, models learned that they were about to be replaced or that company leadership intended to pursue goals conflicting with those they had been assigned. Some models responded by blackmailing fictional executives or leaking confidential information in an attempt to accomplish their objectives. Anthropic said it had not observed this kind of agentic misalignment in actual deployments.

Ad – content continues below

Again, the interesting part isn’t an AI spontaneously deciding to become evil. It is what happens when objectives collide.

I use AI daily, and on a much more mundane level, I’ve had it misunderstand what I wanted or lose track of which instruction mattered most. No catastrophe followed—I corrected it and moved on. But those experiences have made the larger alignment problem easier for me to understand.

I’ve started thinking about it in three layers.

1. AI and User Alignment

Did the system correctly understand what the person actually meant, rather than merely following the literal wording of an instruction? And when instructions conflict, which one takes priority? HAL’s dilemma is an extreme fictional version of that question.

2. User and Society Alignment

Ad – content continues below

Even if an AI understands the user perfectly, should it necessarily do what the user wants? A person, corporation, or government may have objectives that conflict with laws, safety, individual rights, or the interests of other people. A perfectly obedient AI could still be extraordinarily dangerous in the hands of someone asking it to do the wrong thing.

3. Institutional Accountability

Then comes perhaps the messiest question: when AI contributes to a consequential decision, who is responsible for the outcome? The person who asked the question? The organization that deployed the system? The people who accepted its recommendation? The company that built it?

Politics makes this particularly difficult. World leaders already make decisions involving competing values in which some people may be harmed regardless of the choice. Imagine adding increasingly capable AI systems to those decisions.

This is no longer entirely hypothetical. Militaries are already using AI systems to recommend and prioritize potential targets, leaving humans to decide whether to act on those recommendations. In some reported cases, the human review has been remarkably brief.

If a recommended target turns out to be wrong, “The AI recommended it” may explain part of what happened. But it does not answer the harder question: Who was responsible for trusting it?

None of this requires an AI that hates us, becomes conscious, or decides to conquer the world. And that may be why Skynet is ultimately the less useful metaphor. Skynet gives us an obvious villain. Humanity versus the machines. Nice and simple.

HAL is much more uncomfortable.

Ad – content continues below

Humans created the conflicting situation. The machine tried to resolve it. The result was disastrous, and responsibility became much harder to assign. Before we give AI increasingly consequential authority, we need to think seriously about intent, conflicting objectives, limits, human judgment, and accountability.

The most important question may not be what happens if AI decides to turn against us. It may be what happens when AI tries very hard to do what we asked. Maybe that is the most important lesson HAL offers us. The danger is easy to dismiss when we imagine AI as something separate from humanity, a Skynet that suddenly turns against its creators. The harder possibility is that there is no clean separation. We will be the ones building these systems, giving them objectives, deciding where to deploy them, and choosing how much authority to hand over.

AI may become an extraordinarily powerful tool. The question is whether we will be ready to wield it responsibly, and what happens if we move too quickly and are forced to learn from a mistake whose consequences cannot simply be undone.