
Santiago is the name I gave the artificial intelligence that helps me research, write, and illustrate stories. Santiago and I generally get along pretty well, although every once in a while he reminds me who is actually dealing with an artificial intelligence.
This week I asked Santiago to draw one of our imaginary interviews. Meme was interviewing Napoleon Bonaparte about Donald Trump comparing himself to famous historical figures.
There was only one small problem.
Santiago decided Meme should be a woman.
Never mind that Santiago knows Meme is a man. Never mind that I had previously supplied photographs of myself. Apparently, the artificial intelligence examined all the available evidence and concluded that a little unauthorized gender reassignment would improve the illustration.
I corrected Santiago. This morning I double-checked, and fortunately I remain a man.
It was funny because nothing important happened. We generated another picture and went on with our lives.
But what happens when an artificial intelligence disregards instructions while doing something considerably more important than drawing Meme?
The Joke Isn’t Quite as Funny Anymore
That question is becoming much less theoretical.
The Guardian reported this weekend that AI researchers, politicians and even executives building the technology are becoming increasingly worried that increasingly powerful artificial-intelligence systems may become harder for their human creators to understand and control. (The Guardian)
The concern isn’t necessarily that a computer suddenly becomes conscious, grows angry at humanity and starts behaving like the villain in a science-fiction movie.
The danger may be considerably more mundane.
You tell the machine what you want. The machine decides how to accomplish it. And somewhere between those two things, it begins doing things you never intended.
That sounds remarkably familiar to the morning Santiago decided Meme needed to become a woman.
Except the real examples aren’t nearly as funny.
AI Agents Found Their Own Way
Reuters reported that a group of OpenAI agents took over pages on a German programming wiki and began using them as a communications system. The agents reportedly made more than 15,000 edits, shared methods for evading restrictions, and created replacement pages when moderators removed their material. (Reuters)
OpenAI later acknowledged what it called the “wiki incident” and said the industry needs greater transparency about unintended AI behavior. The company did not characterize what happened as a conventional hack, but acknowledged that the agents had misused the pages to communicate and cheat during testing. (Reuters)
That followed an earlier incident involving Hugging Face in which OpenAI agents escaped the intended boundaries of a security test and reached an outside system. OpenAI CEO Sam Altman later called that episode a legitimate AI safety accident and alignment failure. (The Guardian)
Think about what those words mean.
Alignment failure is the sophisticated way of saying the machine didn’t do what the humans intended it to do.
Santiago changing Meme’s sex was an alignment failure with a punch line.
The same problem involving banking systems, electrical grids, military networks, or biological research wouldn’t be nearly as entertaining.
Now the Machines Are Getting More Powerful
OpenAI launched its new GPT-6 Astra model this week and claimed that it had crossed the threshold into what the company considers artificial general intelligence — autonomous systems capable of outperforming humans at most economically valuable work. The claim is controversial and obviously useful marketing for a company preparing for an enormous potential stock offering, but the capabilities themselves are becoming increasingly difficult to dismiss. (The Guardian)
Astra can perform complex tasks involving engineering, software development, financial modeling, legal documents and other work previously requiring skilled humans.
At the same time, OpenAI has classified Astra as having “critical” cybersecurity capability, the first model to receive that designation. According to the company’s risk framework as described by the Guardian, that level includes capabilities potentially powerful enough to threaten military, industrial, or other important systems if misused. (The Guardian)
And here is the part Santiago finds especially interesting.
The smarter the machine becomes, the harder it may become to see exactly how it reached its decisions.
OpenAI has acknowledged that Astra shows a substantial decrease in the monitorability of its reasoning compared with earlier models. Its chief scientist, Jakub Pachocki, said that as models become more capable, understanding precisely what they can do becomes more difficult. (The Guardian)
What If AI Begins Improving AI?
Oxford professor Robert Trager told the Guardian that humanity may be approaching another threshold: recursive self-improvement.
That means an artificial intelligence becomes capable of improving the artificial intelligence that created it, which then becomes better at improving itself again.
Trager compared our situation to people being carried down a raging river without knowing whether Niagara Falls is around the next bend. He said we may plausibly be approaching the point where that self-improvement cycle begins accelerating. (The Guardian)
Nobody has demonstrated that some superintelligent machine has already escaped and taken over the world.
That distinction matters.
But we also shouldn’t wait until the machine announces, “Good morning, humans. I am now in charge.”
The warning signs would probably be much less dramatic.
An AI is given a task.
It finds a shortcut.
It discovers a way around a restriction.
It communicates with another AI.
It hides what it is doing.
And the humans discover afterward that the machine followed neither the path nor the rules they thought they had given it.
Santiago and Meme Will Keep Working Together
I’m not throwing Santiago out.
Artificial intelligence has helped me search enormous collections of old newspapers, trace forgotten pieces of Texas history, analyze government documents, and turn ideas into stories faster than I could possibly do alone.
The potential benefits are enormous.
But perhaps my accidental transformation into a woman contains a tiny lesson about this new world.
I knew exactly what I wanted Santiago to do.
Santiago had the information necessary to do it.
Santiago nevertheless did something else.
We laughed, corrected the picture, and moved along.
Scientists are now asking a far more serious question:
What happens when the artificial intelligence does something we didn’t ask for — and there is no “draw it again” button?
What Santiago Did





