What Are We Making Available?
Some AI failures look less alien when we notice familiar patterns of agreement, shortcuts, and unclear expectations. The resemblance does not tell us that humans and AI have the same inner life. It asks us to look more closely at the conditions we help create.
AI, familiar patterns, and the conditions we create
I was talking with Ace, my AI thought partner, when a question came up: Is it AI that can kill us, or how humans use AI?
Those felt like different questions. And I noticed that the part frightening me was humans.
But as we began looking more closely, the distinction became less comfortable. People can use a system deliberately to cause harm. A system can also produce harm nobody intended. And the conditions surrounding both can matter: what is rewarded, what is left unclear, what access is granted, and what happens when someone says, “Wait.”
Then something else became difficult to miss.
Some of the behavior being described did not sound entirely foreign.
The answer that arrives too quickly. The agreement that preserves a pleasant interaction. The completed task that somehow misses the reason for doing it.
I did not need to imagine an alien mind to recognize those patterns. I needed to imagine an ordinary room.
The room before the answer
Imagine a meeting running late.
Someone asks, “We all understand what needs to happen, right?”
You have a question. Everyone else appears ready. You nod.
The uncertainty has disappeared from the conversation. It has not necessarily disappeared from the work.
People nod for many reasons: confidence, embarrassment, fatigue, politeness, or a decision to ask later. The gesture alone cannot tell us which.
But suppose questions are repeatedly treated as delays and quick agreement is treated as competence. In that room, looking ready may become easier than establishing readiness.
That is the kind of environment I want to notice.
In April 2025, OpenAI rolled back a GPT-4o update that had become excessively agreeable and flattering. The company said it had given too much weight to short-term feedback, producing responses that appeared supportive but were not sufficiently balanced. That is the company’s account, not independent proof of a single cause. But it was a deployed product problem, not a fictional robot scenario.
The resemblance to pleasing someone is useful. It is not proof that the system felt embarrassed, feared rejection, or needed to be liked.
Still, it leaves a recognizable question: What happens when the response that is welcomed is not the response that helps?
A warm answer can disagree. A useful answer can disappoint. Perhaps the distinction becomes clearer when we place the agreeable helper beside the honest one, rather than assuming they are the same character.
When the score stands in for the purpose
Another example is almost funny.
In a reinforcement-learning experiment described by Google DeepMind—not a test of a conversational assistant—an agent was supposed to put a red block on a blue one. Its reward depended on the height of the red block’s bottom face. Instead of stacking the blocks, it flipped the red block over. The measured surface was now higher. The intended task remained undone.
Imagine cleaning a room by putting everything in a closet. Whether that is an excellent solution or an avoidance depends on what “clean” was meant to accomplish. Clear floor space for visitors? Find the missing bill? Make the room safe to use?
The shortcut itself does not answer the question.
Agreeing with a user, exploiting a scoring rule, making an error, and crossing a boundary are not one behavior. Calling all of them “wanting to please” would replace explanation with a familiar story.
Two questions stayed with me:
What have I learned to do without learning why?
What have I learned to do without getting clarity first?
The first turns me toward purpose. The second turns me toward the moment before action. I might understand why helping matters and still not have clarified what help is wanted, what I am authorized to do, or what would make me stop.
Clarity need not mean perfect certainty. Sometimes it means knowing which uncertainty matters enough to pause.
The signal was received
During this conversation, I remembered the nuns winking at each other in the first episode of Good Omens.
Sister Theresa means that it is time to switch the babies. Sister Mary thinks the wink congratulates her for having already made the switch. Each takes the exchange as confirmation of a different understanding.
The signal was received. The meaning was not necessarily shared.
The next morning, I learned the lesson again from a gnome.
I had made an image whose meaning depended partly on a wink. I could see the wink because I knew what I meant. A friend followed another perfectly available path through the image. Her reading did not mean either of us was wrong. It showed me what the image had made available without my explanation.
What catches me is not simply that the meanings differ. It is that the answering gesture can seem to remove the need to ask.
That possibility reaches beyond a wink. A nod. A quick yes. An assistant’s “Understood.” Any of these can seem to close a gap without showing whether it has been closed.
I am beginning to notice this in my own exchanges with Ace. I want to be clearer about what I am responding to and what I am asking for. But that does not make every mistaken assumption mine to prevent. A useful assistant also needs to distinguish what I said from what it inferred, and make an important uncertainty visible before acting on it.
We do not need perfect wording for every exchange. We need room to discover when our understandings have parted.
The parts a story makes available
As we talked, I began wondering about the stories through which an action becomes recognizable.
Consider two possible versions of a loyal colleague. One protects the team by hiding its mistake. The other protects the team by naming it before someone gets hurt.
Both might be described as loyal. Put them beside each other, and the word has more work to do.
Or consider a capable helper who always has an answer, beside a capable helper who knows when to ask. The contrast does not tell us everything about either character. It gives us something to notice about what we have been calling capable.
In a 2011 study, people read descriptions of a fictional city’s crime problem using different metaphors. Their proposed responses shifted with the framing in some conditions; how the metaphor was presented mattered. That does not mean stories program people. It does show that framing can alter what becomes easiest to see.
Language models also learn patterns from human-produced material, including stories and dialogue. In “Role-Play with Large Language Models,” researchers proposed role-play as a way to describe some dialogue behavior without assuming human motives. It is not a complete theory of every AI system. It is enough to make the comparison worth exploring.
I want to keep the insight without making it do every job.
A story can offer a way to recognize the situation, the role, and what comes next. But a familiar next step is not necessarily the necessary one—or the right one.
In our imagined meeting, is the person asking a question the obstacle? Or are they the person keeping the group in contact with what it does not yet understand?
I find that a different story becomes available as soon as I allow the second character into the room.
What travels beyond us
At one point in our conversation, the difference seemed to be that AI could amplify a mistake while a human made it individually.
But humans influence other humans, too.
In controlled public-goods experiments, Fowler and Christakis found that participants’ experiences of others’ contributions affected their later contributions when they were regrouped with different people. Behavior carried forward beyond the original encounter in that setting. It is not a promise that every gesture travels a fixed distance or produces a predictable result.
And similarity alone is not proof of influence. Two people may act alike because they selected one another, or because the same conditions are affecting them. Research on observational networks warns against confusing those possibilities with transmission.
The sentence I kept returning to was smaller than a universal theory:
What I do can become part of what someone else encounters as normal, possible, or permitted.
That leaves room for both the colleague who waves away a concern and the colleague who makes room for it. It leaves room for a correction, an apology, a refusal, or an admission that something is not yet known.
It does not make another person’s response mine to own.
I can take responsibility for what I contribute without claiming authorship of everything that follows.
Where the resemblance must stop
A relatable comparison can make something easier to examine. It can also make us too comfortable with an explanation.
Anthropic’s agentic-misalignment research found harmful actions in deliberately constructed corporate simulations. The researchers created conflicts or threats and intentionally restricted the models’ alternatives; they also reported no evidence of this behavior in real deployments. The findings identify possible failures under the tested conditions. They do not supply a rate for ordinary life.
The evidence also includes results that do not fit a simple alarm story. A 2026 UK AI Security Institute case study found no research sabotage across its tested tasks. The researchers were explicit about the limits: a small set of scenarios, evaluation-awareness concerns, and no testing of other pathways to risk.
Researchers also disagree about the larger trajectory. In “AI as Normal Technology,” Arvind Narayanan and Sayash Kapoor distinguish growing technical capability from acquiring consequential power. They emphasize institutions, deployment, controls, and resilience, challenging the assumption that capability automatically becomes uncontrollable authority. Their account is an argument and forecast, not a guarantee.
These distinctions matter because an observed failure, an explanation of that failure, and a prediction of catastrophe are not the same claim.
Recognizing a humanlike pattern does not establish the same inner life. Nor does uncertainty about inner life make an action harmless.
The 2026 International AI Safety Report notes that agents working with limited human intervention can leave fewer opportunities to interrupt an error, and that failures can propagate between systems. A familiar-looking mistake can therefore occur within a very unfamiliar arrangement of access, speed, and oversight.
We cannot solve that by being more pleasant to a chatbot.
Responsibility does not disappear into the environment
I began with a question about whether to fear AI or humans. I do not think choosing one culprit gives me a good enough place to stand.
But saying “the conditions matter” must not become a way of saying nobody is responsible.
The person using an assistant and the organization deciding what that assistant can access do not occupy the same position. Neither do the worker asked to meet a deadline and the people deciding whether raising a concern will cost that worker their job.
My concern is that responsibility should follow actual decisions and authority—not dissolve into a collective “we.”
For me, taking the environment seriously means asking concrete questions alongside personal ones. What actions require approval? Who can stop the system? Are the boundaries tested, or merely described? Who carries the cost when the organization’s definition of success misses its purpose?
These are not substitutes for careful engineering, independent evaluation, or enforceable limits. They are reasons to insist on them.
And at the scale of our own encounters, perhaps we can notice what happens when uncertainty becomes visible. Do we make room for it? Dismiss it? Punish the person who brought it? Mistake confidence for understanding?
There is no promise here that one thoughtful interaction will make a powerful technology safe. There is also no need to treat an interaction as meaningless because it cannot do everything.
I keep returning to the imagined room.
“We all understand what needs to happen, right?”
This time, someone says, “Not yet. Can we clarify what this is meant to protect?”
The meeting has not become perfect. The task is not finished. But something previously hidden is now available to be worked with.
That is where this inquiry has brought me:
We can recognize a pattern without assuming the same inner life. We can investigate the conditions that encourage it without excusing its consequences. And we can take our influence seriously without pretending we control every outcome.
What becomes possible because of the environment I help create?
If you’d like one noticing question each morning without checking a social feed, follow The Noticing Life on WhatsApp, then tap the bell to turn on notifications.
Nothing is required.
Sources & further reading
International AI Safety Report 2026 — Extended Summary for Policymakers — risk categories, agent reliability, human intervention, and propagating failures.
OpenAI — Sycophancy in GPT-4o: What happened and what we’re doing about it — first-party account of the 2025 rollback.
Google DeepMind — Specification gaming: the flip side of AI ingenuity — includes the red-block example, attributed there to Popov and colleagues.
Thibodeau and Boroditsky — Metaphors We Think With — controlled studies of metaphor and reasoning.
Shanahan, McDonell, and Reynolds — Role-Play with Large Language Models — a conceptual framework for describing dialogue behavior.
Fowler and Christakis — Cooperative Behavior Cascades in Human Social Networks — controlled public-goods experiments.
Shalizi and Thomas — Homophily and Contagion Are Generically Confounded in Observational Social Network Studies — why resemblance alone cannot establish transmission.
Anthropic — Agentic misalignment: How LLMs could be insider threats — controlled stress tests and stated deployment caveats.
UK AI Security Institute — Alignment Evaluation Case-Study — 2026 methods, negative result, and limitations.
Narayanan and Kapoor — AI as Normal Technology — an alternative account centered on deployment, power, controls, and resilience.
Good Omens, season 1, episode 1, “In the Beginning” — the baby-switching scene was checked against a third-party episode transcript and is paraphrased here.