One of the best developers I worked with could help me find a bug without touching the keyboard or giving me the answer.

When I first started developing, I had several generous mentors. The one who stands out most was a quiet Russian developer. Walking over to his desk meant receiving the stare: a long, unreadable look capable of making you reconsider the problem, your career and briefly the meaning of existence.

Then he would ask, “What is the problem?”

I would begin describing the code, what I had changed and all the surrounding detail. He would interrupt: “I asked what the problem was.”

Once I could state it, he would ask what the logs, debugger or breakpoints showed. If I had not looked, the conversation was over until I had. If I had, the next question was usually the one that mattered.

“So what does that mean?”

Sometimes he spotted the fault. Sometimes he highlighted the part of my account that did not agree with the evidence. Often he simply waited until I heard what I was saying and the answer arrived.

I would thank him and walk away, either impressed with myself or quietly recalibrating my opinion of my own intelligence. He would return to work having said almost nothing.

The answer was not the help

For a long time, his ability looked like magic. With more experience, I began to see a repeatable sequence underneath it. He separated three things I was trying to collapse:

  • the problem from all the detail surrounding it
  • the evidence from the story I had started telling
  • what the evidence showed from what I assumed it meant

I had usually approached his desk hoping to borrow an answer. He returned the thinking to me instead.

That was not the same as refusing to help. If I was missing knowledge I could not reasonably discover, he taught me. If the system contradicted everything we could observe, he joined the investigation. The stare worked because he was listening closely enough to know whether I needed information, another question or a few more seconds of silence.

Silence without attention is abandonment. His silence was a form of attention.

Finding other ways to hear the problem

Around the same time, I heard about rubber duck debugging: the practice of explaining a problem to an inanimate listener. I tried it, but talking to an object in an office felt silly and strangely public. I worried that it made me look less capable, not more thoughtful.

I tried holding the conversation internally. That helped while building, although leaning back with my eyes closed could be misread as a different approach to employment.

Voice notes worked better. Near the end of the day, I would record the problem as though I were asking someone for help: what I was trying to do, what I had observed, what I had tried and where my understanding stopped.

By the next morning, I often knew the answer before replaying the recording. Hearing it again still helped. An assumption sounded less like a fact. A detail I had treated as incidental became central. Occasionally the entire explanation revealed that I had built a difficult solution around a problem nobody had asked me to solve.

Research on self-explanation offers a related, though much narrower, mechanism. In a study involving 24 eighth-grade students learning about the circulatory system, 14 were asked to explain each line of a passage to themselves while ten read the passage twice. The prompted group showed a greater gain in understanding (Chi and colleagues, 1994).

That was a small educational study, not an experiment in professional software debugging. It does not prove that a voice note will fix a production fault. It supports the modest idea that generating an explanation can do work that merely receiving or rereading one does not. Explaining requires us to connect information, and the missing connections become easier to notice.

I found other ways to interrupt fixation. I went for walks while trying not to think about the challenge, switched to a different task, or took a tightly timed nap. Box breathing sometimes helped when the emotional weight of the problem had become larger than the problem itself. These practices did not hand me an answer. They reduced the gravity around the current answer long enough for another interpretation to appear.

Learning to give the stare

As my experience grew, I moved from asking questions to being asked them. I soon discovered that the stare was a skill in its own right.

A developer, product manager or stakeholder rarely arrives with the complete history of their thinking. They arrive with its current result. Terms have acquired local meanings. Alternatives have been rejected without being named. An important constraint may have become so obvious to them that it no longer appears in the explanation.

This is close to the distinction I explored in Sometimes the breadcrumbs are the story. We do not need to rediscover every thought from the beginning, but we need enough of the path to understand what the conclusion is connected to.

The difficult part is listening without taking possession of the problem. As a developer, the urge to jump in and fix something is strong. I can stop hearing the person while they are still speaking because I have begun constructing the solution.

Asking “why?” does not automatically repair this. A repeated why can sound like an interrogation, particularly when someone believes they have already explained a simple request. “What does that mean?” can sound like a test. A technically useful question can still close the conversation if it communicates that I am evaluating the speaker rather than trying to understand the work.

I have found it more useful to make my uncertainty visible: “I think I have treated this as a deployment container, but you may mean a container in the code. Which one governs the requirement?” The question no longer asks the other person to defend their competence. It gives them an assumption they can correct.

Good listening is not passive. It translates words into a tentative model, exposes the parts that model supplied for itself and leaves room for the speaker to change it.

Three answers can reveal the question

At some point, from a source I can no longer remember, I picked up a rule that has stayed with me:

If you have not thought of at least three ways to solve a problem, you may not yet have discovered the problem.

It is too much ceremony for a trivial or easily reversed change. For a difficult requirement, architectural decision or disputed request, it has repeatedly earned its cost.

The three approaches have to be valid and materially different. Three versions of the same architecture with renamed components do not count. One might optimise for a quick reversible experiment, another for integration with the existing system, and another for a longer-lived boundary. Each must still meet the intent and acceptance criteria.

Comparing them makes the requirements work harder. A stated constraint may turn out to describe a preferred implementation. An acceptance criterion may only test the first approach. A familiar term may mean different things to different people.

I have seen a proposed solution use a Docker container because the request mentioned “containers”, only to discover that the original speaker meant a structure within the code. One solution could carry that misunderstanding for quite a long time. Three varied solutions dragged it into the light because the word could not keep the same meaning across all of them.

The final answer is sometimes the first approach. That is not surprising: it is usually where our initial conviction and most of our early effort went. The alternatives are not there to make the first answer wrong. They test whether it still holds once its assumptions and trade-offs become visible. Sometimes the result is a hybrid. The value is not the theatre of presenting three boxes. It is the refinement that happens while trying to make all three honest, before the team has invested its identity in one answer.

AI is too willing to complete the thought

This history feels newly relevant when I work with AI. We often describe an assistant as better when it requires less of us: less context, less precision and fewer corrections before it produces something useful. Those are genuine improvements. They can also remove the moment in which an unfinished thought would have revealed itself.

We give AI vague requests, fragments and throwaway comments, expecting the model and its surrounding tools to work out what we mean. As systems improve, they become better at producing a plausible completion. Plausibility is helpful when the missing detail is incidental. It is dangerous when the missing detail decides which problem gets solved.

Typing changes the way I explain something. I can refine the request while composing it, delete a false start and make the account more technically precise before anyone sees it. That can improve clarity. It can also produce a polished version of what I currently think, stripped of the uncertainty and corrections that would help someone understand how I reached it.

Speaking is closer to the way I work through a problem with a colleague. The thought develops while I am saying it. I go too deep into one detail, hear that I have contradicted myself and correct it aloud. The earlier version does not disappear unless I explicitly replace it: “No, that is not quite right. What I mean is…” Those changes are useful information. They reveal where the model in my head is still moving.

Voice notes preserved more of that movement, which is one reason they helped me. They also preserved repetition, digression and more detail than another person necessarily needed. However, speaking to a person or machine throughout an eight-hour day is exhausting for me. Voice helped me recognise what the interaction was missing; it is not how I want every interaction to happen.

The useful question is whether the interaction behaves more like an attentive exchange with a colleague. Does the AI notice when my explanation changes? Does it distinguish a correction from another fact? Can it reflect the model it has heard, make its additions visible and ask about the contradiction instead of silently resolving it? A text interface can do those things too, but only if the system values the unfinished reasoning rather than rushing to clean it up.

AI can otherwise polish the result until the assumption is harder to see. The earlier note about breadcrumbs concerns what happens when the path behind a conclusion disappears. This problem is earlier still: the assistant can help the conclusion arrive before the governing question has been formed.

What does the running system say?

My mentor did not stop once I had stated the problem clearly. His next questions moved the conversation away from my account and towards the system. What had I observed? What did I think the system was doing? What did the logs, debugger and breakpoints show? Only then did he ask what the evidence meant.

That sequence is easy for an AI agent to shorten. Given a bug report or issue ticket, it reads the description, searches the code, finds a plausible path and begins changing it. The ticket becomes the observation, and a static reading of the code becomes an explanation of how the running system behaves.

Both are evidence, but of different things. The ticket contains someone's account of the behaviour. Cold code shows what the implementation appears able to do. An existing log records what the system emitted under conditions we may not fully know. None is the same as reaching the reported path and watching the failure occur now.

A bug ticket is a claim about behaviour. Reproduction turns it into an observation.

That distinction matters because cold code can tell a convincing story. A branch looks reachable but is disabled by runtime configuration. A suspicious function is never called in the reported path. A test passes because it reproduces the implementation's assumptions rather than the user's experience. Occasionally the issue has already disappeared or the reported behaviour meant something different from what the agent inferred.

When it is safe and practical, I want an AI agent to establish a baseline before editing: reproduce the behaviour, record the conditions, and capture a failing check that distinguishes the observed problem from the expected result. Only then should it decide which part of the code deserves suspicion. After the change, it should repeat the same path and show what changed.

Reproduction is not always possible. Production-only failures may be unsafe to trigger. Intermittent faults can disappear when observed. The required data may no longer exist, and external systems may be beyond the agent's reach. In those cases, the honest state is not “verified”. It is a hypothesis supported by the evidence available, with the missing observation stated explicitly.

This is the rest of the stare. First: what is the problem? Then: what have you observed, how do you think the system works, and what does the running system show? The final question—what does that mean?—only becomes useful once there is something current and concrete to interpret.

AI should not only make us explain the evidence we already have. Sometimes it should ask why neither the human nor the agent has tried to produce the evidence the decision requires. Requiring that observation before solutioning is another form of useful friction.

The cost of useful friction

There is some evidence that deliberately slowing AI-assisted decisions can change how people use the advice. In an experiment with 199 participants completing a food-substitution task, researchers compared ordinary AI explanations with three “cognitive forcing” designs. Participants either requested the AI advice on demand, made an initial decision before seeing it, or waited before it appeared (Buçinca, Malaya and Gajos, 2021).

The forcing designs reduced overreliance on incorrect AI suggestions compared with showing explanations directly. The trade-off matters: participants gave the least favourable subjective ratings to the interventions that reduced overreliance most, and the average benefit was greater among people more inclined towards effortful thinking.

This was a constrained decision-support task using simulated AI, not generative software development. It does not establish that an assistant should obstruct every request. It reveals a design tension that resembles my experience: the interaction people like most may not be the one that best preserves their independent judgement.

Three small experiments with the stare

I have been experimenting with a few AI skills. Their names are personal shorthand rather than a finished framework. What matters is the behaviour they are trying to introduce.

/wait-what interrupts solutioning. It asks what I am actually describing, why it matters and what value a change would create. Not every problem needs a solution, and not every unfinished thought needs an enterprise-grade implementation assembled from familiar patterns.

/make-assumption makes an assumption explicit for both the AI and me. Instead of silently filling a gap, one of us states what we currently believe. If it is accurate, we can proceed. If it is wrong or uncertain, the gap becomes a question rather than inherited context.

/query-three-challenge asks the AI to do more work before asking me to choose. Where it sees a material ambiguity, it develops three independent approaches that could satisfy the intent and acceptance criteria, tests them and summarises what differs and why. I can then choose, correct the framing or offer another approach.

Three AI-generated answers are not automatically diverse. They can share the same source, terminology and blind spot. The useful test is whether changing the approach exposes different requirements or consequences, not whether the output contains three numbered headings.

Early trials have been encouraging. I also use these behaviours between agents, not only between an agent and me. An agent can state the assumption it is handing to the next one. Another can test whether three proposed approaches really meet the same intent. A reviewer can ask what the available evidence means before treating a fluent summary as a decision.

I am beginning to suspect that some AI drift starts here: not only when context is compressed or evidence is missing, but when an uncertain assumption passes through a conversation without either participant making it visible. This is an observation from my work, not a general explanation for hallucination. The narrower claim is useful enough: agents cannot challenge an assumption that the interaction encourages everyone to leave implicit.

Challenge must earn the interruption

An assistant that challenges every sentence would be unbearable. It would move the burden of indiscriminate generation into indiscriminate clarification. Low-risk, reversible and familiar work often deserves a direct answer.

The amount of friction should rise with ambiguity, consequence and the cost of reversal. A copy edit does not need three architectures. A decision that establishes a shared boundary may deserve them. A person learning a system may need knowledge, not another question. A team responding to an incident may need action now and reflection after stability returns.

The purpose is not to maximise conversation or make the human perform intelligence for the machine. It is to reveal material uncertainty before hidden choices become implementation. Ask for the problem when the symptoms have taken over. Ask for the evidence when the story is moving faster than observation. State an assumption when a missing detail could change the direction. Create alternatives when one answer has become too easy to defend.

Then help. The stare is a pause, not a permanent operating state.

The assistant that waits

When I imagine better AI collaboration, I do not only imagine a model with more answers. I imagine one with better judgement about when an answer would take the thinking away from the person who still needs to understand it.

I still hear my mentor's sequence. What is the problem? What have you observed? How do you think the system works? What do the logs and debugger show? So what does that mean?

None of those questions is sophisticated. Their value comes from timing and attention. He did not mistake my first account for the problem, the ticket for an observation, cold code for the running system, evidence for interpretation or speed for help.

The useful assistant is not always the one that reaches the answer first. Sometimes it is the one that notices we have not finished forming the question.

The stare was never an absence of help. It was a way of returning the thinking to the person who would have to understand the answer after he had gone back to his work.

Further reading

These books offer related perspectives rather than direct proof of the argument. They are included as a reading list to examine, not as endorsements of every claim they contain.