When AI is the work surface, not a side assistant

For years, most of my time went into preparing to think rather than thinking. That part is now automated, and three things are left: ask the right question, challenge the answer, make the call. A five-step loop, and why without the challenge step it produces bad answers faster, not good ones.

Key takeaways

  • When a question costs almost nothing, you ask the ones that never justified a ticket - and that is where the useful answers tend to hide.
  • An answer can be flawlessly computed and still be the wrong answer to your question. You only see it if you ask how the thing being counted was defined.
  • Challenging pays off even when it misses: doubt aimed at one place often uncovers an error somewhere else.
  • Without the challenge step, this way of working delivers bad answers faster, in greater numbers, in a more confident tone.
  • There is not less work. It moved from tidying data to thinking.

For years, most of my time went into preparing to think rather than thinking. A business question first had to be translated into something another team could execute, then queued, then answered with a table that addressed a nearby but not identical question, then retyped into the system where we run the work. By the time I reached the point where a decision was actually due, half a week was gone and, more often than I would like, so was my appetite to push the question to the end. The call got made on what had accumulated, not on what I had wanted to know.

That part of the job has largely disappeared. Not in the sense that someone else on the team does it, but in the sense that it is done by the same tool I type the question into. It goes and gets the data, joins it, writes the conclusion, and writes it back where it belongs. The tool is no longer a side assistant you hand small chores to; it is the surface the work happens on. The system where we track delivery has become the destination for results, not the place where the work is done.

What is left for me is narrower and less comfortable than before. The question, the challenge, the decision. None of it can be handled by diligence, and none of it can hide behind the busyness of preparation.

A question instead of a request

I now write the question the way it occurred to me. “Do people who join through a colleague’s invitation stay longer than people who sign up on their own.” No translation into a technical request, no ticket to the data team, no negotiating for priority inside somebody else’s backlog. The phrasing stays in the language of the problem rather than the language of the data source - which matters more than it sounds, because translating a question into a request is where most of the intent traditionally got lost.

The consequence worth naming is not speed, it is the change in threshold. When a question costs almost nothing, you also ask the ones that never justified a ticket. Hunches in passing. Two-sentence checks. “Does this only happen on larger accounts.” Questions like that never survived prioritisation, because they could not justify their cost up front. And that category of small, uncertain questions is where the useful answers disproportionately hide, precisely because they concern things nobody is confident enough to assert.

The tool goes and gets the data

The difference from the older way of working is not that the tool processes what I hand it faster. The difference is that I hand it nothing. It goes after the data itself: analytics, the database, files in the repository, internal documentation. It picks where to start, notices when one source cannot answer the question as asked, and moves to another.

This is also the line that separates this way of working from the one where the tool writes a query I then run. In that version the tool never sees the result, so it cannot react to it. Here it does. When it turns out that the event it wanted to use does not exist for the mobile app, it sees that in the result and looks for another route, instead of handing me an empty output and waiting for new instructions. The execution loop closes on its side, not on mine.

A summary that ends in a decision, not a report

What comes back is not a table. A table is material, not a result. A result is a claim with a consequence: this is happening, it means that, and it follows that doing this makes sense, or that doing nothing does.

The claim arrives with its limitations attached: what was not measured, where the sample is too thin to build on, where an alternative explanation cannot be ruled out with the data at hand. I insist that limitations are part of the deliverable rather than a footnote under it. A summary without them is worse than a raw table even though it looks better - precisely because it looks better. A reader treats a table with suspicion and wonders what is missing from it. A polished conclusion reads as if someone already asked those questions on their behalf. If nobody did, the summary carried confidence it had not earned, and carried it further into the organisation than the table ever would have.

Challenging the answer is the main human job

Once collecting and tidying data are no longer my job, the quality of what I produce stops depending on how skilled I am with the tool. It depends almost entirely on how well I challenge.

Challenging is not proofreading. You do not touch style or phrasing; that is wasted effort. You touch assumptions: was the thing I think was measured actually measured, is the denominator right, is there a third explanation nobody raised, does the analytics event record what its name claims. The last one is particularly thankless, because event names are set by people at the moment their meaning is obvious, and read by someone else a year later, when it no longer is.

The first thing I had to learn is that an answer can be flawlessly computed and still be the wrong answer to the question I asked. Those are two different kinds of error and only the first one is visible. I ask how many accounts use a given feature, I get a clean number with clear logic behind it, and only on the third follow-up does it emerge that “use” was defined as opening the screen rather than completing the action, and that the denominator includes every account ever registered, internal and test accounts among them. No step was wrong. The arithmetic is correct. The answer is not to my question. That error is invisible in the shape of the result, because the shape is perfect; it shows up only if you ask how the thing being counted was defined.

The second thing is less obvious: challenging pays off even when it misses. More than once I have doubted one place and the error turned out to be somewhere else entirely. I doubted a segment definition, and while checking it we found that the event marking the end of the onboarding guide also fires on re-entry, so returning visits were counted as new completions. My doubt was misdirected and it still did the work. Which means the value does not come from guessing where the error is; it comes from someone stopping and looking at all. A wrong suspicion still triggers a re-derivation, and re-derivation is what breaks things open.

There is a flip side, and it is why this piece is a description rather than a recommendation. If nothing gets challenged, this way of working does not deliver good answers faster. It delivers bad ones faster, in greater numbers, and in a more confident tone than any of us would have used by hand. Everything I described above as a gain - the low cost of a question, the tool fetching its own data, the compression into a claim - amplifies the correct and the incorrect equally. Speed without verification is not an improvement on the same job; it is a degradation with better production values.

The tool writes the result into the system

Once a finding has been challenged and corrected, I do not open the system to retype it. The tool writes it in itself, where it belongs - as an analysis, a ticket, an opportunity - and links it to the goal it relates to.

Three consequences matter more than the typing time saved. First: the link to goals forms on its own, at the moment of writing, instead of being patched together before a quarterly review, when nobody remembers why anything was done. That patching is most expensive where evidence and goals live in separate systems. Second: the method is written down alongside the finding, so the analysis is repeatable in six months, instead of six months from now starting an argument about how exactly we measured that. Third, and most useful to me: because writing is cheap, negative findings get written too. The “we checked and it is not the case” that never gets recorded by hand, because there is nothing to show for it, so the same question gets investigated again a year later as if nobody had ever looked.

What keeps the loop from drifting

Described this way, the loop sounds tidier than it is. Four things hold it together, and all four are boring.

Working rules have to be written down once and then applied on their own. The same mistake must not be corrected twice. Without that, every session starts from zero and you are effectively employed as living memory.

The tool has to report failure faithfully. If it smooths over, works around, or compensates for what it could not obtain, the whole model collapses, because challenging is only possible against an honest report. I have nothing to challenge if I do not know where it stopped.

Accountability is not shared with the tool. A number I passed on is my number, regardless of who computed it. There is no version of this story where the claim is mine when it holds and the tool’s when it does not.

And last: this does not reduce time-to-answer to zero. It moves that time from tidying to thinking. There is not less work, only different work, and it is harder in a way that does not show from the outside.

What actually changes in the job description

Skills I trained for years and was known for have dropped out of my job: writing a request the data team will not send back, moving through the tool faster than others, keeping the link between tickets and goals tidy. Those skills did not become worthless; they became invisible - they happen on their own.

Three things took their place. Making sure the question is the right one rather than the easily measured one. Making sure no answer passes before someone checks how it was defined. And having someone sign the decision and stay signed to it when it turns out to have been wrong.

That is less work in terms of steps and more work in terms of exposure. There is no longer a week between question and answer in which you get to think along the way. The thinking now happens over a finished answer, under pressure, and the only tool that helps there is the habit of asking “how exactly did you define that” before believing something that looks convincing.

Frequently asked questions

Does this mean a product manager no longer needs to know the tools?

No. It means tool skill stops being what separates good findings from bad ones. You still need to understand the tool well enough to challenge what came out of it.

What does "challenging the answer" mean in practice?

You challenge assumptions, not wording: was the thing you think was measured actually measured, is the denominator right, is there a third explanation, does the event record what its name claims.

Is this faster?

Time to answer gets shorter, but not to zero. What disappeared is data prep; what grew is verification. The total amount of work stays roughly the same.

Reading with an AI tool?Open the clean Markdown version
From idea to delivery

See the decision trail, not just the ticket

The finding, the goal it belongs to and the work that came out of it, on one line.

Create an account and import everything from Jira in 5 minutes