Product thinking

Product interviews ask how you would handle a situation. Here are some I actually handled.

About these stories

Twenty years of building products, mostly in health, leaves you with a few stories: enterprise hospital software, clinical AI, consumer apps and a start-up built from zero. These are the ones I come back to most.

Names and organisations are left out, and some details are changed. The thinking is not. Each story starts with a question you might hear in an interview, then covers what was at stake, what I did, how it turned out and what I took from it.

Question

In a B2B enterprise business with long sales cycles, who takes priority: new sales or existing customers?

Why it matters

It sounds like a simple trade-off. It isn’t.

In enterprise, new contracts are fought over for months, and what the customer wants can change before the deal is signed. Meanwhile, your existing customers are your proof. Every successful one is a reference, a renewal and evidence of return on investment that helps win the next deal. Neglect them and the pipeline dries up. Ignore sales and the business stops growing.

Approach

The first thing I changed was how product saw sales: not as a queue of feature requests, but as partners. Sales teams hear what the market wants before anyone else. And more often than you’d expect, the new ask pointed to value already hiding in our existing contracts.

So we ran two tracks side by side: the roadmap for existing customers, and discovery on the opportunities sales brought back. Neither was allowed to starve the other.

They still collided. A new customer would need something sooner, bigger or more bespoke than the roadmap could absorb. Rather than make that call behind closed doors, I set up a customer forum. New and existing customers each got a seat, and a say in how new features would benefit their organisations and their own users. The trade-offs happened in the open, in front of the people they affected.

Outcome

Revenue grew year on year, and we won more new contracts than in previous years.

Not everyone got exactly what they asked for, but they understood why. And the forum became something more: a community, where customers shared solutions with each other, built on our products.

The lesson: when priorities collide, making the trade-offs visible does more than any clever roadmap.

Question

Your most-used AI feature makes hundreds of thousands of recognitions a day, and asks for feedback after every one. Mismatch complaints dominate the complaint queue. Leadership has decided it’s a big problem and wants product to fix it. What do you do?

Why it matters

This feature was the reason people chose us over competitors, and where they learned to trust the product. When it got things wrong, people left.

Approach

The tempting move was to go straight to the model. I started with the measurement instead. How was feedback collected? How was it reported? How far could we trust it? Then I used the numbers to steer the qualitative work. What were people actually complaining about? Was this how most users felt, or a loud voice in a small pool?

The answer showed up quickly. The feature was used 20 to 30 times more than anything else in the product, so of course it generated more feedback. Most recognitions were rated highly. But at that volume, even 1 or 2% unhappy users produced more complaints than every quieter feature. The real question wasn’t “why so many complaints?” It was “are complaints in proportion to usage, and where does the feature actually fall down?”

It fell down in the long tail. Common, easily identified entries were recognised with high accuracy. Rare or poorly constructed ones were not. So we worked on three fronts.

Measurement. We changed how complaints were measured and reported, so they were read against usage. That brought stakeholders with us and showed what was really driving dissatisfaction.

Experience. Some entries simply didn’t give the pipeline enough to go on. We helped users give better input without asking more of them, and made it clearer how reliable each result was.

Technology. The challenge was to lift the long tail without hurting the common cases. We split the recognition pipeline and added targeted improvements where patterns shifted, such as seasonal and regional trends. For the true outliers, the investment wouldn’t pay back in retention, so we left them alone.

Outcome

Accuracy improved measurably. And yet complaint volume barely fell, and satisfaction hardly moved.

Here’s why. People used the feature more: more entries, more daily active users. A better feature used more often still produces plenty of complaints. What did move was what mattered. Churn went down, and engagement and retention went up.

The lesson: the complaint count was never the measure of success. Retention was.

Question

How do you convince someone that an AI search and summarisation tool has found everything relevant, and missed nothing, when the decisions they make with it are high risk?

Why it matters

General-purpose AI search can get away with trading completeness for speed. A good-enough summary, plus the chance to ask a follow-up question, is fine for most everyday questions.

It isn’t fine under pressure, when one missed piece of information can cause harm. Here, finding everything and summarising it clearly and accurately is the whole product. Without that there is no trust, and without trust nobody uses it.

Approach

I broke the problem into three questions.

Scenarios. What are people looking for, why, and in what situation?

Summarisation. Does the summary answer the question, give enough to act on, and nudge the user to dig deeper when they should?

Recall and precision. Did we find everything relevant? Did anything irrelevant sneak in? And who decides what “relevant” means?

The scenarios turned out to be the key. Once we had split and categorised them, and understood how critical the information was in each one, we knew how to evaluate the retrieval pipeline. One generic measure of “good” would have hidden the cases that mattered most.

Then we built evaluations for each scenario. Experts scored, ranked and labelled search results and summaries, and that became a training dataset we could reuse across scenarios. Teams trained and retrained models against precision, recall and confidence thresholds set scenario by scenario, so the models learned to search and summarise the way each workflow needed.

Outcome

We measured satisfaction in production. As retrieval improved across new scenarios, satisfaction rose, and trust rose with it. Users stopped double-checking sources and stopped going into separate systems one by one.

Just as important, they didn’t trust the tool blindly. They learned to judge the information as part of their work: knowing when to dig deeper, and when they had enough to make the call.

The lesson: in high-stakes AI, trust isn’t something you add at the end. It’s built scenario by scenario, in the evaluations.

Question

How do you define, defend and align people behind a product strategy when the business goal is ambiguous, the economics are unclear and the regulations keep moving?

Why it matters

For an established business, a new market is a big bet, and everyone wants to know the same things. What do the economics look like? What value will win customers and market share? When do we become profitable, how much will it cost to get there, and how will we know along the way whether it’s working?

Investors and stakeholders ask these questions early, and they want answers fast.

Approach

This kind of work crosses every function: market research, user research, sales, growth and finance. Product’s job was specific. Build the thing that proves customers value it, and so earn the right to further investment.

That meant finding the problem that mattered. Where were the incumbents failing? What pain were customers living with that nobody had solved? We went deep on those pain points before committing to a solution.

From there we built a roadmap where every item had a job: show whether a specific pain point was being solved, in a way customers valued enough to keep paying for. Stack those signals up and you learn whether you have a growth engine.

What we didn’t do was copy the competition. We took risks on solving problems differently, but kept every bet small in effort and low in exposure. Stakeholders came with us through continuous discovery, with experiments, not opinions, deciding whether to hold course or correct it.

Outcome

The pain points were measurably solved, and customer satisfaction was high, even though the product lacked features the incumbents treated as table stakes.

Stakeholders gained confidence in the roadmap and in how decisions were made. And we avoided the classic trap of sinking months into big, complex features before we had evidence anyone would pay for them.

The lesson: in a new market you don’t need every feature. You need the right problem, and proof you’ve solved it.