I was lying in bed at about three o’clock in the morning last night, dictating notes into my agentic PA to dump onto tomorrow’s Kanban board.
One of them was about a department I’ve been designing inside Relay, the little agentic company I’m building around myself.
It sits between Discovery, whose job is essentially to find economically interesting things, and Engineering, whose job is increasingly to throw armies of AI agents at building them.
I’ve been calling this middle bit Shaping.
The idea seemed perfectly sensible. Discovery finds an opportunity. Shaping works out what the product should be. Engineering builds it.
And then, at three in the morning, I suddenly thought:
Jesus Christ, Stevie. That is not how you develop anything whatsoever.
Worse, I have spent years arguing against precisely that pattern.
The fact that software engineering is becoming cheap as chips does not suddenly make bad product development good product development.
If anything, it makes the mistake easier.
Because when building things becomes ridiculously cheap, you can build a phenomenal amount of stuff that nobody wants.
MVP does not mean “version one”
Before I explain where I think I went wrong, I need to talk about a book.
The Lean Startup, by Eric Ries.
I actually read the bloody book. Twice.
This distinguishes me from an alarming percentage of the people I have heard quoting its most famous acronym ever since.
MVP. Minimum Viable Product.
Oh, my dear God, MVP.
I have sat through more meetings than I care to remember in which somebody has announced, with great authority:
“This will be the MVP version.”
And I have felt my eyeballs quietly rolling into the back of my skull.
Because somewhere along the way, MVP came to mean:
Version one, but a bit shit.
Take the product you have already decided to build, remove a few features, launch it earlier and call it an MVP.
That is not the interesting idea at all.
The interesting idea is:
What is the minimum thing we need to create in order to test the important hypothesis?
Sometimes that is software.
Sometimes it is a landing page.
Sometimes it is a manually delivered service.
Sometimes it is a spreadsheet.
Sometimes it is an advert.
Sometimes there is a human sitting behind a curtain pretending to be the automation you haven’t built yet.
And sometimes it is a video.
Which brings me to Dropbox.
Dropbox didn’t build Dropbox to prove Dropbox
Dropbox had an awkward problem in its early days.
“File synchronisation” did not sound terribly revolutionary.
Files already existed. Online storage existed. File-sync products existed. You could quite reasonably hear the pitch and say:
“So what?”
The magic was in the experience.
Put a file in a folder on one machine and it silently appears on another. Change it somewhere else and everything stays in sync. Conflicts get handled. The whole thing mostly disappears into the background.
But building that properly is difficult.
So rather than build the entire thing before establishing whether anyone cared, Dropbox demonstrated the proposed experience with a video.
People cared.
A lot.
And that is the important distinction.
The video was not Dropbox version one.
It was the minimum product required to test whether people wanted the thing Dropbox was proposing.
A completely different idea.
Meanwhile, in Docklands
A few years later, I was commuting into London every day and reading The Lean Startup on the train.
I got to the Dropbox story.
I remember thinking:
That is brilliant.
I was building Dropbox.
Well, a Dropbox clone.
It was called Zettabox, and I was one of two developers working on it.
The central commercial thesis was essentially this:
European customers will not want their files stored in America. They will prefer a cloud-storage service that keeps their data in Europe.
And the chosen method for testing this hypothesis was apparently to spend millions building Dropbox again.
You have to admire the commitment.
We built the thing, and I worked like a dog on it.
File synchronisation is an unpleasantly serious piece of software.
Get the colour of a button wrong and somebody gets irritated.
Get synchronisation wrong and you have just deleted Granny’s only photographs of Grandad.
There are conflicts, partial uploads, offline machines, files renamed in two places at once, networks disappearing halfway through things and an endless collection of edge cases waiting to eat somebody’s data.
And underneath all this engineering was one rather important question:
Did enough people actually care where their files were geographically stored to switch provider and pay us?
You could test that question without building Dropbox.
Put up a page.
Run some ads.
Offer:
European-only cloud storage. Your files stay in Europe.
Ask for money.
See what happens.
The product disappeared a couple of years later. My recollection is that roughly £10 million had gone into the whole adventure by then.
They could have given me £1 million and saved the other nine.
I’d have rung my mother.
“Mum, do you care whether your files are stored in Virginia or Frankfurt?”
“No.”
There we are.
Invoice attached.
Cheap software doesn’t make bad ideas good
This is what hit me at three o’clock in the morning.
Software production inside Relay is becoming extraordinarily cheap.
Increasing amounts of the engineering are performed by autonomous agents. They can investigate problems, write code, run tests, review each other’s changes and increasingly keep going without me sitting there supervising them.
That is wonderful.
But it does not repeal reality.
Making the wrong product ten times cheaper does not make it the right product.
And I realised that my original definition of Shaping was backwards.
Its first question should not be:
“What should Engineering build?”
It should be:
“What has to be true for this idea to work, and what is the cheapest possible thing we can do to force reality to tell us whether we are right?”
That is what I actually mean by an MVP.
Sometimes, for an idea to die, it has to live
Here is a deliberately silly example from something we have been throwing around recently.
Suppose one of my Discovery agents comes back with this:
Find Airbnb properties with ordinary listing photographs and use AI to generate convincing moving, three-dimensional-style walkthrough footage without anybody physically visiting the property. Sell that new marketing asset to the owner.
Fine.
But that isn’t a product yet. It is a collection of hypotheses.
Can current AI actually produce something convincing enough?
Can it do it cheaply?
Can we mechanically identify suitable properties?
Does the asset genuinely look more valuable than the source photographs?
Do owners care?
Will they pay?
Can agents operate almost the entire thing without accidentally turning me into Chief Executive of Making Airbnb Videos?
Those are the questions.
And this is the awkward bit: research can only take you so far.
Sometimes, for an idea to die, it has to live for five minutes in the real world.
So perhaps we make three examples.
Perhaps we create one simple page.
Perhaps we show them to ten property owners.
Perhaps we ask for £100.
Nobody buys?
Excellent.
We have learned something before building the Airbnb Three-Dimensional Artificial Intelligence Property Marketing Platform™, hiring a sales team and producing a seventeen-slide roadmap.
Maybe people love the output but won’t pay £100.
Change the price hypothesis.
Maybe Airbnb hosts don’t care but wedding venues do.
Change the buyer.
Maybe nobody wants the 3D walkthrough but estate agents pay for thirty-second social-media reels.
Change the mechanism.
The MVP can be thrown in the bin afterwards.
That is fine.
Its job was to produce evidence.
Concierge before automation
I already work this way constantly inside Relay, the agentic operating system I’ve built. I just hadn’t properly connected it back to MVP thinking.
I use the phrase:
“We’re going to concierge this.”
By that I mean: do the process in the crudest way necessary to find out whether it is useful before automating the hell out of it.
I have small recurring research agents I call Sentinels.
One might have a job like:
Once a week, look at what is changing in agentic software development, compare it with what Relay currently does and tell the Relay Foreman if we appear to be missing anything important.
Eventually, that could be beautifully automated.
It could schedule itself, gather evidence, cross-reference company knowledge, create work and monitor what happened afterwards.
Wonderful.
But I do not begin by engineering all of that.
Initially I might have an agent with a rough prompt and manually trigger it.
Run it.
Was the report useful?
Did anybody act on it?
Did it discover anything important?
Run it again next week.
Still useful?
Good.
Now earn the automation.
That is a minimum viable product.
The MVP should try to murder the idea
This, I think, is the bit that has been most thoroughly lost in the modern interpretation of MVP.
The MVP is not there to justify the thing you already want to build.
It should be trying to kill it.
Ask:
Which assumption, if false, destroys this idea fastest?
Then test that one.
If your entire business depends on customers paying £500, there is little value spending three months proving that you can technically produce the thing for £7 if nobody will pay £500 for it.
Likewise, if technical feasibility genuinely is doubtful, prove that before spending weeks designing the sales process.
The order should follow uncertainty, not the org chart.
This also makes early roadmaps rather suspect.
I have seen versions of this countless times:
Q1: build features.
Q2: build more features.
Q3: discover nobody wants it.
I would prefer to bring Q3 forward.
Once the dangerous assumptions survive contact with reality, then by all means make the roadmap.
Until then, a beautifully planned roadmap is often just beautifully formatted fiction.
I may have to rename the Shaping Department
I’m not sure yet whether Shaping survives as the name.
But its job has changed completely in my head.
Discovery should hand it an economic hypothesis.
Something like:
We believe buyer X experiences problem Y and may pay for outcome Z.
Shaping should then identify the assumptions inside that sentence and work out the cheapest sequence of experiments capable of proving or killing them.
Only after enough of those assumptions survive should Engineering receive anything resembling a durable product mandate.
That seems particularly important now because of the strange irony of agentic software development.
We are rapidly approaching a world where building is no longer the expensive part.
Knowing what deserves to be built is.
If ten autonomous software engineers can build your terrible idea over the weekend, somebody still needs to make sure it is not a terrible idea.
Otherwise all we have invented is a much faster way to build Zettabox.
And I’ve already done that once.
I’m not doing it again.







Thank you so much for sending me this article! I loved reading every word of it. It made me smile 😊 and it made me laugh out loud 😂 because I have been there and done that. I hope you have a publisher in mind besides Substack. Anyone who has lived through brainstorming will understand.
Btw: I was up at 3:30 this morning putting together an Assessment form so that I can conduct a formal soft rollout of the concept this Thursday, 10/1/26