startup-stack

Two-Week Experiment

One risky assumption. One test. One number. Two weeks.

One risky assumption. One test. One number. Two weeks.

The point of an experiment is that it can fail. If you cannot describe a result that would make you change your mind, you are not running an experiment — you are building something and hoping.


The assumption

We believe that: [the risky belief, stated so plainly it could be wrong]

If this is false, then: [what breaks — which part of the plan collapses]

Type:Desirability — do they want it? ☐ Feasibility — can we build or deliver it? ☐ Viability — does the money work? ☐ Compliance — are we allowed to?

Why this one first: [of everything unproven, why is this the assumption most likely to kill you?]

Test the riskiest thing first, not the easiest. Founders systematically test feasibility — the part they enjoy — while desirability, the part that actually kills companies, stays untested for a year.

The test

We will: [the specific thing you will do]

Starts
Ends
Owner
Cost
Time it will take

Sample size: [how many people, shops, calls, impressions] Is that enough to conclude anything? [be honest — 4 data points is an anecdote]

The measure

Metric[one number]
How it is measured[precisely — two people should compute the same figure]
Baseline today
Success threshold[≥ X — decided now, before you see the result]
Failure threshold[≤ Y]
Ambiguous range[between — means run it again differently, not "it kind of worked"]

Set the threshold before you run it. A number decided afterwards will always be met, because you will find a reading of the data that meets it. This is not dishonesty; it is how everyone's brain works, which is why the discipline exists.

What we do with each result

IfWe will
Success[the next commitment, with its own number]
Failure[what specifically changes — pivot the offer, pivot the segment, drop the channel, stop]
Ambiguous[what the next test looks like]

The failure row is the one that matters. "We would look into it further" is not a decision — it is the absence of one, and it is how companies spend two years on an assumption nobody ever falsified.


Result

Ran: [dates] · Actual sample: [n]

Number: [ ] versus threshold [ ] → ☐ Success ☐ Failure ☐ Ambiguous

What happened:

What surprised us:

What we are doing about it:

Where this got recorded in the stack: [section]


Common ways this goes wrong

Testing what you were going to build anyway. If the answer does not change what you do next, it was not an experiment. It was work with a report attached.

Moving the threshold. See above. Write the number down, in this file, before you start.

Too small a sample. Four conversations is an anecdote. Twenty is a signal. Where the sample has to be small — enterprise sales, institutional pilots — say so, and treat the result as directional rather than conclusive.

Testing three things at once. If you change the price, the creative and the channel in the same fortnight, you learn nothing about any of them.

Not writing down the result. The single most common failure. The test runs, the founder forms a vague impression, and six months later nobody can remember what the number was — so the same assumption gets re-litigated from scratch.

Generated from worksheets/experiment.md in the repository. Edit the markdown, not this page.