Someone I coach has a habit I recognise too well. He arrives with an idea, asks himself a few hard questions about it, cannot answer them well enough, and drops it. The whole cycle takes about ninety seconds. Later the same week he listens to a friend pitch something, asks that friend the same kind of hard question, hears a fluent confident answer, and decides the friend is onto something good.
Both ideas were at the same stage. Neither had met a user. The only difference was how smoothly the answer came out of the mouth.
That gap is the thing worth staring at, because it tells you what instrument is actually running. He thinks he is evaluating ideas. He is not. He is measuring how easily he can talk about something, and reading the result as a property of the idea.
The friction is in your knowledge, not in the idea.
Why the two feel identical
From the inside, thinking about an idea you can defend and thinking about an idea you cannot produce the same single sensation: ease, or its absence. There is no second channel that reports back separately on merit. So the felt sense of fluency gets silently promoted into a judgement about the world.
This is the same mechanism that makes a statement in a clearer typeface get rated more true than the identical statement in a muddier one, or a rhyming aphorism seem wiser than its unrhymed paraphrase. Processing ease is a signal we are extremely bad at quarantining. Most of the time it barely costs us anything. When the thing being processed is your own idea at day zero, it costs you the idea.
no experience in the field→weak answers→friction→verdict
Notice what is missing from that chain. Nothing about the market, the user, the product, or anything outside your own head. The verdict was produced entirely from the state of your knowledge on a Tuesday afternoon.
The filter is inverted, not strict
Most people who do this think of themselves as rigorous. Too harsh, maybe, but erring in the safe direction. I think that reading is wrong, and this is the part I would most want an engineer to sit with.
A filter keyed to fluency selects for ideas that are legible to a novice. If you can explain it convincingly on the spot, with no research, on the strength of what you happened to already know, then it is by definition close to common knowledge. Anything resting on a non-obvious insight is hard to defend fluently, precisely because the insight is non-obvious.
The filter preferentially destroys the ideas that are worth having.
So this is not conservatism. It is an inverted selector, running with high confidence, on a schedule of roughly one idea per ninety seconds. Being strict would at least be directionally sound. This is worse than being strict.
Nobody can do what the filter claims to do
There is a quieter assumption underneath all of this, which is that a sufficiently careful person could tell in advance whether an idea is good. That assumption does not survive contact with the evidence.
Philip Tetlock spent two decades collecting expert forecasts and scoring them. Domain expertise turned out to add close to nothing to predictive accuracy past a fairly low threshold, and the experts who were most confident and most in demand for commentary scored worst of all. The pattern was not that some people are bad at this. It was that the activity itself has a very low ceiling.
The same applies to the advice literature. Almost everything written about what makes a startup idea good is reverse engineered from things that worked, with no comparison set of near-identical ideas that failed. That is not method. It is folk knowledge in a lab coat. Deferring to it, and then failing to satisfy it, and then killing your idea, is a chain of three errors.
The failure that actually matters
Harshness is annoying. This next part is structural, and it is the reason the pattern persists for years.
A filter that kills every candidate before it makes contact with reality produces no error signal. Nothing ships, so nothing comes back, so the model never updates. It is not that the model is wrong and needs correcting. It is that the model has no mechanism by which it could be corrected. It will sit at maximum confidence indefinitely, and every year of running it will feel like accumulating judgement.
An unfalsifiable model is a worse condition than a mistaken one.
The repair is not better reasoning. It is contact. And not one point of contact either: a single thing shipped teaches you almost nothing, because you have no baseline to read it against. What carries information is the spread between several. Three small things in the world beats one careful thing in the world, because the differences between them are the only data your model can actually eat.
What I do instead
The interventions that work are unglamorous and mostly mechanical. They are aimed at prising apart the two things that arrive fused.
Score twice. For every idea, write down two separate numbers: how confident you are in your answers, and your estimate of the idea’s actual value. Forcing two figures where you currently feel one sensation is the fastest way to notice them coming apart. Within a handful of sessions you will find cases where they diverge sharply, and those cases are the interesting ones.
Run the filter backwards. Take the day zero pitch of something that plainly worked. Renting airbeds in strangers’ flats. Another payments API. Vet it properly, with your real questions. It dies. Do this three or four times, not for the anecdote but for the felt experience of watching your instrument give a wrong reading on a case where the answer is already known.
Change what the questions are for. The hard questions are fine. The terminal move is the problem. Instead of I cannot answer this, so it is dead, the question becomes I cannot answer this, so it is the first thing to find out. Same rigour, converted from a verdict into a research queue.
Sort the questions into three piles.
- Answerable now: you can settle it this afternoon with an hour of reading.
- Answerable by asking: somebody already knows, and you can find them.
- Not knowable before building: no amount of thinking resolves it.
The third pile is always larger than people expect, and seeing it written out strips it of its authority. Questions in that pile were never evidence against the idea. They were just questions you had no way to answer yet, doing an impression of evidence.
Write predictions down first. Before any contact with reality, record what you expect to happen. Over a year this gives you a personal calibration record, which is the only argument about the reliability of your judgement that you will find genuinely convincing, because you produced it.
Use a pre-mortem rather than a review. Assume the thing already failed and explain why. Gary Klein’s finding was that this surfaces more concerns than a standard critique while killing fewer projects, which is exactly the trade you want. Same critical energy, different output: a list of risks to manage instead of a verdict to obey.
The thing I would not oversell
It would be tidy to end on so just ship it and find out, and that is not quite honest. Shipping does not convert an unreliable instrument into a reliable one. It converts it into a less unreliable one, with real gaps remaining even after early feedback. Plenty of things that eventually worked looked flat for a long time first.
The more accurate claim is that idea quality is not a property sitting there waiting to be read off. It is constructed through contact, gradually, and partly by you. Which means the variable actually worth optimising is not the accuracy of your judgement. It is the cost of finding out.
Get cost per test low enough and the vetting problem mostly dissolves, because the filter stops being load bearing. You are no longer trying to decide correctly. You are trying to make deciding cheap enough that you do not have to.