How do you validate product market fit before you build?
Learn how to validate product market fit before building through interviews, landing pages, pre-sales, and honest concierge demand tests.

You cannot prove full product-market fit before a product exists. You can, however, validate the chain of assumptions that must hold before building: a specific group has a recurring problem, they already spend time or money on it, your promise makes sense to them, they will commit to a solution, and you can deliver the result at a workable cost. That is enough evidence to decide whether to build the smallest real product.
Founders get into trouble when they collapse all of those claims into one vague question: "Do people like my idea?" A friendly interview can support the problem claim. A signup on a landing page can support the message and channel claims. A paid pre-sale can support willingness to buy. A concierge service can show whether the promised outcome is possible. None proves the whole business, and treating one signal as proof is how teams spend six months polishing something nobody adopts.
You can validate the path to product-market fit
To validate product market fit before building, test the riskiest assumptions with progressively stronger customer commitments. Start with past behavior in interviews, move to a public offer, ask for money or another meaningful commitment, then deliver the outcome manually. Build software only after the evidence survives that sequence.
The wording matters because product-market fit describes a relationship between a product and a market. Before customers can use, keep, recommend, or abandon a product, you do not have that relationship to measure. You have evidence about demand. That evidence can make a build decision rational, but it cannot make future retention certain.
I separate early evidence into four layers:
- Problem evidence shows that a defined customer experiences the problem often enough to act.
- Demand evidence shows that your offer causes that customer to take a meaningful step.
- Delivery evidence shows that you can produce the promised result, even with manual work.
- Fit evidence appears after use through retention, repeated purchase, expansion, referrals, or distress at losing the product.
Each layer answers a different question. Twenty people describing an ugly spreadsheet process may justify a demand test. They do not justify a claim that people will buy your proposed workflow. Ten deposits may justify delivering the service manually. They do not tell you whether customers will keep paying after the novelty wears off.
This distinction also prevents a common argument between co-founders. One founder says, "The interviews were great, so we should build." The other says, "Nobody has paid, so we learned nothing." Both are overstating the case. The interviews may have reduced problem risk while leaving purchase risk untouched. Name the remaining risk and design the next test around it.
Write the failure rule before the experiment
A demand test is useful only when a disappointing result can change your decision. Before recruiting anyone, write the customer, claim, observable action, time window, pass rule, and consequence. If you choose the rule after seeing the results, you will quietly lower the bar for an idea you already love.
Use a single experiment record like this:
Customer segment:
Riskiest claim:
Test and acquisition channel:
What the customer sees:
Commitment requested:
Pass rule and time window:
What we will do if it passes:
What we will do if it fails:
Known sources of bias:
Result, with raw counts:
Decision and date:
The pass rule is your decision policy, not an industry benchmark. A founder selling a $40 workshop and a founder selling a $40,000 annual contract should not expect the same funnel. Set the rule from the economics and access available to your business. If you need ten annual customers to support the first year, a test that reaches only students with no purchasing authority has the wrong sample even if its conversion rate looks lovely.
Raw counts belong beside rates. "Ten percent converted" hides whether one person acted out of ten visitors or 100 acted out of 1,000. Record where participants came from, how many saw the offer, how many were eligible, and what each person actually did. Small samples are normal at this stage. They call for restraint, not statistical theater.
Strategyzer's work on testing business ideas recommends identifying the assumptions that combine high importance with weak evidence, then choosing experiments for those assumptions. I agree with the sequence but add a sharper constraint: test the assumption that could kill the business soonest. If customers care but you cannot reach them without spending more than the sale is worth, more interviews about the problem will not save the idea.
Interviews reveal behavior, not demand
Customer interviews should establish how people handle the problem today, how often it occurs, what it costs, and who controls a purchase. They are poor tools for predicting whether someone will buy an imagined product. People try to be encouraging, and founders make the problem worse by explaining too much.
Recruit people who match a behavioral definition, not a broad identity. "Independent US accountants who prepare more than one multistate filing a month" is recruitable and falsifiable. "Busy professionals who value efficiency" is advertising copy. Include people outside your network; friends and warm introductions often grant attention because they trust you, not because the problem hurts.
Ask about a recent event. "Tell me about the last time you closed the books" produces names, steps, files, delays, and workarounds. "Would an automated close assistant help?" produces a guess about a future self. Stay with the story long enough to learn what triggered the work, what the person tried, where the process broke, who noticed, and what happened next.
A useful interview has a simple spine:
- Ask the person to reconstruct the last specific occurrence.
- Trace the current workflow and every handoff.
- Identify what they have already tried, paid for, or stopped using.
- Quantify the consequence in the units they use, such as hours, missed revenue, rework, or risk.
- End by asking for an introduction to someone with the same problem.
Do not pitch during the evidence portion. You can show a concept after you understand the current behavior, but label reactions to the concept as a separate result. A compliment such as "I would use that" costs nothing. Sending a redacted spreadsheet, introducing the budget owner, scheduling a follow-up, or agreeing to a pilot requires effort and carries more weight.
Look for convergence in behavior, not repeated adjectives. Five people may call a task "annoying" while only one has tried to fix it. That one person may be the better early customer. Conversely, different job titles may describe the same expensive workaround, which can reveal a tighter segment based on the triggering event rather than the title on a profile.
Interview evidence is strong enough to revise the customer and problem hypotheses. It is not a purchase forecast. Move on when new interviews mostly repeat known workflows and you can recruit the segment without explaining who qualifies. If every conversation uncovers a different problem, narrow the segment or admit that the idea is still a collection of anecdotes.
A landing page tests an offer and a channel together
A landing page tests whether a particular audience, reached through a particular source, understands an offer well enough to act. It does not test the product because visitors cannot use the product. It also does not estimate total market size unless the traffic represents the market, which founder networks almost never do.
Build the page around one customer, one situation, one promised outcome, and one call to action. Show enough detail for a visitor to reject it intelligently: who it is for, what changes, what the delivery looks like, what it costs if you know the price, and when it will be available. A vague "Join the future of finance" page may collect curiosity while teaching you nothing about the actual offer.
The call to action should match the risk you need to test. An email signup asks for a small commitment. A booked qualification call asks for time and exposes whether the visitor fits. A refundable deposit or purchase asks for money. Do not compare those actions as though they carry equal evidentiary weight.
Strategyzer's reference guide for landing pages makes a useful qualification: an email submission proves that a person invested time and shared an address. It does not prove that the person will pay. Founders often cite the guide's first point and forget the second. Treat the conversion as evidence of interest at the price and level of detail shown, nothing stronger.
Traffic quality can reverse the apparent result. A page sent to 200 former colleagues measures goodwill inside your professional circle. A page reached through the search phrase your intended buyer already uses measures the combination of message and channel. Paid traffic can buy a clean test, but only if targeting resembles the acquisition channel you could actually afford later. Record results for each source rather than blending every visitor into one rate.
Keep the page stable during the declared test window. If you change the headline, audience, price, and button after every few visits, the final number describes several offers mixed together. For a small audience, run separate rounds and keep the raw result for each. The goal is a decision, not a conversion graph that always slopes upward.
A failed page has several possible meanings. Qualified visitors who reach the page and leave may reject the promise, price, proof, or timing. Nobody reaching the page may mean the channel failed before the offer got a fair test. Talk to a few eligible non-converters, review where they dropped, then change one major claim in the next round. Do not rescue the result by counting likes, time on page, and compliments as substitutes for the action you requested.
Pre-sales expose the gap between interest and purchase
A pre-sale is the strongest demand test available before full delivery because the customer accepts a real price, terms, and delay. Money changes the conversation. Buyers ask about implementation, refunds, security, procurement, deadlines, and who will do the work. Those objections are part of the evidence, not friction to edit out of your notes.
Make the offer honest. State what exists now, what the buyer will receive, when you expect to deliver, what remains manual, and how cancellation or refunds work. A "Buy now" button that quietly turns into a waitlist measures surprise and damages trust. If you only want to test checkout intent, do not capture a charge; explain what happens before the customer confirms.
For a consumer physical product in the US, the Federal Trade Commission's Mail, Internet, or Telephone Order Merchandise Rule deserves attention. The FTC says a seller needs a reasonable basis for the stated shipping time. If no shipping time appears, the rule generally uses 30 days; delays require notice and a chance to consent or cancel for a prompt refund. A founder taking pre-orders should read the current FTC guidance and get advice for the specific product rather than treating "pre-sale" as a legal exemption.
B2B commitment may take a different form. A paid pilot is excellent evidence when the price resembles the eventual model and the buyer has authority. A signed letter of intent is weaker because wording ranges from a polite expression of interest to a negotiated commitment. Record whether it names price, scope, decision maker, start date, and conditions. Never put "five LOIs" in an investor deck as if every letter carries the same obligation.
Discounts can contaminate the test. A modest founder offer may compensate early customers for delay and rough edges. A 90 percent discount may validate a different product with impossible economics. Likewise, a deposit that is instantly refundable proves more than an email but less than a payment the buyer expects you to earn. Describe the commitment precisely.
Read objections by frequency and consequence. If prospects want the outcome but procurement requires a certification you cannot obtain, the sale may still be blocked. If one prospect requests an unusual feature in exchange for a large contract, decide whether that request represents your segment or turns the company into custom consulting. Revenue is evidence, but revenue from the wrong customer can point the build in the wrong direction.
Concierge delivery shows what the product must accomplish
A concierge test delivers the promised outcome manually to a few real customers before you automate it. It tests whether the result matters, what inputs the work requires, where trust breaks, and whether customers return. It also exposes the labor and exceptions that a prototype can hide.
Suppose you want to build software that prepares a weekly cash forecast for small agencies. Instead of building integrations and a forecasting engine, ask two agency owners to provide the same exports the future product would use. Create the forecast by hand, deliver it on the promised schedule, and observe what decisions follow. Tell customers that people perform the work; do not stage manual labor as finished automation.
Paul Graham's "Do Things that Don't Scale" argues that founders can sometimes act as the software and automate later. The advice is popular because it gets a result in front of a user quickly. It becomes bad advice when founders mistake a service with no fixed scope for evidence of a repeatable product. The manual version should follow a consistent promise and capture the steps, time, inputs, exceptions, and outcome each time.
Charge for the concierge test when payment is part of the risky claim. Free work can still reveal workflow and outcome, especially when regulation or procurement delays payment, but it cannot answer the same demand question. If you waive the fee, ask for another commitment that matters: access to real data, a scheduled decision meeting, a reference call after success, or a signed pilot plan.
Track the full cost even while you ignore margin for the first few deliveries. Separate work that software could plausibly automate from judgment that requires an expert and from customer acquisition or support. If each delivery needs a founder to interpret an ambiguous goal for three hours, the work may support a premium service. It does not yet support a product that customers use alone.
The best concierge result is not "the customer liked it." It is an observable change followed by another commitment. The customer uses the forecast in a staffing decision, requests the next report, pays the next invoice, invites the finance lead, or hands over data without repeated chasing. If the deliverable sits unopened, polish will not repair the weak outcome.
Grade every signal by the cost of saying yes
Honest interpretation starts by asking what the customer risked. Words carry little cost. Time, reputation, data access, process change, and money carry more. A signal gets stronger when the person matches the segment, understands the offer, has alternatives, and still chooses the commitment without personal pressure from the founder.
Classify each observation after the round. A prospect who describes a recent workaround confirms that the problem exists for her, but not that she wants your solution. A qualified visitor who submits an email gives the offer a small commitment, but says nothing yet about payment or continued use.
A budget owner who books a sales call gives the problem and message some of her time; procurement can still fail. A customer who pays a realistic price demonstrates willingness to buy under the stated terms; repeat use and scalable delivery remain open. A customer who returns after manual delivery shows that the outcome may deserve repeated use, while the economics of reproducing it in software remain untested.
Founder effort can inflate every signal. People introduced by a close friend may take meetings as a favor. A sale led by a founder can succeed because the buyer trusts the founder's expertise. Concierge customers may stay because the founder performs bespoke work. Do not discard those results, but mark the dependency and test whether it survives with a colder channel, standard scope, or someone else delivering.
Negative evidence deserves the same precision. A landing page with no purchases does not prove the problem is imaginary if the ads reached the wrong people. An interview with no urgent pain is more informative when the participant clearly matches the segment. Ask whether the test exercised the claim before declaring the claim dead.
Beware of evidence laundering. This happens when a weak action changes names as it moves through a team: clicks become leads, leads become interested buyers, and interested buyers become pipeline. Keep the original event visible. "Four of 37 qualified visitors paid a $50 refundable deposit" may sound less grand than "strong pipeline," but it lets you make a decision.
The same discipline applies to silence. Nonresponse is data about that message and channel, but you cannot infer why each person ignored you. Follow up within the protocol you set, then close the round. Endless follow-up can manufacture replies while hiding that acquisition will require founder persistence you cannot sustain.
Mixed results should narrow the next bet
Most serious tests produce mixed evidence, and the right response is usually to narrow the claim. You may find a painful problem with no budget, a clear offer with an expensive acquisition channel, strong pre-sales with impossible manual delivery, or enthusiastic use by a segment too small to support the planned company. "Validated" and "invalidated" are often too blunt.
Read results across the assumption chain:
- If interviews show no repeated behavior, change the customer or problem before testing copy.
- If interviews converge but the page fails with qualified traffic, revise the offer, price, or promised outcome.
- If the page converts but buyers will not pay, reduce the gap between the page's promise and the commercial terms.
- If customers pay but concierge delivery does not change behavior, fix the outcome before automating.
- If customers pay and return but delivery costs too much, redesign the operation or choose a price that can support it.
Consider a founder testing a service that prepares investor update drafts for CEOs at the seed stage. Eight interviews reveal that updates slip because financial and hiring data live in separate documents. A landing page sent through a founder newsletter gets 18 email signups, but only two match the intended stage. The founder should not report an 18-lead win. The channel reached a broad audience and the qualification failed.
She narrows the page to CEOs who already send monthly updates and offers a paid, manually prepared first draft using their existing files. Six qualified people visit, three book calls, and one pays. During delivery, that customer refuses to share the promised data and wants general writing help instead. The payment supports willingness to buy something, but the delivery contradicts the original workflow claim.
The next test should address data access, not add writing features. The founder can offer the same scope to another small cohort, show the exact input checklist before payment, and require sample files during onboarding. If customers who accept those terms use and reorder the update, the evidence chain gets stronger. If they keep withholding inputs, the concept needs a different data path or a different promise.
This is why a single conversion target rarely settles a build decision. The chain is only as convincing as its weakest necessary assumption. A high page conversion cannot pay for labor. A profitable manual service cannot prove software adoption. A founder's job is to locate the weakest link without pretending the stronger ones disappeared.
Build when the next uncertainty requires a product
Start building when repeated commitments support the customer, problem, offer, price, and outcome, and when the largest remaining risk can only be tested through product use. That risk might be whether users complete onboarding alone, whether an automated result matches the manual one, whether a team returns weekly, or whether the acquisition economics hold beyond the founder's network.
The first build should isolate that uncertainty. If the concierge test proves customers want a recurring report, build the smallest path that lets them provide inputs and receive it with less manual work. Do not add team permissions, a settings center, and six export formats because mature products have them. Every extra feature delays the behavior you need to observe.
After launch, measure fit with behavior appropriate to the product's natural cycle. A daily workflow tool may reveal retention quickly. Annual tax software cannot. Look for continued use after the initial push, repeated purchases, expansion to colleagues, unsolicited referrals, and a shrinking need for founder rescue. Y Combinator's Startup School guidance calls retention the best measure of product-market fit because retained users keep choosing the product.
The Sean Ellis survey asks active users how they would feel if they could no longer use the product, with "very disappointed" as the strongest answer. The often repeated 40 percent threshold came from his comparison of startups, but it is a heuristic, not permission to ignore cohort quality or behavior. Run it among users who have experienced the main value, segment the answers, read their explanations, and compare the result with retention. A survey of curious signups before use measures expectation, not fit.
Stop when the central problem does not recur, qualified people refuse meaningful commitments across more than one credible test, or the delivery economics fail under prices customers accept. Narrow when one segment acts and the rest merely praises. Continue testing when the experiment missed the intended customer or asked for a commitment unrelated to the business model.
Founders often keep building because code feels cumulative while rejected offers feel like lost work. The opposite is true. A clean rejection preserves months of runway and gives the next idea a better question. Building earns its place only when customers have already done enough to make product use the next honest test.
Inside Sisters, founders can ask women who have handled customer development, sales, and go-to-market work to challenge an experiment plan or give candid feedback on an offer before the build begins. Bring the raw counts and the failure rule; a polished story gives peers less to work with.
FAQ
Can you validate product-market fit without a product?
You can validate the assumptions leading to product-market fit, but you cannot prove fit before people use a product. Interviews, pre-sales, and manual delivery can establish problem, demand, and outcome evidence; retention and repeated use come later.
How many customer interviews are enough to validate an idea?
There is no universal interview count. Continue until qualified participants repeat the same recent behaviors and constraints, then use a demand test because more conversation cannot prove willingness to buy.
Do landing-page clicks prove demand?
No. A click shows that a message earned attention from that traffic source. A qualified signup, booked call, deposit, or purchase asks for progressively stronger commitment, but none proves retention.
How large should a waitlist be before I build?
A raw waitlist total has little meaning without the eligible visitor count, acquisition source, offer, and eventual economics. Set a pass rule from the number and type of customers your first version needs, then verify that some waitlist members accept a stronger commitment.
Is it ethical to take pre-orders before building?
Yes, if buyers know what exists, what remains to be built, when delivery is expected, and how cancellation or refunds work. US sellers of covered merchandise should also check current FTC shipping rules and get advice for their specific offer.
What is the difference between a concierge test and an MVP?
A concierge test delivers the outcome mainly through human work, with the manual process disclosed to the customer. An MVP is a product someone can use, so it can test product behavior that manual delivery cannot.
What if interviewees love the idea but nobody pays?
Treat the interviews as problem evidence and the refusal to pay as demand evidence. Check whether you reached the buyer and made a concrete offer, then change the segment, promise, price, or model instead of counting compliments as sales.
Does a letter of intent validate demand?
It can support demand, but its strength depends on the terms. A letter that names price, scope, authority, timing, and conditions carries more weight than a polite nonbinding statement with no next action.
When should I stop validating and start building?
Build when customers repeatedly make meaningful commitments and manual delivery shows that the outcome changes their behavior. The largest unanswered question should require actual product use, such as independent onboarding, automation quality, or retention.
How should I measure product-market fit after launch?
Use retention or repeat purchase over the product's natural usage cycle, then examine expansion, referrals, and the need for founder intervention. The Sean Ellis disappointment survey can add context among experienced active users, but it should not replace behavioral evidence.

