Testing a New Product Before Committing the Company
By Mark Bold | emerging-tech | 8 min read
A disciplined evidence ladder lets you test a new product without betting the company. Define what each experiment proves, what would disprove it, and how to protect your core customers while you learn.
Testing a New Product Before Committing the Company
Most new product failures are not caused by bad ideas. They are caused by the gap between what a founder believes is true and what has actually been demonstrated. Enthusiasm, a few encouraging conversations, and a plausible market narrative are often mistaken for evidence. By the time the gap becomes visible, the company has already committed engineering capacity, redirected sales attention, and told existing customers to expect something that may never arrive.
The discipline that protects you is an evidence ladder. Each rung tests a specific claim, has a defined cost, and produces a result that can either support or disprove the thesis. The ladder is not a product plan. It is a sequence of experiments that lets you spend small amounts of money and attention to answer increasingly expensive questions before you commit the company.
Begin With the Claim You Must Prove
Before writing a specification or hiring a designer, write down the single claim the new product depends on. Not three claims. One. A useful claim is specific about the buyer, the problem, the moment the problem is felt, and the alternative the buyer is currently using. A weak claim says that mid market operations teams would benefit from better reporting. A stronger claim says that operations leaders at distribution companies with multi warehouse footprints will pay a defined annual fee to replace the custom spreadsheets they maintain to reconcile inventory across sites.
The second half of the exercise is harder. Write down what evidence would disprove the claim. If you cannot describe a result that would cause you to stop, you are not running an experiment. You are rationalizing a decision you have already made. Disproof should be concrete. For example, if fewer than a defined share of qualified prospects agree to a paid trial after a structured conversation, the thesis is weakened enough to pause.
This matters because the later rungs of the ladder are expensive. If the initial claim is vague, every subsequent test will be interpreted generously, and the project will acquire momentum that is unrelated to actual demand.
Build the Ladder From Interest to Delivery
The evidence ladder has four rungs, and each one tests a different question.
The first rung is expressed interest. Do prospects, when presented with a clear description of the problem and a proposed solution, recognize the problem as theirs and want to hear more. Expressed interest is cheap to generate and cheap to misread. People are polite. They say encouraging things to founders. Expressed interest is useful only to filter which conversations to continue. It does not predict purchase behavior and should never be treated as validation.
The second rung is trial usage. Will qualified prospects invest their own time to use a working prototype, a manual service dressed as a product, or a scoped pilot. Trial usage reveals whether the problem is painful enough to overcome switching friction. It also surfaces the integration, data, and workflow realities that no conversation uncovers. The point is not to make the trial smooth. It is to see whether prospects persist through imperfection because the underlying value is real.
The third rung is paid demand. Will the prospect sign an order, pay a deposit, or commit a budget line before the product is fully built. Money changes the nature of the conversation. It forces the buyer to involve procurement, legal, and often a superior, which exposes whether the champion has real authority and whether the problem has organizational priority. A pattern of paid commitments from independent buyers is the first evidence that cannot be explained by goodwill.
The fourth rung is reliable delivery. Can you fulfill the promise at a cost structure, support load, and quality level that would sustain a business. Early customers tolerate failures that later customers will not. If delivery requires heroics from a few people, the economics do not generalize. This rung tests whether the opportunity is a product or merely a well paid service.
Most companies conflate these rungs. They treat expressed interest as paid demand, or they ship broadly before confirming reliable delivery. The ladder exists to force sequence.
Define Spending Limits and Stop Conditions in Advance
Each rung should have a budget, a time box, and a decision rule written before the work begins. A spending limit is not a target. It is the maximum you are willing to lose to answer the question at that rung. If the question can be answered for less, that is a success, not a reason to spend the rest.
Stop conditions are the hardest part of the discipline. In practice, teams rarely stop. They reframe disappointing results as learning, extend the timeline, and seek a different cohort of prospects. Some of that is legitimate iteration. Much of it is sunk cost reasoning. The protection is to decide, in advance and in writing, what pattern of results would end the experiment, and to have someone other than the project champion confirm whether that condition has been met.
It is also worth defining what a positive result does not entitle you to. A successful trial rung justifies moving to a paid rung. It does not justify hiring a sales team, building a brand campaign, or publicly committing to a launch date. Advancing one rung at a time preserves the option to stop with limited damage.
Distinguish Learning From Selling
Founders and senior executives are often the people running early experiments, and they are usually effective sellers. This is a problem. When the person testing demand is also capable of persuading a reluctant buyer, the signal from the experiment is contaminated. You learn what you can sell, not what the market will pull.
The practical correction is to separate the two activities. In a learning conversation, the goal is to understand how the prospect currently handles the problem, what they have tried, what they would need to see to switch, and what would make them hesitate. The goal is not to close. In a selling conversation, the goal is to secure a commitment. Both are legitimate. Mixing them produces false positives.
One useful test is to review a sample of your early customer wins and ask whether an average sales professional, without the founder in the room, would have closed them. If the answer is no, you have not yet proven product demand. You have proven founder led service demand, which is a different and smaller business.
Protect the Existing Product and the Existing Promise
The cost of a new product experiment is not only the direct spend. It includes the engineering attention diverted from the core product, the executive bandwidth absorbed by the new initiative, and the risk to customer promises that the current business depends on.
Before starting, name what the existing business will not get because of the experiment. Which roadmap items will slip. Which customer commitments will be harder to meet. Which support obligations will compete for the same people. If these cannot be named, the real cost is being hidden, usually from the board and often from the executive team itself.
There is also a reputational dimension. Early access programs, pilots, and announcements create expectations among existing customers. If the new product is quietly abandoned, those customers will remember. The least damaging path is to frame experiments internally and externally as experiments, with explicit language about what is being tested and the possibility that it will not continue. This feels less bold. It preserves credibility, which is a scarce asset in a company that intends to try more than one thing over its lifetime.
Contract terms, data handling commitments, and any representations made to pilot customers should be reviewed with counsel familiar with your jurisdiction and your existing agreements. The right structure depends on specifics that general advice cannot cover.
Executive Imperatives
Executive Imperative: Write down the single claim your new product depends on, in language specific enough that a disinterested reader could judge it, and write down the result that would disprove it. If you cannot do both, you are not ready to spend money on the experiment.
Executive Imperative: Build the evidence ladder explicitly, with budget, time box, and stop condition for each rung from expressed interest through reliable delivery. Advance one rung at a time, and do not let a positive result at one rung authorize commitments appropriate to a higher rung.
Executive Imperative: Separate learning conversations from selling conversations, and audit your early wins to confirm the demand is not dependent on founder persuasion. Pattern of independent paid commitments is the first evidence worth acting on.
Executive Imperative: Before starting any experiment, name what the existing product and existing customers will give up to make it possible, and frame external pilots as experiments rather than commitments. Reputation and attention are the quietest costs and the hardest to recover.