How to validate an AI app idea before you build it
The short answer
1
question: what does a wrong answer cost
Validating an AI product means testing tolerance for wrong answers, not excitement about right ones. Ask what an error costs the person who acts on it. A demo that impresses people who then keep doing the task by hand has told you nothing, because trying something once is not the same as relying on it.
This applies where the model produces something a person acts on. If the output is entertainment or a first draft nobody ships unread, error tolerance is high and the harder question is whether anyone returns.
01
Five things to check before you build
AI ideas attract more enthusiasm per conversation than any other category right now, and that enthusiasm is the least reliable signal in this document. These checks exist to get underneath it.
- 01
What a wrong answer costs the person who acts on it
This is the question the whole idea turns on. A wrong suggestion in a brainstorming tool costs a second of attention. A wrong figure in a filing costs a penalty and somebody's job. The second case needs a reliability current models may not reach, and no amount of prompt work closes that gap. Get the answer in consequences, not adjectives.
- 02
Whether they would check the output anyway
If a careful person would verify every result before using it, you have not removed the work, you have added a review step to it. That can still be valuable when reviewing is faster than doing. Ask them to estimate both, because a tool that saves nothing after verification is a demo rather than a product.
- 03
What happens when the model provider ships this
A meaningful share of AI product ideas are a system prompt away from being a feature of something the buyer already pays for. Ask what tools they already use that have started adding AI features, and what would happen to your product if the nearest one added this next quarter. An answer of nothing much is a finding.
- 04
What it costs you to answer one query
Unlike ordinary software, serving a user has a real marginal cost here, and heavy users cost the most while paying the same. Estimate cost per query against the price you have in mind before you interview anyone, so you know which answers about willingness to pay are actually viable and which are pleasant but unaffordable.
- 05
Who does this today, and whether they are unhappy
Most AI ideas automate something a person currently does and is often fine at. The buyer may be perfectly satisfied with that arrangement. Find out who holds the task now, whether anyone has complained about it, and what the person doing it thinks. Displacing a competent human is a much harder sell than filling a gap nobody covers.
Ask what a wrong answer costs them. If the honest answer is a lot, enthusiasm for the demo is not evidence.
02
What counts as evidence here
The pattern to watch for is the second week. Almost everyone will try an AI tool once. What matters is who was still using it after the novelty wore off, and for AI specifically that gap is unusually wide.
- A stated consequence for a wrong answer, in money, time or liability
- Someone describing an AI tool they adopted and then quietly stopped using, and why
- An estimate of how long verifying the output would take them
- Evidence they already pay for the manual version of this work
- A specific account of the last time this task went wrong without AI
03
Questions founders ask next
How do I validate an AI idea when people are excited about anything with AI in it?
Stop asking about your product and ask about their last month. What did they try, what did they keep, and what did they abandon after a week. Adoption history for AI tools is unusually informative right now precisely because so many people have tried several and kept almost none, so the pattern of what survived tells you more than any reaction to your idea.
Is a thin wrapper around a model a real business?
Sometimes, but not because of the model. It works when the value sits in something the model does not give you, such as access to data nobody else has, a workflow that has to integrate with systems the provider will never touch, or a regulated context where somebody has to be accountable for the output. Test which of those you have before building, because the wrapper itself is not defensible.
How accurate does my AI product need to be?
That is set by the cost of an error, not by a benchmark. Where a mistake is cheap and visible, people tolerate a surprising amount of wrongness because they catch it themselves. Where a mistake is expensive or invisible until later, the bar rises steeply and often past what is currently achievable. Establish the consequence first, then judge whether you can meet it.
Should I build the AI part before validating?
No, and this is the category where that mistake is most expensive. You can test the value of the output by producing it manually for a handful of people, which is slow and does not scale and is exactly the point. If nobody changes what they do when the answer is correct and delivered by a human, the model was never the missing piece.
How do I test willingness to pay for an AI tool?
Anchor on what the manual work costs rather than on what other AI tools charge. Comparing against other AI pricing imports assumptions from products with different unit costs and different buyers. The hours currently spent on the task, at the rate of whoever spends them, is the number that survives contact with a budget conversation.
What if my AI idea already exists?
That is usually good news, because it means somebody validated the demand for you. The question becomes why the existing tools are not enough for the specific people you would serve. Find users of the closest competitor and ask what they still do by hand despite paying for it, and treat a shrug as evidence that the gap you imagined is not felt.
What you get back
A look inside a real brief
Not a screenshot of an interface. This is the artifact a round produces, with every score attached to the evidence behind it. It is an extract; the full one goes further on all of it.
The idea
Booking and no-show fees for independent dog groomers
Sources read
37
Confidence
Medium
The signal is there. Go build it.
68% of 19 people confirmed they have this problem.
How we got to that number
Thirteen of nineteen described losing money to no-shows without being prompted, and six already pay for something to reduce it. The score is held below seventy because the free workaround is tolerated rather than hated, which is the switching cost you have not tested yet.
Section 1 of 4
An extract from one worked example. The figures are illustrative, not a customer’s result.
Test it before you build it
Describe what the product decides and who acts on it. Research and the full scored report cost nothing.
Or create an account first. Nothing you type here is lost when you sign up.