Goal: find one business that works, expecting several to fail first. Method: cheap parallel signal tests, then real-money demand tests on the survivors. Prepared: 15 August 2026 · candidates drawn from the 331-idea scored database
You have 21 candidates clearing the quality bar and budget for roughly 8 tests.
Candidates are nearly free to produce — the loop generates and screens them automatically, costing you no money and no time — so a deep bench is worth keeping. It has genuine option value: when four of the first four tests fail, you want candidates 5 through 12 already screened and waiting rather than starting a research cycle from cold. Breadth also buys diversity of bet type, which is what stops you testing four versions of the same wrong assumption.
But cheap-to-produce is not the same as valuable-at-the-margin. Testing is the bottleneck, and it is the expensive step — real money, real weeks, and the only step that produces truth rather than estimates. Two research runs have already shown the ceiling holding steady at 91.8, so a bigger bench mostly widens your options rather than raising your best one.
So: keep generating, since it costs nothing. Just stop treating bench size as the finish line. The plan below is built to fail fast and fail cheap against the bench you already have.
Running eight $500 tests spends $4,000 and leaves almost nothing to scale a winner. Splitting each test into a cheap signal stage and a real-money stage is far more capital-efficient, because most candidates die at the cheap stage.
| Stage | Spend | Duration | What it answers | How many |
|---|---|---|---|---|
| 1 — Signal | $150 | 5–7 days | Does this buyer exist, can I reach them, and what does a click cost? | 8 candidates |
| 2 — Demand | $350 | 10–14 days | Will they actually pay? | ~3 survivors |
| Line | Amount |
|---|---|
| Stage 1 × 8 | $1,200 |
| Stage 2 × 3 | $1,050 |
| Infrastructure — domains, hosting, Stripe, disclaimers, LLC | $500 |
| AI API credits for fulfilment | $250 |
| Spent finding the answer | $3,000 |
| Left to scale the winner | $2,000 |
Compare that with eight full $500 tests: same number of answers, $1,000 more spent, and nothing left to press the advantage when something works.
What it is: a real landing page with a real price and a real checkout button. When someone clicks Buy, they reach a short form and a message that you'll confirm their order within one business hour. You are measuring intent, not collecting money yet.
What it is NOT: a fake door with no price. Hiding the price is the single most common way these tests produce meaningless results — you learn that people like free things.
Setup, per candidate:
Measure exactly four things:
| Metric | Meaning |
|---|---|
| CPC | Can you afford this buyer at all? |
| Click → landing page engagement | Did the page match the promise? |
| Intent rate — % of visitors who click Buy | The number that matters |
| Cost per intent | CPC ÷ intent rate |
Pre-committed thresholds — write these down before you launch:
| Result | Decision |
|---|---|
| CPC over $20 | Kill. The buyer is unaffordable at your order value. |
| Intent rate under 3% | Kill. The offer does not land. |
| Intent rate 3–7% | Rework the offer once, then retest or kill. |
| Intent rate over 7% and CPC under $15 | Advance to Stage 2. |
| Zero clicks on the ads at all | Kill immediately — no search intent exists. Stop spending on day 2. |
Run four at once. Stage 1 requires no fulfilment capability whatsoever, so there is no reason to run these sequentially. Four simultaneous tests at $25/day is $100/day and gives you demand data on four businesses inside a week.
Only for Stage 1 survivors. Now the checkout actually charges. You fulfil every order by hand.
| Result | Decision |
|---|---|
| Zero sales | Kill. Suggestive, not conclusive at this sample size — but it is what the budget buys. |
| 1–2 sales, CAC above AOV | Continue cautiously. Fix the funnel before spending more. |
| 3+ sales, CAC below AOV | Scale. Move budget from the other tests into this one. |
| Any customer who buys twice | Stop testing others. This is the strongest signal available and it outranks everything above. |
At $8–15 per click, $350 buys 23–44 clicks. Zero sales from 40 clicks is directional, not proof. You are ranking candidates against each other, not establishing statistical certainty. Do not over-read a single test — but also do not keep funding something that produced nothing while another produced buyers.
Test these simultaneously in week one. They are deliberately chosen to be different kinds of bet, so that whatever happens you learn something that generalises.
construction estimating services · takeoff services · outsourced construction estimator · concrete takeoff serviceFTC claim substantiation supplement · health claim substantiation file · supplement advertising compliance review · structure function claim substantiationHOA document review service · condo document review buyer · reserve study reviewCAM reconciliation audit · CAM charges audit tenant · operating expense reconciliation review · lease audit commercial tenantlease audit intent, which runs year-round; the sharp-trigger read comes in Q1. Do not over-read a weak August number.Grant Readiness Scoring (nonprofits pay too slowly for a 14-day test), Bid Leveling (test only after Construction Takeoff proves the buyer), Supplement Claim Substantiation. CAM Reconciliation Audit left this list for the Stage 1 order in R004, taking the slot of Xactimate Supplement Review, which R004 killed on 26 August 2026 — see the Run Log.
You can do all of this yourself, which is the whole reason these businesses were selected.
Do not build the AI fulfilment pipeline yet. Deliver the first ten orders by hand, whatever it costs you in time. The knowledge of what customers actually want in the report is the real asset, and building the pipeline first guarantees you build the wrong one.
You are not trying to be right first time. Assume four of these fail.
| Milestone | Target |
|---|---|
| Candidates through Stage 1 | 8 |
| Candidates through Stage 2 | 3 |
| Businesses producing a paying customer | 1 |
| Businesses producing a repeat customer | 1 |
| Total spend to get there | under $3,000 |
| Elapsed time | 6–8 weeks |
The programme succeeds the moment one customer buys twice. Everything else — database size, verification coverage, rubric scores — was only ever a way of deciding what to put in front of a customer first.
Every candidate here is "upload documents, get analysis back." If frontier models make that trivial and free, the capability stops being worth anything. That is not a reason to avoid these businesses — it is a reason to build them a particular way, starting on day one. Five defences, in rough order of durability:
1. Sell liability, not analysis. A security program written by someone carrying errors-and-omissions insurance who will stand behind it in an examination is a different product from the same text out of a chatbot. The buyer is purchasing someone to be responsible, and a model cannot hold insurance. Carry real E&O, say so on the page, and price it in. This is the single strongest defence and it costs $500–1,000/yr.
2. Own the distribution, not the capability. If you are the vendor a state dealer association lists, or the one a tax-software company points its users to, the model improving is irrelevant — the buyer isn't shopping. Channel relationships survive capability shifts. Spend disproportionate effort here.
3. Accumulate inputs a general model cannot have. Every job produces something proprietary: regional pricing actuals, what a specific examiner asked for last time, which clauses a particular franchisor actually concedes. Keep it structured from order one. After 200 jobs it is a real asset; a frontier model starts from zero on it.
4. Sell the outcome, not the document. "A WISP" is commoditisable. "You stay compliant — annual refresh, vendor reassessments, and I show up if you are examined" is a retainer. Move up the stack as soon as the one-off sale is proven.
5. Choose buyers who will not do it themselves. A tax preparer with 300 clients and thirty years in practice is not going to learn to prompt an LLM, however cheap it gets. Free capability is irrelevant to a buyer who will not use it. This is a selection criterion — apply it when picking the next candidate off the bench.
How the current picks score on this. FTC Safeguards holds up best: the buyer is explicitly purchasing accountability, and defences 1, 4 and 5 all apply cleanly. Grant Readiness sits in the middle — the rubric knowledge is proprietary-ish, but a motivated nonprofit could self-serve. Construction Takeoff is most exposed, being the most purely computational; its defence is almost entirely #2 and #3, so if you pursue it, treat trade-association distribution as the actual business rather than a channel.
A reasonable counter-position: none of these need to last ten years. If one throws off $120k a year for three years on a $3,000 test spend, that is a good outcome even if it is eventually competed away. Time-boxing the bet is a legitimate answer to commoditisation risk, as long as it is a decision rather than an accident.