newcoStatusSummaryProformaLeaderboardIndustriesDocumentsDataFull Report
Snapshot 14 Aug 2026Version 2, written after R001. Rankings inside it are superseded by the Run Log and the Leaderboard.

AI Document-Upload → Report Businesses

A 252-Idea Exploration, Scored, Verified, and Ranked

Prepared for: John Mullin · Version 2 — revised after research run R001 Constraint set: under $5,000 capital · solo operator · strong design, marketing, and paid-ads capability in-house · 30-day window to first traction · must be cheap to invalidate

⚠️ What changed in v2 — read this first

The first automated research run tested v1's claims and two of the three headline recommendations moved.

Construction Plan Takeoff: 96.6 → 91.8. Still #1, but the wedge is much narrower than I claimed. V1 asserted that no done-for-you AI takeoff service existed and the only alternatives were DIY software or offshore labor. That is no longer true — and may not have been true when I wrote it. Beam AI (Attentive.ai) sells exactly the service I described as missing: QA-reviewed, 24–72 hour turnaround, accuracy guaranteed within ±1%, 500,000+ takeoffs completed. The business survives on a specific and defensible remainder, not on open field.

HIPAA Security Risk Analysis: dropped out of the top 3 entirely (84.8 → 81.8, now #11). V1 built half its case on "an updated Security Rule is expected to finalize in 2026." That was wrong. The NPRM published 6 January 2025 has not been finalized, the May 2026 target slipped, and no updated rule is in force. I presented a pending regulation as a scheduled event. The standing obligation and OCR's enforcement pattern remain real, but the deadline catalyst — the thing that made it a timely business — does not exist.

Grant Readiness Scoring: 88.4, unchanged, independently re-confirmed. It is now the most robust pick of the three.

New #3: FTC Safeguards Rule WISP (87.3). Entered the database at #4 overall on its first appearance.

Nine of v1's top 20 were killed outright, 8 of them to the free-lead-magnet filter. If you read v1, discard its rankings.


Executive Summary

I built and scored 252 businesses that all share one shape: a customer uploads documents, AI returns a valuable report, they pay once with a satisfaction guarantee, and a subscription extends the relationship.

Three findings matter more than the ranking itself.

1. The best idea is construction preconstruction — but the case is narrower than it first looked. Plan Takeoff & Material Estimate leads at 91.8, three points clear of second. Most criteria point the same direction: high order value, brutal and verified customer pain, repeat purchase, an obvious subscription, trivial regulatory risk, and a buyer who can be reached on Google the day they need it. The exception is incumbency — Beam AI already sells a guaranteed done-for-you takeoff service, so what remains is the occasional bidder it prices out, not the whole market.

2. Research destroyed four of my own top-ten picks. Before verification, Ad Account Waste Audit ranked 3rd and Technical SEO Audit ranked 8th. Both collapsed for the same structural reason: the deliverable is already given away free as a sales lead-magnet. Every PPC agency offers a free audit to book a call. Every payment processor offers free statement analysis to win a switch. Every bookkeeper offers a free books review. You cannot charge $500 for the thing your competitors use as bait. This turned out to be the most valuable filter in the entire exercise, and I did not have it when I started — it emerged from the research.

3. Your stated advantages point at a specific kind of buyer. You can design, market, and run ads yourself. That is worth the most where the incumbent is a tired offshore service with a bad website, the buyer searches with clear commercial intent, and the purchase repeats. It is worth the least in consumer niches where you compete against free tools and one-and-done transactions. The ranking reflects that.

The Top 3

#BusinessScorePriceWhy it wins
1Construction Preconstruction Desk — takeoff and estimate for the occasional bidder91.8$299–$1,499Estimators are among the scarcest roles in a labor-starved industry, and 59% of specialty trades are under $1M revenue — too small for Beam AI's ~$8k/yr per-trade subscription, and unserved on a per-report basis
2Grant Readiness Scoring — score a nonprofit's draft against the funder's rubric before submission88.4$399–$1,499Priced against grant writers at $1,500–15,000; the only pre-submission scoring tools built are for funders, not applicants; dense referral network
3FTC Safeguards Rule WISP — the independent security program covered businesses must have87.3$1,200–$4,000Auto dealers, tax preparers and mortgage brokers are all covered non-bank financial institutions; the IRS requires a WISP for every PTIN holder; MSPs bundle it free, which is exactly why an independent one is the product

Part 1 — Exploration: The Idea Tree

Method

I built the taxonomy top-down from industry → sub-vertical → product, forcing 14 ideas per branch across 18 branches. Forcing a fixed count per branch matters: it stops the exercise from over-indexing on the industries I find easiest to imagine, and it pushes into the long tail where underserved niches actually live. Several strong ideas — royalty statement underpayment audits, demurrage dispute packets, experience-mod worksheet audits — only appeared because a quota forced me past the obvious four.

Every idea was specified along five dimensions so it could be scored consistently rather than vibes-ranked:

The 18 Branches

Ordered by branch average score, which is itself a signal about where this business model works best:

BranchAvgWhat's in it
Finance, Accounting & Tax74.2Missed deductions, merchant fees, nexus exposure, QoE lite, bill audits
Construction & Trades74.1Takeoffs, bid leveling, change orders, submittals, liens, certified payroll
Marketing, Brand & Creative Ops74.0Ad audits, agency audits, review mining, pitch deck teardowns
Government, Grants & Nonprofits72.8Grant readiness, funder match, 990 review, FAR clause mapping
Procurement, RFP & Sales72.8RFP drafting, security questionnaires, SaaS spend audits, battlecards
Consumer Life-Admin71.5Subscriptions, credit disputes, car deals, timeshare exits
Compliance, Safety & Regulatory71.2HIPAA, OSHA, SOC 2, ADA/WCAG, food safety, DOT
Insurance70.6Coverage gaps, claim underpayment, total-loss disputes, experience mod
Energy, Agriculture & Environment70.5Utility tariffs, solar contracts, mineral royalties, water rights
Supply Chain, Logistics & Trade70.0Freight audits, HTS classification, tariff exposure, drawback
IT, Security & Tech Diligence70.0Cloud cost, license compliance, pen test remediation, tech DD
Real Estate — Residential69.8Inspection reports, HOA docs, closing disclosures, tax appeals
Education & Academia69.1Aid letters, IEP review, literature synthesis, proposal review
HR, Employment & Benefits67.4Handbook gaps, classification risk, severance review, I-9 audits
Healthcare & Medical66.7Bill audits, denial appeals, prior auth, payer contracts
Immigration & Visa65.1RFE strategy, EB-1 evidence, PERM audits, pathway analysis
Real Estate — Commercial64.4Lease abstraction, CAM audits, rent roll underwriting, estoppels
Legal & Litigation60.7Contract review, depositions, discovery, demand letters

Read the bottom of that table carefully. Legal (60.7), Real Estate — Commercial (64.4), Immigration (65.1), and Healthcare (66.7) rank lowest — despite containing some of the highest-pain, highest-willingness-to-pay ideas in the database. Two different forces put them there:

This is the rubric earning its keep. A naive "biggest market × highest price" ranking would have pushed you straight into commercial real estate or medical billing, and both would have burned your $5,000 before the first sale.

The interactive tree map lets you explore all 252 by branch, filter by risk level and score, and click any idea for its full specification. ---

Part 2 — Research: The Grading Rubric

Design principle

You gave me six things to optimize for. I kept all six, added four that your constraints imply but you didn't name, and pulled regulatory risk out of the average so its cost stays visible instead of being diluted.

#CriterionWeightA 5 looks likeA 1 looks like
1Near-Term Profitability15%Cash-positive inside 30 days, >85% gross marginMonths of build before a first dollar
2Low Upfront Cost10%Under $500 to a sellable v1Needs $10k+ in data, licenses, or specialists
3Average Order Value13%$800+ per reportUnder $75
4Demand Size & Growth12%Millions of buyers, growing >15%/yrTiny or shrinking
5Underservice / Whitespace12%Buyers openly say nothing good existsCrowded and well-funded
6Social Proof of Pain7%Loud public evidence — Reddit, review sites, pressNo visible complaints anywhere
7Paid-Ads Reachability11%High commercial intent, affordable CPCNo search intent, unreachable
8Word-of-Mouth / Referral Engine9%Buyer or a repeat referrer tells others every timePrivate, embarrassing, one-and-done
9Subscription Attach Potential5%Ongoing need, monthly plan is obviousPure one-off
10Cheap Invalidation Speed6%Provable or killable in <14 days for <$500Needs months and real money to learn anything

Score = (weighted average of the ten, scored 1–5) × 20 − risk penalty, where the penalty is 3.5 points per risk point above 1 on a 1–5 scale. A risk-5 idea therefore loses 14 points — enough to knock a strong idea out of contention without pretending it's worthless.

The four criteria you didn't name, and why I added them

Paid-Ads Reachability (11%). You said you can run ads. That's only an advantage where a buyer types a purchase-intent query. Nobody Googles "commercial lease abstraction service" the week they need one — those deals move through brokers. Contractors absolutely Google "construction estimating service." Same business model, completely different acquisition reality. Weighting this at 11% is what separates ideas you can actually reach from ideas that merely sound good.

Word-of-Mouth (9%). You asked me to prioritize the buyer who spreads the product. I scored this on structural referral, not enthusiasm. The highest scores go where a repeat referrer exists — an inspector who sends every client, a state nonprofit association that lists you, a trade group where twelve contractors share one supplier list. That's compounding distribution. A delighted one-time consumer is not.

Subscription Attach (5%). You want a monthly tier. Weighted low deliberately — forcing a subscription onto a genuinely one-off product is a common way to poison an otherwise good business. It should be a bonus, not a requirement.

Cheap Invalidation Speed (6%). You said you want to invalidate cheaply. This is a real criterion: an idea you can kill in 10 days for $300 is worth more to you right now than a marginally better idea that takes three months to evaluate.

Weights are levers, not laws

The Rubric tab in the spreadsheet drives every score by live formula. Change a weight and all 252 scores recalculate. If you decide AOV matters more than I assumed, or that you'd rather not touch anything above risk 2, re-rank in ten seconds and see what surfaces.


Part 3 — Ranking

The Top 10

All ten are Tier 2 — competitor pricing and a demand statistic sourced — except where noted. Nothing in the database has reached Tier 3, because Tier 3 requires real CPC data and I still have no source for it.

#BusinessScorePriceRiskThe one-line case
1Plan Takeoff & Material Estimate91.8$249–$1,4991Estimators are the scarcest role in a labor-starved industry; Beam AI serves the steady bidder, the occasional bidder is unserved
2Grant Application Readiness & Scoring88.4$299–$1,4991Pre-submission scoring tools exist only for funders, not applicants; alternative costs $1,500–15,000
3Bid Leveling & Comparison87.6$199–$8991Same buyer and same upload motion as #1 — this is SKU 2, not a second business
4FTC Safeguards Rule WISP87.3$1,200–$4,0002Mandated for a huge, search-reachable buyer set; MSPs bundle it, so independence is the product
5Xactimate Scope & Supplement Review86.0$199–$9993Roofers pay eagerly; held back only by public-adjuster licensing, which positioning can solve
6Merchant Processing Fee Audit86.0$149–$7991Real savings, real pain — but you fight against free analysis from every processor rep
7HOA / Condo Document Risk Review84.5$129–$3492Sharp trigger, terrifying downside ($5k–100k+ special assessments), agents refer repeatedly
8CAM / Operating Expense Reconciliation Audit83.1$499–$2,499240% of reconciliations contain material errors; incumbents take 30–50% contingency, flat fee is the wedge
9Supplement & Health Claim Substantiation Dossier82.8$1,500–$7,5003Published price band, FTC penalty $53,088 per violation, ~700 marketers already on notice
10HIPAA Security Risk Analysis81.8$499–$2,4993Standing legal obligation and real enforcement — but the 2026 deadline I claimed in v1 never materialised

Gone from v1's top 10:

A note on how I picked the three to detail

Ranks 1, 3 and 5 are all construction, and 2 and 8 are both nonprofit grants. Presenting "the top 3 scores" would hand you three flavors of the same bet. I treat the construction cluster as one business with a SKU ladder (takeoff → bid leveling → supplement review), then detail the next two distinct businesses: Grant Readiness at #2 and FTC Safeguards WISP at #4.

That means Merchant Processing Fee Audit (86.0) is again passed over, this time for FTC Safeguards (87.3) — which at least now matches the score order rather than contradicting it. The reasoning is unchanged and still worth stating: the merchant audit's core weakness, competing against the free statement analysis every processor rep offers, is structural and unfixable. Every other candidate's weakness is a positioning problem.


🥇 #1 — The Construction Preconstruction Desk

Score 91.8 (was 96.6) · Tier 2 · Risk 1 · $299–$1,499 per report · $299–$799/mo subscription

The thesis narrowed in v2. V1 claimed no done-for-you AI takeoff service existed. Beam AI (Attentive.ai) sells precisely that — QA-reviewed, 24–72 hour turnaround, ±1% accuracy guarantee, 500,000+ takeoffs completed, roughly $8,000/yr per trade. A funded, US-based, accuracy-guaranteed incumbent now occupies the wedge I said was empty. Whitespace score cut from 5 to 3.

What survives is narrower but sharper: Beam prices as an annual per-trade subscription aimed at contractors who bid constantly. 59% of specialty trade contractors are under $1M in revenue and bid occasionally — for them an $8,000 annual commitment is absurd, and there is no credible per-report option between "free DIY software you still have to operate" and "annual enterprise subscription." That gap is the business. It is a real gap, but it is a remainder, not an open field, and you should size your ambition to it.

What it is

A contractor uploads a plan set PDF. Within hours they get back a quantity takeoff and a priced material and labor estimate they can turn into a bid. Later SKUs level competing subcontractor bids and review insurance scopes.

Why this wins, in evidence

The pain is severe, verified, and getting worse. 92% of US construction firms cannot find enough workers. The industry needs roughly 349,000 net new workers in 2026 and 456,000 in 2027 just to hold even. Critically, estimators are specifically named among the scarcest roles, alongside superintendents and project managers — contractors report full pipelines they cannot start because the right people don't exist to run them.

Demand for the deliverable is structurally rising. Backlogs run 8–9 months and 84% of contractors cite aggressive price competition. When margins compress, contractors respond by bidding more jobs to win the same volume — and every additional bid requires another takeoff. Your unit of demand grows precisely when the industry gets harder.

The incumbents leave an obvious hole.

AlternativeWhat it costsWhy it disappoints the occasional bidder
Hire an estimator$80k–120k/yrImpossible to find; absurd for a 9-person shop
DIY AI software — Togal.AI $299/user/mo, Kreo Pro $175/user/mo (entry $35), STACK from $249/user/moMonthly subscriptionStill requires the contractor to do the work. Real accuracy is 90–95% on clean residential plans, 80–90% on commercial, and no vendor's accuracy claim has been independently verified
Beam AI (Attentive.ai) — done-for-you, QA-reviewed, ±1% guarantee~$8,000/yr per tradeThe strongest incumbent. Annual per-trade commitment sized for contractors bidding constantly — not for someone who needs four takeoffs a year
BuildingConnected Pro$5,000–15,000/yrEnterprise-priced, same commitment problem
Offshore estimating services$250–$2,500/estimateSlow, variable quality, poor communication, ugly deliverables

The remaining gap: a credible per-report option for the contractor who bids occasionally. Every serious alternative demands an annual commitment; the only per-report options are offshore shops with 2011 websites. That is where your design and marketing skill converts into pricing power — and it is a narrower claim than v1 made, deliberately.

The buyer is perfectly sized for you. Specialty trade contractors are ~60% of all construction firms. 59% generate under $1M in revenue and average 9 employees — far too small to justify a staff estimator, far too busy to do takeoffs at 10pm after a full field day. There are hundreds of thousands of these businesses.

Word-of-mouth is structural. Contractors are clustered — by trade, by supply house, by local association, by union hall. They compare notes on subs, suppliers, and services constantly. And unlike a consumer buying once, a contractor who likes your work uses you again next week.

Unit economics

LineFigureNote
Blended AOV$499Mix of $299 single-trade / $749 multi-trade / $1,499 full commercial
AI inference cost~$3–8Per plan set
Your review time20–40 minNon-negotiable early — see risks
Gross margin (pre-labor)~97%
Repeat rate1–2×/month for an active bidder
Year-one value of one retained contractor$3,000–$9,0006–18 orders

CAC sensitivity. Rather than assert one number, here's the range that matters. At a $8 CPC:

Landing page conversionLead → saleCACVerdict at $499 AOV
5%25%$640Loses money on order 1, profitable by order 2
8%30%$333Profitable on order 1
12%35%$190Strongly profitable

The middle row is the realistic target for a well-designed B2B page with a free-sample offer. Even the pessimistic row works because the purchase repeats — which is the whole argument for this niche over a one-and-done consumer product.

Your $5,000, allocated

ItemAmount
Domain, hosting, Stripe setup$60
AI API credits (covers ~400 reports)$250
Terms of service, disclaimers, LLC$400
Google Ads — validation phase (days 1–14)$1,200
Google Ads — scale phase (days 15–30)$2,000
Buffer / unallocated$1,090

The 30-day plan

Days 1–5 — Prove you can produce the artifact. Pull 10 real plan sets from public bid boards (free). Produce takeoffs manually with AI assistance. Compare against published estimates where available. You are not building software yet — you are proving the output is good enough that a contractor would pay. If it isn't, you've spent $0 and five days.

Days 6–10 — One landing page, one offer. Single trade, single geography. Concrete or framing in one metro. The offer: "Send us your plans. Get a priced takeoff back in 24 hours. If it's not useful, don't pay." Your satisfaction guarantee is the conversion mechanism — it removes the entire risk of trying an unknown vendor, which is the #1 objection for a service like this.

Days 11–20 — Buy intent, deliver manually. $60–80/day on exact-match search: construction estimating services, takeoff services near me, [trade] estimating outsourcing. Deliver every order by hand. Do not automate. You are learning what contractors actually need in the output — and that knowledge is the real asset.

Days 21–30 — Decide. Kill criteria below.

Kill criteria — decide with your head, not your hope

Signal by day 30Action
Fewer than 5 paying ordersKill or pivot the trade/geo. Not enough intent.
5+ orders but CAC > $600 and no repeatsPivot the offer. Try bid leveling as the wedge instead.
5+ orders, any repeat purchaseContinue. Repeat purchase is the signal that matters most.
10+ orders and 2+ repeat buyersScale. Add trades, add the subscription tier.

Honest risks


🥈 #2 — Grant Readiness Scoring for Small Nonprofits

Score 88.4 · Risk 1 · $399–$1,499 per review · $199/mo subscription

What it is

A nonprofit uploads their draft application plus the funder's RFP. They get back a simulated reviewer score against that funder's actual rubric, with every weak section flagged and specific fixes — before they submit.

Why this wins

The price anchor is extraordinary. Grant writers charge $1,500–$5,000 for a foundation proposal and $5,000–$15,000 for a federal one, with hourly rates of $50–$150 ($150–250+ for senior federal specialists). Against that, a $399 pre-submission score that materially raises win probability is an easy yes. You aren't asking them to replace their grant writer — you're selling insurance on work they've already paid for.

The whitespace is verified and specific. I checked the tooling landscape carefully:

That third gap is your entire business. It's also the highest-anxiety moment in the whole grant cycle — the night before a deadline, with a document they cannot un-send.

Word-of-mouth is the best in the database. Nonprofit executive directors are structurally networked in a way almost no other buyer is: state nonprofit associations, funder convenings, capacity-building cohorts, shared board members. One ED telling five peers at a state association meeting is worth more than any ad you'll buy. This is exactly the "spreads by word of mouth" buyer you asked me to prioritize.

Ads should be affordable here — with one caveat I want to flag rather than hide. Nonprofit keywords don't attract the bidding wars that legal ($9.87 average CPC) and insurance do. But the Google Ad Grants program gives qualifying nonprofits substantial free search spend, and thousands of them bid on adjacent terms with money that costs them nothing. That could push nonprofit-sector CPCs up, not down. I did not verify actual CPCs for these keywords — check this in Keyword Planner before you spend a dollar. It's the weakest assumption in this write-up.

Unit economics

LineFigure
AOV$399 foundation / $999 federal
Cost per report~$4 AI + 25 min review
Gross margin~96%
Subscription$199/mo for unlimited scoring during a grant cycle
Natural repeat4–12 applications/year per active nonprofit

The 30-day plan

Days 1–7. Build the scoring engine against 3 real public RFPs (federal RFPs are public — free training data). Score 5 historical winning applications and 5 losers if you can obtain them. Your rubric must correlate with actual outcomes or you have nothing.

Days 8–14. Landing page + a genuinely free "first section scored free" offer. This is the right lead magnet here because — unlike the ad-audit trap — scoring a specific document against a specific rubric is not something competitors give away.

Days 15–30. Two channels in parallel: (a) $50/day on grant writing help, grant application review, federal grant proposal review; (b) direct outreach to 3–5 state nonprofit associations offering a free workshop. The second channel is slower but is where the compounding is.

Kill criteria

Signal by day 30Action
Fewer than 4 paid reviewsKill. Nonprofits are too slow or too poor at this price.
4+ reviews but no referralsContinue cautiously. The WOM thesis is the reason to be here.
4+ reviews and ≥1 unprompted referralScale hard. Thesis confirmed.

Honest risks


🥉 #3 — FTC Safeguards Rule WISP Build & Gap Report

Score 87.3 · Tier 2 · Risk 2 · $1,200–$4,000 · $249/mo subscription

Replaces HIPAA Security Risk Analysis, which fell to #11 — see the retraction at the end of this section.

What it is

An auto dealer, tax preparer, or mortgage broker uploads their IT inventory, vendor list, and existing policies. They receive a Written Information Security Program — the document the FTC Safeguards Rule legally requires them to maintain — plus a gap report against each of the Rule's required elements.

Why this wins

The covered population is enormous and almost nobody realizes they're in it. The Safeguards Rule applies to "non-bank financial institutions," which is far broader than it sounds: auto dealerships, tax preparers, mortgage brokers and originators, payday lenders, collection agencies, investment advisers, and finance companies. Separately, the IRS requires a written security plan for every PTIN holder — that alone is hundreds of thousands of tax preparers, most of them one- and two-person shops.

Independence is the product — and this is the subtle part. MSPs and IT providers hand out a WISP free to win the managed-services contract. On the surface that looks like kill-filter #1, and I want to be explicit that I checked: it isn't, because the free version is written by the party being assessed. An MSP's WISP documents the MSP's own controls and never concludes the client should replace them. For a dealer whose examiner asks who wrote it, that's a real problem. You're not competing with the free document — you're selling the thing the free document structurally cannot be.

High AOV with an unusually clean sale. $1,200–$4,000 for a mandated document with a named regulator behind it. The buyer isn't deciding whether; they're deciding who.

The buyer is search-reachable and clustered. Auto dealers and tax preparers both search actively for compliance help, and both sit inside dense professional networks — state dealer associations, NATP and NAEA chapters, franchise dealer groups. That's the structural referral pattern the rubric rewards.

Risk is only 2. Producing a security program document is not a licensed activity, unlike claim adjusting or legal advice. Disclaim that you are not certifying compliance and you are on solid ground — meaningfully safer than the HIPAA play it replaces.

Unit economics

LineFigure
AOV$1,200 solo tax preparer · $4,000 multi-rooftop dealer group
Cost per report~$8 AI + 45–60 min review
Gross margin~95% pre-labor
Subscription$249/mo — annual refresh, vendor reassessment, incident-response plan upkeep
RepeatAnnual review is required by the Rule

The 30-day plan

Days 1–7. Build against the Rule's published required elements (16 CFR Part 314) — free and specific. Produce one complete sample WISP. Have someone with security or compliance background review it before you sell.

Days 8–14. Pick one buyer segment. Tax preparers are the better beachhead than dealers: far more of them, much lower sophistication, an unambiguous IRS requirement, and a hard annual rhythm you can time campaigns against.

Days 15–30. $60–80/day on WISP tax preparer, FTC Safeguards Rule compliance, written information security plan IRS. Verify CPC in Keyword Planner first — I have no CPC data for these terms and will not guess.

Kill criteria

Signal by day 30Action
Fewer than 3 salesKill. At $1,200+, three sales is a low bar for a mandated document.
Sales, but buyers say "my MSP already gave me one"Sharpen the independence message before spending more — that objection is the whole game.
3+ sales with any subscription attachScale. Add auto dealers as segment two.

Honest risks

Retraction: what happened to HIPAA

V1 ranked HIPAA Security Risk Analysis third at 84.8, and half of that case rested on a claim I got wrong. I wrote that "an updated HIPAA Security Rule is expected to finalize in 2026" and called it the buying trigger. The NPRM published 6 January 2025 has not been finalized, the May 2026 target slipped, and no updated rule is in force. I presented a pending regulation as a scheduled event.

What survives is real: the standing obligation under 45 CFR 164.308(a)(1)(ii)(A), and 76% of 2025 OCR enforcement actions citing risk-analysis failures. What doesn't survive is the deadline — the thing that made it urgent. Score corrected to 81.8, now #11.

The generalizable lesson, which is now written into the research loop: never score a pending regulation as though it will land on schedule. Regulations slip constantly. A rule creates a buying trigger the day it is final, not the day it is proposed.


Red Team: The Strongest Argument Against Each Pick

I built the case for these three. Here is the best case against them, stated as forcefully as I can — because the failure mode of a report like this is that it only argues one direction.

Against #1 (Construction Takeoff): You have no construction knowledge, and that is disqualifying. This is not a business where a good website wins. A contractor's first question will be "who's doing the estimate?" and "I have an AI" is a worse answer than "a guy in Karachi who's done 4,000 of these." The offshore services you're dismissing as ugly and slow have something you don't: estimators who know that a 2,400 sq ft slab needs a specific amount of rebar and will notice when the number is wrong. Your 30-minute review can't catch errors you lack the knowledge to see. And this argument got stronger in v2 — Beam AI now guarantees ±1% accuracy across 500,000+ completed takeoffs. You cannot make that promise, and the occasional bidder you're targeting is exactly the customer least able to catch your errors. The counter: start in one simple trade, buy a retired estimator's time for QA, and let the guarantee absorb early mistakes. But if you can't stomach learning the domain, this is now the wrong pick.

Against #2 (Grant Readiness): Nonprofits are the worst-paying customers in the economy and you're selling them a nice-to-have. Every dollar they spend on you is a dollar not spent on the mission, and they feel that acutely. Your product is unprovable — they won't know for 6–9 months whether it helped, so you can never build a results-based case study. The dense referral network cuts both ways: one ED saying "we paid $399 and still got rejected" travels just as fast as praise. The counter: the price anchor against grant writers is genuinely strong and the whitespace is real. But the slow-payment risk is why this is #2 and not #1.

Against #3 (FTC Safeguards WISP): The independence argument is a story I told myself, and it may not survive contact with a buyer. I argued this escapes the free-lead-magnet filter because the MSP's free WISP isn't independent. That's a sophisticated distinction — and sophisticated distinctions are exactly what buyers ignore. A dealer who already has a WISP in a folder has no felt problem, and "yours isn't independent" is an argument you have to win in a Google ad against a document that cost them nothing. There's also a demand question I flatly did not resolve: I established the Rule exists and who it covers, but not that anyone is being penalized. "Legally required but never enforced" describes a large graveyard of compliance products. The counter: the IRS PTIN requirement is unusually concrete, and tax preparers have an annual rhythm you can time against. But this is the least-verified idea of the three and the bear case is genuinely open.

Note what just happened to the previous #3. In v1 this slot held HIPAA, and my argument for it leaned on a 2026 Security Rule that never finalized. The red team I wrote for v1 flagged "enforcement statistics prove OCR fines people, not that practices know they're exposed" — and then I recommended it anyway. The same unresolved demand question now sits under FTC Safeguards. I've flagged it again. Resolve it this time before spending money.

The meta-argument against all three: they are all "sell a report to a business that should already have this." That's a real pattern, but it means all three share a failure mode — if businesses are willing to keep not having the thing, none of them work. Your $500 test is the only way to find out. ---

Part 4 — Self-Evaluation

What actually worked

Forcing a fixed count per branch. The quota of 14 per industry pushed me past the obvious ideas into the long tail. Royalty statement underpayment audits, demurrage dispute packets, and experience-mod worksheet audits all scored in the top third and none would have occurred to me in a free-form brainstorm.

Pulling risk out of the weighted average. Folding regulatory risk into the blend would have quietly buried it. As an explicit visible penalty you can see exactly what a licensure problem costs an idea — and overrule me if you disagree.

Verifying before ranking. This was the highest-value step by a wide margin. Four of my pre-research top ten collapsed on contact with evidence, including my #3 pick. Without this step I'd have confidently handed you an ad-audit business that competes against a service every agency gives away free.

What was weak — the honest list

1. False precision. This is the biggest problem. I report scores to one decimal place across hundreds of ideas. That implies a resolution I do not have. A gap of one or two points is noise — I said so explicitly when choosing between the third-place candidates, but the spreadsheet's tidy ranking invites you to trust distinctions that aren't real. Treat the top 15 as a set, not a sequence. R001 proved the point: ideas moved by up to 14.6 points on a single piece of evidence, which is roughly ten times the gap between adjacent ranks.

2. Only ~8% of the database was verified. I researched roughly 20 ideas properly. The other 232 carry scores that look identical in authority but are unverified priors. Given that verification moved ideas by up to 14.6 points, the unverified scores could each be off by that much in either direction. There are almost certainly ideas ranked 40–80 that belong in the top 10, and ideas in the top 20 that would collapse the way the ad audit did.

3. I generated the ideas and I scored them. That's a conflict. No independent adversary ever argued against my picks. Ideas I found interesting to write about plausibly received warmer scores. This is the single easiest thing to fix in a rerun.

4. The "free lead-magnet" filter was applied unevenly. It emerged mid-research and only got applied to ideas I happened to check. I'd bet real money that several unverified ideas in the database have the same fatal flaw and are sitting there with inflated underservice scores.

5. Reachability scores are inference, not data. I never pulled actual search volume or CPC for a single specific niche. I reasoned from general industry benchmarks. Since reachability carries an 11% weight, this is a meaningful soft spot — and it's cheap to fix with Keyword Planner.

6. Survivorship bias in the competitor research — and this one cuts both ways. I found competitors by searching for them. Companies that market well are visible; niches where nobody markets look "underserved." But sometimes nobody markets because the business doesn't work — the buyers won't pay, or the deliverable can't be made good enough. I likely scored some genuinely bad ideas as attractive whitespace. An absence of competitors is ambiguous evidence and I treated it as positive evidence.

7. Salience bias toward well-documented industries. The construction labor shortage is heavily covered in the press, so I found abundant, vivid, quantified evidence for it. A quieter niche with an equally good opportunity but less written about it would score lower purely from evidentiary thinness. My #1 pick may be partly an artifact of good press coverage. I still believe the case, but you should know the mechanism.

8. I let you set the count at 250 without pushing back. You asked for 14 per branch and I complied. In hindsight I should have argued that 80 ideas verified properly beats 250 ideas mostly unverified. The extra 170 rows added surface area, not decision quality. I gave you volume when you needed confidence.

9. Zero contact with reality. No landing page, no fake-door test, no ad data, no conversation with a single contractor, nonprofit director, or practice manager. Everything here is desk research. The most valuable next hour is not more analysis — it's one phone call with a subcontractor.

10. I scored a pending regulation as though it were a scheduled event. (Added in v2 — found by run R001.) I wrote that an updated HIPAA Security Rule was "expected to finalize in 2026" and made that the buying trigger for my #3 pick. It didn't finalize. The NPRM from January 2025 is still pending and the May 2026 target slipped. Regulations slip constantly and I treated a proposal as a date. This is now a standing rule in the research loop: a rule creates a buying trigger the day it is final, not the day it is proposed.

11. I asserted whitespace that a single competitor search would have disproven. (Added in v2.) V1's headline claim was that no done-for-you AI takeoff service existed. Beam AI has completed 500,000+ takeoffs with a published ±1% accuracy guarantee. This is the survivorship-bias failure from item 6 happening in the opposite direction — I searched for category information (pricing benchmarks, labor statistics) and never ran the one obvious query: "who sells exactly this?" That query is now mandatory before any idea reaches Tier 2.

The prompt I'd write instead

If we rerun this, restructure it as a funnel with a verification gate, not a single pass:

Stage 1 — Generate wide, cheap (target 150–200). Build an industry → sub-vertical → product tree. For each idea capture only: buyer, input documents, output report, and the price of the existing manual alternative. Do not score yet.

Stage 2 — Apply kill-filters BEFORE scoring. Eliminate on any single disqualifier: (a) the deliverable is already given away free as a lead-magnet; (b) it requires a professional license; (c) demand is seasonal and the season is more than 60 days out; (d) an entrenched incumbent owns the distribution channel; (e) the buyer can't be reached by search intent. Report what died and why — the kill list is as informative as the survivors.

Stage 3 — Score only the survivors (expect 40–60). Use the weighted rubric. Score bands, not decimals: A (85+), B (75–85), C (below). Never report a score I can't defend to the nearest 5 points.

Stage 4 — Verify the top 20 to a fixed evidence standard. Each requires: 3+ named competitors with actual pricing, one quantified demand statistic with a source, one piece of primary complaint evidence (Reddit, G2, forum, review site), and real CPC and search volume from Keyword Planner. No claim survives without a citation. An unverifiable claim gets deleted, not softened.

Stage 5 — Red-team the top 5 with a separate agent that has never seen my reasoning and is instructed to kill each idea. Only survivors get written up.

Stage 6 — Write up 3, each with explicit kill criteria and a $500 test to run this week.

Three specific instruction changes I'd add:


Part 5 — The Research Loop System

The goal is a system where research runs continuously and cheaply, the database improves every week, and you spend your time on the three ideas that matter instead of re-deriving the other 249.

The loop

   ┌──────────────────────────────────────────────────────────┐
   │                                                          │
   ▼                                                          │
GENERATE ──► SCREEN ──► VERIFY ──► TEST ──► DECIDE ───────────┘
 (cheap)    (kill-     (expensive,  (real    (scale / pivot /
            filters)   top 20 only)  money)   kill / re-enter)

The discipline is that each stage costs more than the last, so each stage must eliminate aggressively. Most research projects fail by verifying everything shallowly instead of killing most things instantly and verifying a few things properly.

Stage 1 — Generate (weekly, automated, ~free)

A scheduled run each Monday that adds 5–10 new ideas from live signal rather than from imagination:

Stage 2 — Screen (automated, instant)

Every new idea passes the five kill-filters before it earns a score. This is the cheapest, highest-leverage step and it's the one I skipped this round:

  1. Is this deliverable given away free as a lead-magnet? → kill
  2. Does it require a professional license? → flag risk 4–5
  3. Is demand seasonal and the season >60 days out? → park with a wake-up date
  4. Does an entrenched incumbent own the distribution channel? → kill
  5. Can the buyer be reached by search intent? → if no, kill

Stage 3 — Verify (weekly, top candidates only)

Fixed evidence standard, no exceptions. An idea is not "verified" until it has:

RequirementSource
3+ named competitors with real pricingTheir pricing pages
1 quantified demand statisticIndustry report, government data
1 piece of primary complaint evidenceReddit, G2, forums, review sites
Real CPC + monthly search volumeGoogle Keyword Planner
A stated reason the gap existsYour own reasoning, written down

That last row is the antidote to the survivorship problem. If you can't articulate why nobody serves this niche, you haven't verified it — you've just failed to find the competitor.

Stage 4 — Test (you, with real money)

Every idea that survives verification gets the same $500 / 14-day test:

  1. One landing page, one offer, one geography (day 1–2)
  2. $30/day exact-match search ads (day 3–12)
  3. Deliver every order 100% manually
  4. Measure only three things: cost per lead, lead→sale conversion, and whether anyone comes back

Repeat purchase is the single most informative signal. One repeat buyer tells you more than fifty visitors.

Stage 5 — Decide

Pre-committed thresholds, written before you see the data so you can't rationalize:

ResultDecision
CPL < $50 and ≥1 saleScale — increase budget 3×
CPL < $50, no salesFix the offer, not the traffic
CPL > $150Kill the channel, try referral/direct
Any repeat purchaseContinue regardless of CAC — you'll fix CAC later
Nothing after $500Kill. Return to the database, next candidate.

How to run it

Three pieces, and I can set all of them up:

1. A weekly scheduled research run. Every Monday morning: hunt new regulatory forcing functions and complaint threads, run the kill-filters, score survivors, and append to the database with sources attached. You wake up to 5–10 new pre-screened ideas and a note on anything that changed for your current top 3.

2. A live dashboard. A page you re-open any time showing the current top 20, which are verified vs. unverified, what changed this week, and the live status of whatever test you're running.

3. A reusable skill. The rubric, the kill-filters, and the evidence standard packaged so any future idea gets evaluated the same way — including ideas you bring to it. Type the idea, get it scored and screened consistently instead of re-explaining the framework each time.

The most important thing in this document

Everything above is desk research, including my confident-sounding #1 pick. The rubric is a tool for deciding what to test, not a substitute for testing.

The highest-value hour you can spend this week is not reading this report again. It's calling three subcontractors and asking who does their takeoffs, what it costs, and what annoys them about it. Thirty minutes of that will teach you more than my remaining 249 ideas.


Sources

Construction & estimating

Grants & nonprofits

HIPAA & compliance

Ideas that were downgraded on research

Other verticals referenced in the database


Version 2, revised after research run R001. Database as of 11 Aug 2026, 18:21: 331 ideas · 279 active, 51 killed, 1 parked · 134 (40.5%) at Tier 1 or above · 31 at Tier 2 · none yet at Tier 3. Scores are structured judgement, not measurement — R001 moved them by up to 14.6 points on evidence.