Skip to main content
NELL Builder
Methodology

How NELL avoids hallucination.

Every report is engineered to be defensible. Here is what that means in practice - the mechanisms baked into how we generate, the failure modes we refuse, and what we are still shipping to make the data even harder to dispute.

Last updated: May 12, 2026

What every NELL report does, by design

Web-grounded retrieval, not training-data guesses

Every claim is sourced from public material we retrieve at generation time - competitor pricing pages, funding announcements, review sites, regulatory filings, founder forums. Training-data recall is not allowed as a primary source.

Disciplined report schemas

Validation follows the 5-pillar structure (problem real · market worth it · can we win · can we monetize · verdict). GTM follows the 5-step plan. Structured output dramatically reduces the surface area where hallucination breeds.

Per-section confidence ratings

Each section in every report carries an explicit confidence label (High / Medium / Low) tied to the strength and recency of the underlying evidence. You see when we are sure, and when we are inferring.

Brutal-truth verdicts

NELL is built to tell you when an idea will not work. Weak evidence, inflated TAM, commodity-risk AI wrappers, missing willingness-to-pay - these are all flagged, not glossed over. Sycophancy is what most users distrust about AI.

Automatic credit refund on failed reports

If a report fails to generate cleanly, is interrupted, or comes back unusable, the credit is automatically returned within seconds. You see a toast notification on the dashboard. We do not charge for output we cannot stand behind.

What we refuse to fabricate

Most AI hallucination shows up in a small set of failure modes. We have designed explicit behaviors for each.

Failure mode
Made-up TAM/CAC numbers when no source exists
What NELL does instead
We label the field "Insufficient public evidence" and explain what we looked for, so you can decide whether to commission custom research.
Failure mode
Citing a URL that does not actually back the claim
What NELL does instead
Citations are inline and clickable. If a citation does not support the claim, flag it and we refund the credit.
Failure mode
Averaging conflicting sources silently
What NELL does instead
When two reputable sources disagree on a number, we surface the conflict and explain which one we lean toward and why.
Failure mode
Generic "this looks great" verdicts
What NELL does instead
Every verdict includes falsifiable kill criteria - specific signals that, if observed, would invalidate the thesis. Strategy you cannot disprove is not strategy.
Failure mode
Recycling stale market data
What NELL does instead
Web retrieval happens at generation time, not from cached training data. Sources carry retrieval timestamps so you know how fresh the evidence is.

Anatomy of a NELL claim

Every load-bearing claim in a report carries four pieces of metadata. If any are missing, do not trust the claim.

  • 1
    The claim itself
    Specific, testable, framed in plain language. Not marketing prose.
  • 2
    Cited source
    Inline link to the URL we retrieved. Clickable. Goes to the page that actually supports the claim.
  • 3
    Confidence rating
    High / Medium / Low. Tied to source quality and the recency of the evidence.
  • 4
    Kill criterion (where applicable)
    The specific signal that would invalidate the claim. Strategy you cannot disprove is not strategy.

What we are shipping next

We publish the roadmap so you can see where the trust bar is going, not just where it stands today.

"Insufficient evidence" labels everywhere

Shipping next

Every quantitative field in every report will carry an explicit "Insufficient evidence" state when sources cannot back the number, instead of a fabricated estimate.

Source-quality tier on every citation

Q3 2026

Citations will be tagged Primary (G2 reviews, regulatory filings, founder forums, GitHub) vs. Secondary (analyst summaries, press coverage) vs. Tertiary (SEO content), with the tier visible on the report card.

Conflict flagging on disputed claims

Q3 2026

When two reputable sources disagree on a number (e.g., TAM = $5B vs $20B), the report will surface both and explain the basis for each.

"Flag this claim" button on every section

Q4 2026

One-click reporting on any claim that looks wrong. Routes to the NELL team for blind audit; a public correction log will be published quarterly.

Falsifiable 7-day tests on every claim

Q4 2026

Every load-bearing claim will be paired with a specific 7-day experiment you can run to disprove it. Validation becomes a hypothesis, not an assertion.

Evidence-drift detection on re-runs

2027

Re-run a report 30+ days later and NELL will surface what has changed (new competitor funded, new regulation, new published research) so the picture stays live, not frozen.

Found a claim that looks wrong?

Tell us. We treat every flagged claim as a real audit, not a support ticket - and we publish a quarterly correction log so the work is visible.

Reply to your report email
Every report we deliver has a "Flag a claim" link in the email footer.
Or book a call

See related: Trust & accuracy FAQ · Privacy Policy ·