Home
Method 1.1-beta

Methodology

StartupBench evaluates the idea visible from public evidence—not revenue, traction, team quality, or private company data. The website is scored separately.

Final score
S = round(2I + W3, 2)

Each area is a shifted weighted geometric mean. This makes weak criteria matter instead of letting one exceptional criterion fully cancel them out.

G(r, w) = 2[(ri + 1)wi − 1]

i = 1, …, n·ri ∈ [0, 5]·Σwi = 1·G ∈ [0, 10]

Idea score

overall
Problem relevance25%
Differentiation25%
Solution strength20%
Market opportunity15%
Feasibility & scalability15%

Website score

overall
Messaging clarity25%
UX & conversion25%
Visual hierarchy20%
Trust & evidence15%
Technical quality15%

Method 1.1 requires a decoded, usable screenshot for all three website judges. A failed capture or access challenge stops evaluation. Legacy results without verified visual evidence remain readable but are excluded from the board and recent activity. Area calculations use full precision; rounding happens at the final score.

ScrapeBadger retrieves public page content with JavaScript and anti-bot handling. Screenshots are captured separately. Incomplete content or unusable images stop evaluation; capture sources are recorded in the audit. External context comes from ScrapeBadger search results; snippets are not full-page verification. Research gaps are disclosed, not treated as proof of uniqueness.

How a score is produced

01

Evidence profile

The site is read as untrusted source material. Public competitor research adds context; unsupported assumptions are excluded.

02

Three judges per area

Three Luna judges per area use medium reasoning and the same anchored 0–5 rubric. Idea judges do not see visual-design signals.

03

Deterministic math

The median judge score is taken per criterion. The published weights and formula produce the final score.

Anchored ratings

Every criterion uses the same evidence scale. Missing evidence is uncertainty—not an invitation to invent facts.

0No supporting evidence or direct contradiction
1Very weak, vague, or largely unsupported
2Plausible but incomplete or weakly evidenced
3Clear and credible public evidence
4Strong, specific, and well-supported
5Exceptional evidence for the criterion

Research basis

The structure draws on research about opportunity evaluation, multi-item measures, and composite indicators. It does not claim to predict company success. Human validation and reliability testing are still required, so the method remains beta.