Demand22% of the score
Is there measurable appetite in this category and market?
Every idea gets a 0–100 score from a deterministic engine: measured App Store data in, the same number out every time, no hand-tuning. The lens is deliberately indie: it rewards what a small team can win and monetize, not raw market size. Five sub-scores, one weighted sum, one penalty.
raw mapped absolutely onto 40–100: raw 32 reads 40, raw 72 reads 100, straight-line between. Arithmetic max raw is 92.Interactive
Take an idea straight out of this month’s report and watch its score get rebuilt from its five sub-scores, then open the sandbox and move the engine’s real inputs yourself.
The idle/flow lane is admitting brand-new indie entrants right now, so a solo dev can ship a differentiated theme this month without fighting a moat. WHO/WHY NOW: Colony Flow! debuted at #10 (and #8 in the UK) and Lamar - Idle Vlogger entered at #40 — fresh, small, un-moated apps proving idle demand is open on the free chart (16 of top-100 rotate in weekly).
These five sub-scores are the ones the engine wrote into the report for this idea. The panel on the right rebuilds its score from them, step for step. To move the inputs yourself, open the sandbox.
That 62 is the raw composite. Mapped absolutely onto the 40–100 band (raw 32 → 40, raw 72 → 100), it reaches the report as 85.
The five factors
Each sub-score is built from measured signals, normalized as percentiles across every category × storefront cell: a 70 means “better than 70% of lanes”, never a gut feeling. Percentiles are taken 75% against every cell and only 25% against the idea’s own storefront: grade a market purely on its own curve and a thin one looks as good as a deep one. Each card carries its exact weighted recipe.
Is there measurable appetite in this category and market?
Can a small team actually take the lane from whoever holds it?
Do people pay here, and does the idea's model fit how they pay?
How fast can one developer ship a credible MVP?
Is the lane heating up right now?
…then two multipliers on the finished composite
The same app is not worth the same everywhere. A win in a small, cheaper market is a smaller win, but it is no easier to build for, and the leader is exactly as hard to beat.
Going head-to-head with an entrenched leader is the classic indie mistake. Anchor size already sinks Winnability on its own, so this multiplier is a deliberate nudge on top, not a second full penalty.
Calibration
Raw composites cluster in a narrow band, so the displayed score is a fixed, absolute map of the raw: a raw of 32 reads 40, an exceptional raw of 72 reads 100, straight-line in between. Drag the slider (or hover the line) to read any raw.
Because the map is absolute, the same raw always shows the same score, across categories and across runs. A stronger run reads higher and a weaker one lower, so a run’s median moves with its quality instead of being pinned to one number. (It used to be rank-calibrated, which forced every run’s median to ~62 regardless.)
The pipeline
Scoring is the fifth of six monthly stages. Four are deterministic code; the two in the middle (reading reviews, writing ideas) are the judgement calls. The ideas stage is paired with an integrity judge whose only job is to cut what the strategist can’t back up. The review-reading stage is not: its complaint clusters are one analyst’s tally, and we say so on the stage rather than implying a check that isn’t there.
A month of daily chart snapshots for every category × storefront collapses into one cell of measured signals: rating velocity, grow potential, chart churn, grossing depth, indie proof. Trends are least-squares fits over each app’s daily rating series, not first-vs-last guesses.
Per storefront, the leaders worth studying give up ~200 recent reviews each: top-10 grossing (where the money is) plus the top-10 declining (where users are leaving), pulled from that storefront so Turkish complaints about Turkish charts aren’t missed. Each pull is digested down to the evidence an analyst actually cites: every ≤2★ review, the ≥3★ reviews naming something missing, a small happy sample.
An analyst clusters the complaints by root cause, counts them, keeps verbatim quotes, states each leader’s sample window and build span, and flags gaps that repeat across several leaders. Only leaders past the complaint gate (≥25 recent ≤2★ reviews and ≥25% of the sample) survive as openings. These cluster counts are the analyst’s own tally and are not independently recounted: read them as the analyst’s reading of the reviews, not as verified figures. The measured numbers beside them (review counts, neg share, grossing rank, grow potential) are copied from the data and are exact.
A strategist writes one idea per genuinely distinct opening, each naming who switches, why now, the leader complaint it exploits, and the anchor: the exact charted app it competes with. An integrity judge then cuts any idea whose anchor, cited numbers, or effort claim doesn’t hold up, even if that leaves a market thin. An honest two beats a padded three.
Every surviving idea goes through the five sub-scores, the weighted sum, the giant penalty and the rank calibration above, plus the planning metrics (build time, modeled agency quote, run cost, reach).
The scored Markdown is committed to the repo and mirrored into the database. That mirror is what this report page reads: the numbers you see are the numbers the engine wrote.
Inputs
All first-party: a month of daily chart snapshots across every category and seven storefronts, plus live customer reviews for the charting leaders. No paid data, no download estimates bought in.
| Signal | What it measures | Source |
|---|---|---|
| newRatingsPerMonth | Regression-smoothed monthly rating growth summed across the cell’s free top-100: the demand pulse. | Daily chart snapshots + app stats |
| growPotential | Least-squares acceleration of an app’s daily rating trend (0.5 steady, >0.5 speeding up): is the lane heating up? | Daily app-stats series |
| medianRatingsTop10 | Median rating count of the cell’s top-10 grossing apps: how fortified the money end is. | Top-grossing charts |
| appsOver20kRatings | Number of top-10 grossing apps with >20k ratings: proof the category monetizes at all. | Top-grossing charts |
| avgWeeklyChurn | Average weekly turnover of the free top-100 across the month: fluid charts admit newcomers. | Daily chart snapshots |
| appsUnder5kRatings | Top-10 grossing apps with <5k lifetime ratings: proof small apps earn here. | Top-grossing charts |
| rating drop | The leader’s lifetime average vs. current-version average: a live quality slide. | Lookup API |
| neg share | Share of the leader’s recent reviews at ≤2★ (gates the Review-opportunities section). | Live customer-reviews feed |
Read this before you trust a number
The formula above is honest about what it computes. It is not, on its own, honest about what it misses, so here are the three places the number is thinner than it looks. Every figure is measured against this month’s run.
Mining reviews is the slowest stage in the pipeline, and it is what gives an idea its wedge: the specific broken thing a new app exploits. But it reaches the score through exactly one term, the leader is hurting slice of Winnability, capped at 12.2 of 100 points, and damped further when the leader is too big to take. Against a giant it bottoms out near 4.3.
An idea with damning reviews at a huge leader still scores modestly. That is the lens working, not failing: the fight is real but it is uphill. Read the score as “how winnable”, never as “how strong is the evidence”.
36 of 259 ideas this month carry anchor: none, meaning the strategist named no single charted app to take the lane from. Those ideas cannot use the real mechanism: Winnability falls back to a flat 0.55 beatable prior instead of a real leader’s measured rating count.
Their Winnability is a default, not a finding. Trust an anchored score far more than an unanchored one, and check which app it names before you read the number.
Demand pays for a big anchor (a big leader proves the market is real). Winnability pays for a small one. They read the same number in opposite directions, so no idea can max both: the arithmetic ceiling is a raw 92, reachable only at a corner that cannot exist, and the best real idea this month scored raw 62.
A mid-70s score is not a mediocre idea. It is close to the practical top of the scale. Judge an idea against the month’s field, never against 100.