Failure is recorded years late
The dataset marks a company's status as Active, Inactive, Acquired, or Public. That status is not audited on a schedule; it changes when someone at the directory notices and updates it. So the rate at which companies are marked Inactive depends heavily on how long ago they started, independent of whether they actually failed sooner.
Measured directly from the data, the share of each cohort marked Inactive or Acquired by cohort age:
| Cohort age | n | Marked Inactive | Marked Acquired | Flagged “top company” |
|---|---|---|---|---|
| 0 y | 679 | 0.7% | 0.1% | 0.0% |
| 1 y | 570 | 1.6% | 1.4% | 0.0% |
| 2 y | 497 | 5.6% | 2.8% | 0.0% |
| 3 y | 494 | 11.5% | 9.9% | 0.0% |
| 4 y | 632 | 12.5% | 7.8% | 0.0% |
| 5 y | 727 | 17.2% | 10.9% | 0.3% |
| 6 y | 437 | 23.3% | 16.9% | 0.5% |
| 8 y | 277 | 24.5% | 18.1% | 2.2% |
| 10 y | 224 | 33.5% | 20.5% | 5.4% |
| 14 y | 149 | 43.6% | 30.2% | 8.7% |
A company that quietly stops working in month 14 typically stays labeled Active in the directory for years. Which means the most recent cohort looks nothing like the truth:
The 3-year cohort (Summer 2023 to Summer 2026), n = 1,966:
| Batch | n | Active | Inactive | Acquired | Dead % |
|---|---|---|---|---|---|
| Summer 2023 | 220 | 181 | 21 | 18 | 9.5% |
| Winter 2024 | 249 | 219 | 22 | 8 | 8.8% |
| Summer 2024 | 248 | 236 | 6 | 6 | 2.4% |
| Fall 2024 | 94 | 90 | 3 | 1 | 3.2% |
| Winter 2025 | 167 | 163 | 2 | 2 | 1.2% |
| Spring 2025 | 143 | 137 | 3 | 3 | 2.1% |
| Summer 2025 | 166 | 163 | 1 | 2 | 0.6% |
| Fall 2025 | 148 | 144 | 3 | 1 | 2.0% |
| Winter 2026 | 198 | 196 | 2 | 0 | 1.0% |
| Spring 2026 | 196 | 196 | 0 | 0 | 0.0% |
| Summer 2026 | 137 | 137 | 0 | 0 | 0.0% |
| Total | 1,966 | 1,862 (94.7%) | 63 (3.2%) | 41 (2.1%) | 3.2% |
94.7% of the last three years is labeled “Active.” There is no ground truth to backtest against in that window. A model trained on it would learn to predict survival, score around 95% accuracy, and know nothing. Worse, it would look validated, which is the most dangerous outcome a bad backtest can produce.
The rule that follows is simple and non-negotiable for anyone trying to use this kind of data: any backtest needs cohorts at least five years old, and “Active” has to be treated as unknown, never as success.
What a properly aged window actually shows
Widen the window to companies old enough to have a real outcome. Winter 2014 to Summer 2021, 16 batches, n = 2,643:
| Outcome | n | Share |
|---|---|---|
| Inactive (dead) | 633 | 24.0% |
| Acquired | 484 | 18.3% |
| Public | 14 | 0.53% |
| Active (unknown) | 1,512 | 57.2% |
| Flagged “top company” | 50 | 1.89% |
Three things worth sitting with:
- About 2% become notable, about 0.5% go public, and that is after a roughly 1% acceptance filter and institutional funding already did their work. The power law here is not a metaphor; it is the actual shape of the only outcome data that exists.
- 57.2% are “Active,” which is a directory-maintenance label, not a success label. It covers thriving companies, quiet lifestyle businesses, and zombies identically. The majority outcome in this dataset is, plainly, unmeasured.
- Acquired (18.3%) outnumbers Public (0.53%) by roughly 35 to 1. For a small check, acquisition is the realistic liquidity path, not an IPO.
These are base rates for companies that were already selected by a roughly 1% filter and already received institutional funding. For any population that has not passed that filter and has not received that funding, the base rates are unknown and almost certainly worse. No honest return projection for early, unfunded builders can be constructed from this dataset, or from any other public dataset, because no public dataset has ever measured that population. That is a limitation worth stating outright rather than working around.
The strongest-looking predictor in the dataset is fake
The single strongest correlate with failure in the whole dataset is team size:
| Team size (current) | n | Dead % | Acquired % | Top % |
|---|---|---|---|---|
| 1 | 68 | 55.9% | 8.8% | 1.47% |
| 2 | 313 | 59.7% | 18.2% | 0.32% |
| 3–5 | 356 | 37.6% | 20.8% | 0.00% |
| 6–15 | 702 | 22.6% | 22.5% | 0.43% |
| 16–50 | 588 | 10.5% | 16.8% | 0.68% |
| 51+ | 527 | 3.8% | 13.3% | 7.78% |
These bands cover 2,554 of the mature window's 2,643 companies. The other 89 carry no usable team-size value and are excluded rather than bucketed as zero, which would have inflated the smallest band with exactly the dead companies the leakage warning below is about.
A 15× spread in failure rate between two-person teams and 51+-person teams. It looks like the best predictor available in the dataset. It is worthless, because the team_size field records the company's size today, not its size when it raised. Companies that were about to fail shrank first. The feature does not predict the outcome; it partly is the outcome, recorded after the fact and pointed backward.
This is not only a problem with someone else's dataset. Any signal that is read after the fact carries the same defect if it is used as though it were known before the fact, and that risk applies directly to the kind of forward-looking signals a diligence or scoring product would want to use:
| Forward-looking signal | The same leakage risk |
|---|---|
| A streak of consecutive proof-of-work checks | The streak stops because the company stopped, not the other way around. Reading it at scoring time predicts the past, not the future. |
| Spend velocity on a monitored card | Spending stops when the company stops. Same defect. |
| Uptime or deploy freshness on monitored infrastructure | Uptime ends with the company. Same defect. |
| Count of evidence or documentation submitted over time | People stop uploading evidence when they give up. Same defect. |
The fix is a discipline, not a trick: every feature used for scoring has to be frozen at the moment the decision was made, and the label it is judged against has to be measured on a strictly later, fixed horizon (for example, “still active at day 180”). Anything read after the fact and used as if it were known before the fact will produce a backtest that looks excellent and forecasts nothing. It is the single easiest mistake to make in this kind of work, this dataset makes it visible in public data, and it would be naive to assume a private product's own telemetry is immune to it just because it is proprietary.
Category is a filter, not an underwriter
Restricting to the aged window (2018 to mid-2021, n = 1,812) and mapping companies into ten broad categories by keyword:
| Category | n | Dead % | Acquired % | Top % |
|---|---|---|---|---|
| consumer-web3 | 148 | 43.9% | 14.9% | 0.7% |
| decentralized-apis | 28 | 32.1% | 14.3% | 0.0% |
| commerce-logistics | 133 | 29.3% | 7.5% | 2.3% |
| devtools | 189 | 23.8% | 20.1% | 0.0% |
| other | 28 | 21.4% | 10.7% | 7.1% |
| ops-automation | 333 | 21.0% | 18.9% | 0.6% |
| robotics-deeptech | 197 | 18.3% | 9.1% | 1.5% |
| legal-compliance | 95 | 17.9% | 15.8% | 0.0% |
| vertical-ai | 281 | 17.1% | 11.7% | 0.0% |
| fintech-payments | 324 | 13.9% | 14.8% | 0.6% |
| agentic-infra | 56 | 3.6%* | 10.7% | 0.0% |
* Biased upward, see below. Rows sum to 1,812. An earlier version of this table omitted the “other” category; corrected 26 July 2026.
Companies are categorized using today's tags and description text. A 2019 company that survived by pivoting into AI is tagged as AI now; a company that died in 2020 was never re-tagged. The category feature has partly absorbed the outcome it is supposed to predict, the same look-ahead problem as Section 3, showing up a second way.
Ignoring that biased row, category alone moves failure probability from about 13.9% to about 43.9% around a roughly 21% base rate. That is real, and it is weak: roughly a 2× lift in either direction. Two conclusions follow directly:
- A mandate keyed only on category buys a roughly 2× tilt on risk, and nothing more. It is a filter, not an underwriting model. An investor who believes “I only back developer tools” is adjusting risk by a factor of two on a base rate they do not actually know, not managing it.
- “devtools” is the most bimodal category in the table, with both the highest mortality among the large categories (23.8%) and the highest acquisition rate (20.1%). That shape belongs in front of a founder or backer as a distribution, not compressed into a single “fit” score.
The punchline: 1,966 companies, and you cannot tell them apart
Go back to the most recent three years, all 1,966 companies, and ask what an outside observer sees using only what is publicly knowable at a glance: category, team size, region, a one-liner, a live site.
Roughly 46% of them sit in just two categories (ops-automation and agentic-infra). About 63.8% are B2B. The median team is 3 people (25th percentile 2, 75th percentile 6); 37.3% have two people or fewer, 73.6% have five or fewer. 99.6% have a live website. Every one of them has a competent one-liner, because writing one is a condition of being listed at all.
Below is every one of those 1,966 companies, rendered as a single uniform tile: no label, no color, no distinguishing mark. That is deliberate. It is what this dataset actually contains about each of them before an outcome is known: a name, a category, a team-size bucket, a live product, a one-liner. Nothing in that description tells you which roughly 63 of these are already gone and which roughly 41 were already bought. The tiles are identical because, on the information a public directory carries, the companies are functionally identical.
[1,966-tile visualization omitted for print. See the 3-year cohort table in Section 1 for the same distribution.]
1,966 tiles, 1,966 companies, zero visual difference. That is the honest information content of a public startup directory before an outcome is known.
This is the honest verdict of the exercise, and it runs backward from how the question is usually asked. The dataset cannot tell you who made it. What it demonstrates, with real clarity, is that nothing observable at entry tells you who made it. Category, team size, region, tags, and a one-liner do not discriminate between the company that compounds and the one that stops. They never did. Public startup metadata was never built to answer that question, and no amount of reformatting it changes that.
What a valid backtest would require
Not the three-year window above, which Section 1 already disqualifies. A design that could honestly be run and honestly published:
- Population: batches old enough that labels have had at least five years to mature (here, Winter 2014 to Summer 2021, n = 2,643).
- Label:
Inactiveat snapshot counts as failure.AcquiredorPubliccounts as success.Activeis excluded as unknown, not counted as success, which shrinks the usable set to roughly 1,131 companies. That is the conservative, correct choice, even though it throws away most of the data. - Features, entry-time only: batch, a category derived from the original one-liner text rather than current tags, at-batch team size where it can be recovered, region, and whether the product was live at the time of the batch. Anything that was not knowable on day one is excluded by construction.
- Model: something simple and inspectable, such as logistic regression or a small gradient-boosted tree. The point of the exercise is calibration, not squeezing out accuracy.
- What gets published: AUC, a calibration curve, and the confusion matrix at the chosen operating threshold. All of it, not just the flattering parts.
The expected result, stated in advance so it cannot be rationalized after the fact: AUC of roughly 0.60 to 0.65. Public metadata alone should be barely better than a coin flip weighted by the base rate. If a result like this comes back meaningfully higher, the first response should be to go looking for leakage, not to celebrate.
Pre-registering a modest expectation like this is itself a piece of methodology. It is also the honest way to make a claim like “public metadata predicts pre-seed failure at roughly this AUC; an evidence-based signal, measured at decision time on a fixed horizon, adds this much on top,” instead of a claim that simply asserts precision nobody has actually measured.
Limitations
Stated plainly, because the caveats are the point of publishing this rather than a headline number.
- Right-censoring is the dominant issue. Recent cohorts have not had time to accumulate labels; Section 1 covers this in detail.
statusis maintained, not audited. “Active” means “not currently known to be dead,” which is a weaker claim than it sounds like.team_size,tags, andindustryare current values, not values at the time of the batch. All three leak outcome information, as shown in Sections 3 and 4. The 3-year composition figures in this piece describe what the floor looks like today, not what it looked like at the moment of funding, and should not be read as predictive.- The “top company” flag is retrospective, assigned after a company already became large. It is never available at the time a decision would have to be made.
- The dataset is admits only, never applicants. It says nothing about the roughly 99% of applicants who were not admitted, which is most of the population anyone actually making decisions about early builders would care about.
- Acquisition price, funding raised, and revenue are absent from the dataset. “Success” here is binary and coarse: dead, acquired, public, or unknown. It cannot support anything more granular, such as a return multiple.
Method and reproduction
Source. yc-oss.github.io/api/companies/all.json, a public mirror of the Y Combinator company directory, yc-oss.github.io/api/. Snapshot taken 25 July 2026: 6,080 launched companies across 47 batches, Summer 2005 through Winter 2027. Fields used: batch, status, industry, subindustry, tags, team_size, regions, top_company, one_liner, long_description, website, launched_at.
Windows. Batches were converted to a decimal year (Winter = .0, Spring/X = .25, Summer = .5, Fall = .75). The 3-year cohort is batch time between 2023.5 and 2026.5 inclusive (n = 1,966; Fall 2026 and Winter 2027, n = 5, were excluded as not yet started). The mature reference window is 2014.0 to 2021.5 (n = 2,643). The category-outcome table in Section 4 uses 2018.0 to 2021.5 (n = 1,812) to hold the era roughly constant.
Category mapping. A keyword heuristic built for this study: agent and LLM-agent signals, then web3, then hardware and deep tech, then finance, then legal and compliance, then developer tooling (by tag and by a subindustry of “Engineering, Product and Design”), then commerce and logistics, then vertical industries such as health, education, and real estate, then operations and automation, then consumer, with everything left over marked “other.” The ordering of the resulting categories is reliable; the exact percentages should be read as approximate, since this mapping is a research heuristic rather than a canonical taxonomy.
Reproduction. Download the JSON at the source URL above and group by the derived batch time. Every number in this piece traces to that snapshot.