WIREGENT RESEARCH

What 6,080 startup outcomes can and cannot tell an investor

A study of the only public startup dataset that carries outcome labels, and an honest account of what it does and does not prove.

Published 26 July 2026 · Data snapshot 25 July 2026

Before anything else: this dataset is a proxy, not a market

Y Combinator's public company directory is used in this study for one reason: it is the only public startup dataset that attaches an outcome label (active, inactive, acquired, public) to a dated cohort of companies. It is measured here because it is available, and the first finding below explains at length why it is the wrong population to draw general conclusions about startups from.

Every company in this dataset already passed a roughly 1% acceptance filter and received institutional funding before day one of the clock being measured here. That makes the sample post-selection and post-funding at once: these are the survivors of the industry's most selective filter, not a cross-section of people trying to build something. This piece is not about Y Combinator, is not aimed at accelerator applicants, and should not be read as a description of any company's target market. The dataset is used the way an astronomer uses the one star close enough to study in detail: not because it represents the population, but because it is the only one bright enough to measure.

In brief

94.7%of the last 3 years is still labeled “Active”
15×failure-rate spread by team size, and it's leakage
~2×lift on failure odds from category alone
1,966companies, indistinguishable on public metadata
01

Failure is recorded years late

The dataset marks a company's status as Active, Inactive, Acquired, or Public. That status is not audited on a schedule; it changes when someone at the directory notices and updates it. So the rate at which companies are marked Inactive depends heavily on how long ago they started, independent of whether they actually failed sooner.

Measured directly from the data, the share of each cohort marked Inactive or Acquired by cohort age:

Cohort agenMarked InactiveMarked AcquiredFlagged “top company”
0 y6790.7%0.1%0.0%
1 y5701.6%1.4%0.0%
2 y4975.6%2.8%0.0%
3 y49411.5%9.9%0.0%
4 y63212.5%7.8%0.0%
5 y72717.2%10.9%0.3%
6 y43723.3%16.9%0.5%
8 y27724.5%18.1%2.2%
10 y22433.5%20.5%5.4%
14 y14943.6%30.2%8.7%

A company that quietly stops working in month 14 typically stays labeled Active in the directory for years. Which means the most recent cohort looks nothing like the truth:

The 3-year cohort (Summer 2023 to Summer 2026), n = 1,966:

BatchnActiveInactiveAcquiredDead %
Summer 202322018121189.5%
Winter 20242492192288.8%
Summer 2024248236662.4%
Fall 20249490313.2%
Winter 2025167163221.2%
Spring 2025143137332.1%
Summer 2025166163120.6%
Fall 2025148144312.0%
Winter 2026198196201.0%
Spring 2026196196000.0%
Summer 2026137137000.0%
Total1,9661,862 (94.7%)63 (3.2%)41 (2.1%)3.2%

94.7% of the last three years is labeled “Active.” There is no ground truth to backtest against in that window. A model trained on it would learn to predict survival, score around 95% accuracy, and know nothing. Worse, it would look validated, which is the most dangerous outcome a bad backtest can produce.

The rule that follows is simple and non-negotiable for anyone trying to use this kind of data: any backtest needs cohorts at least five years old, and “Active” has to be treated as unknown, never as success.

02

What a properly aged window actually shows

Widen the window to companies old enough to have a real outcome. Winter 2014 to Summer 2021, 16 batches, n = 2,643:

OutcomenShare
Inactive (dead)63324.0%
Acquired48418.3%
Public140.53%
Active (unknown)1,51257.2%
Flagged “top company”501.89%

Three things worth sitting with:

  • About 2% become notable, about 0.5% go public, and that is after a roughly 1% acceptance filter and institutional funding already did their work. The power law here is not a metaphor; it is the actual shape of the only outcome data that exists.
  • 57.2% are “Active,” which is a directory-maintenance label, not a success label. It covers thriving companies, quiet lifestyle businesses, and zombies identically. The majority outcome in this dataset is, plainly, unmeasured.
  • Acquired (18.3%) outnumbers Public (0.53%) by roughly 35 to 1. For a small check, acquisition is the realistic liquidity path, not an IPO.

These are base rates for companies that were already selected by a roughly 1% filter and already received institutional funding. For any population that has not passed that filter and has not received that funding, the base rates are unknown and almost certainly worse. No honest return projection for early, unfunded builders can be constructed from this dataset, or from any other public dataset, because no public dataset has ever measured that population. That is a limitation worth stating outright rather than working around.

03

The strongest-looking predictor in the dataset is fake

The single strongest correlate with failure in the whole dataset is team size:

Team size (current)nDead %Acquired %Top %
16855.9%8.8%1.47%
231359.7%18.2%0.32%
3–535637.6%20.8%0.00%
6–1570222.6%22.5%0.43%
16–5058810.5%16.8%0.68%
51+5273.8%13.3%7.78%

These bands cover 2,554 of the mature window's 2,643 companies. The other 89 carry no usable team-size value and are excluded rather than bucketed as zero, which would have inflated the smallest band with exactly the dead companies the leakage warning below is about.

A 15× spread in failure rate between two-person teams and 51+-person teams. It looks like the best predictor available in the dataset. It is worthless, because the team_size field records the company's size today, not its size when it raised. Companies that were about to fail shrank first. The feature does not predict the outcome; it partly is the outcome, recorded after the fact and pointed backward.

This is not only a problem with someone else's dataset. Any signal that is read after the fact carries the same defect if it is used as though it were known before the fact, and that risk applies directly to the kind of forward-looking signals a diligence or scoring product would want to use:

Forward-looking signalThe same leakage risk
A streak of consecutive proof-of-work checksThe streak stops because the company stopped, not the other way around. Reading it at scoring time predicts the past, not the future.
Spend velocity on a monitored cardSpending stops when the company stops. Same defect.
Uptime or deploy freshness on monitored infrastructureUptime ends with the company. Same defect.
Count of evidence or documentation submitted over timePeople stop uploading evidence when they give up. Same defect.

The fix is a discipline, not a trick: every feature used for scoring has to be frozen at the moment the decision was made, and the label it is judged against has to be measured on a strictly later, fixed horizon (for example, “still active at day 180”). Anything read after the fact and used as if it were known before the fact will produce a backtest that looks excellent and forecasts nothing. It is the single easiest mistake to make in this kind of work, this dataset makes it visible in public data, and it would be naive to assume a private product's own telemetry is immune to it just because it is proprietary.

04

Category is a filter, not an underwriter

Restricting to the aged window (2018 to mid-2021, n = 1,812) and mapping companies into ten broad categories by keyword:

CategorynDead %Acquired %Top %
consumer-web314843.9%14.9%0.7%
decentralized-apis2832.1%14.3%0.0%
commerce-logistics13329.3%7.5%2.3%
devtools18923.8%20.1%0.0%
other2821.4%10.7%7.1%
ops-automation33321.0%18.9%0.6%
robotics-deeptech19718.3%9.1%1.5%
legal-compliance9517.9%15.8%0.0%
vertical-ai28117.1%11.7%0.0%
fintech-payments32413.9%14.8%0.6%
agentic-infra563.6%*10.7%0.0%

* Biased upward, see below. Rows sum to 1,812. An earlier version of this table omitted the “other” category; corrected 26 July 2026.

Companies are categorized using today's tags and description text. A 2019 company that survived by pivoting into AI is tagged as AI now; a company that died in 2020 was never re-tagged. The category feature has partly absorbed the outcome it is supposed to predict, the same look-ahead problem as Section 3, showing up a second way.

Ignoring that biased row, category alone moves failure probability from about 13.9% to about 43.9% around a roughly 21% base rate. That is real, and it is weak: roughly a 2× lift in either direction. Two conclusions follow directly:

  • A mandate keyed only on category buys a roughly 2× tilt on risk, and nothing more. It is a filter, not an underwriting model. An investor who believes “I only back developer tools” is adjusting risk by a factor of two on a base rate they do not actually know, not managing it.
  • “devtools” is the most bimodal category in the table, with both the highest mortality among the large categories (23.8%) and the highest acquisition rate (20.1%). That shape belongs in front of a founder or backer as a distribution, not compressed into a single “fit” score.
05

The punchline: 1,966 companies, and you cannot tell them apart

Go back to the most recent three years, all 1,966 companies, and ask what an outside observer sees using only what is publicly knowable at a glance: category, team size, region, a one-liner, a live site.

Roughly 46% of them sit in just two categories (ops-automation and agentic-infra). About 63.8% are B2B. The median team is 3 people (25th percentile 2, 75th percentile 6); 37.3% have two people or fewer, 73.6% have five or fewer. 99.6% have a live website. Every one of them has a competent one-liner, because writing one is a condition of being listed at all.

Below is every one of those 1,966 companies, rendered as a single uniform tile: no label, no color, no distinguishing mark. That is deliberate. It is what this dataset actually contains about each of them before an outcome is known: a name, a category, a team-size bucket, a live product, a one-liner. Nothing in that description tells you which roughly 63 of these are already gone and which roughly 41 were already bought. The tiles are identical because, on the information a public directory carries, the companies are functionally identical.

1,966 tiles, 1,966 companies, zero visual difference. That is the honest information content of a public startup directory before an outcome is known.

This is the honest verdict of the exercise, and it runs backward from how the question is usually asked. The dataset cannot tell you who made it. What it demonstrates, with real clarity, is that nothing observable at entry tells you who made it. Category, team size, region, tags, and a one-liner do not discriminate between the company that compounds and the one that stops. They never did. Public startup metadata was never built to answer that question, and no amount of reformatting it changes that.

06

What a valid backtest would require

Not the three-year window above, which Section 1 already disqualifies. A design that could honestly be run and honestly published:

  • Population: batches old enough that labels have had at least five years to mature (here, Winter 2014 to Summer 2021, n = 2,643).
  • Label: Inactive at snapshot counts as failure. Acquired or Public counts as success. Active is excluded as unknown, not counted as success, which shrinks the usable set to roughly 1,131 companies. That is the conservative, correct choice, even though it throws away most of the data.
  • Features, entry-time only: batch, a category derived from the original one-liner text rather than current tags, at-batch team size where it can be recovered, region, and whether the product was live at the time of the batch. Anything that was not knowable on day one is excluded by construction.
  • Model: something simple and inspectable, such as logistic regression or a small gradient-boosted tree. The point of the exercise is calibration, not squeezing out accuracy.
  • What gets published: AUC, a calibration curve, and the confusion matrix at the chosen operating threshold. All of it, not just the flattering parts.

The expected result, stated in advance so it cannot be rationalized after the fact: AUC of roughly 0.60 to 0.65. Public metadata alone should be barely better than a coin flip weighted by the base rate. If a result like this comes back meaningfully higher, the first response should be to go looking for leakage, not to celebrate.

Pre-registering a modest expectation like this is itself a piece of methodology. It is also the honest way to make a claim like “public metadata predicts pre-seed failure at roughly this AUC; an evidence-based signal, measured at decision time on a fixed horizon, adds this much on top,” instead of a claim that simply asserts precision nobody has actually measured.

07

Limitations

Stated plainly, because the caveats are the point of publishing this rather than a headline number.

  1. Right-censoring is the dominant issue. Recent cohorts have not had time to accumulate labels; Section 1 covers this in detail.
  2. status is maintained, not audited. “Active” means “not currently known to be dead,” which is a weaker claim than it sounds like.
  3. team_size, tags, and industry are current values, not values at the time of the batch. All three leak outcome information, as shown in Sections 3 and 4. The 3-year composition figures in this piece describe what the floor looks like today, not what it looked like at the moment of funding, and should not be read as predictive.
  4. The “top company” flag is retrospective, assigned after a company already became large. It is never available at the time a decision would have to be made.
  5. The dataset is admits only, never applicants. It says nothing about the roughly 99% of applicants who were not admitted, which is most of the population anyone actually making decisions about early builders would care about.
  6. Acquisition price, funding raised, and revenue are absent from the dataset. “Success” here is binary and coarse: dead, acquired, public, or unknown. It cannot support anything more granular, such as a return multiple.
08

Method and reproduction

Source. yc-oss.github.io/api/companies/all.json, a public mirror of the Y Combinator company directory, yc-oss.github.io/api/. Snapshot taken 25 July 2026: 6,080 launched companies across 47 batches, Summer 2005 through Winter 2027. Fields used: batch, status, industry, subindustry, tags, team_size, regions, top_company, one_liner, long_description, website, launched_at.

Windows. Batches were converted to a decimal year (Winter = .0, Spring/X = .25, Summer = .5, Fall = .75). The 3-year cohort is batch time between 2023.5 and 2026.5 inclusive (n = 1,966; Fall 2026 and Winter 2027, n = 5, were excluded as not yet started). The mature reference window is 2014.0 to 2021.5 (n = 2,643). The category-outcome table in Section 4 uses 2018.0 to 2021.5 (n = 1,812) to hold the era roughly constant.

Category mapping. A keyword heuristic built for this study: agent and LLM-agent signals, then web3, then hardware and deep tech, then finance, then legal and compliance, then developer tooling (by tag and by a subindustry of “Engineering, Product and Design”), then commerce and logistics, then vertical industries such as health, education, and real estate, then operations and automation, then consumer, with everything left over marked “other.” The ordering of the resulting categories is reliable; the exact percentages should be read as approximate, since this mapping is a research heuristic rather than a canonical taxonomy.

Reproduction. Download the JSON at the source URL above and group by the derived batch time. Every number in this piece traces to that snapshot.