Research Data

Free, downloadable datasets with real industry statistics from publicly verifiable sources. All data licensed under CC BY 4.0: use freely, just link back.

📊 Why publish research data? We believe industry benchmarks and statistics should be openly accessible. Every dataset below is curated from public reports, government databases, and market research, properly cited, never fabricated. Use them in your articles, pitch decks, and research. Just include a link back here.

Startup M&A Deal Flow Statistics 2026

Key statistics on startup acquisitions, median deal sizes, time-to-exit, and acquisition multiples. Compiled from public market reports and industry data....

9 data points CC BY 4.0 · CSV + JSON CSV JSON
GitHub Startup Momentum Index →

Weekly ranking of notable startup & dev-tool repos by real GitHub engineering-activity signals. Free dataset, CC BY 4.0, refreshed weekly.

Explore by topic

All Data Pages

The full data library: engineering velocity dataset, momentum index, funding trends, valuation trends, and M&A statistics.

Archive and Reference

Annual archives and quick reference pages for the dataset.

Why This Page Exists

GitDealFlow is a public deal flow signal dataset: 350+ startup GitHub organizations across 15 sectors, refreshed weekly, with breakout teams surfacing 21 to 47 days before their round is announced. This page makes one part of that system legible: what it measures, how it is computed, and how to use it in a live sourcing workflow. The method is published end to end and falsifiable by design, with the working paper on SSRN and the dataset downloadable under CC BY 4.0.

Start Here

A practical read-through of Research Data: the dataset behind this page refreshes weekly across 350+ organizations and 15 sectors, and every figure shown traces to a public GitHub REST API pull. That matters for two reasons. Reproducibility: any number here can be re-derived from primary sources, which is the standard the published methodology sets for itself. Timeliness: engineering acceleration precedes announcements, so this page follows the data cadence rather than the news cycle, and the freshness endpoint always reports the exact pull date.

If Research Data is your entry point, the fastest next steps are fixed: skim the glossary for the three or four terms that anchor the topic, open the research dataset to see the raw weekly snapshots behind the summary numbers, and run one live query against the free momentum checker with a company you already know well. Seeing the signal fire on a familiar name is the quickest way to judge whether code-side sourcing belongs in your own workflow.

One caveat worth stating plainly on Research Data: momentum is a leading indicator, not a verdict. A repository can accelerate for reasons that never become a fundraise, and a quiet quarter does not mean a team is failing. The disciplined use of this page is as one input in a stack, a way to rank where scarce diligence time goes, and a way to notice change early. The methodology page documents every limitation, including the bot filter, the two-period confirmation rule, and the sectors where coverage is thinnest.

Putting Research Data into practice comes down to one discipline: verify against primary data. Every figure on this page is reproducible from the public dataset behind it, which tracks 350+ startup GitHub organizations across 15 sectors and refreshes weekly. Because engineering acceleration is a leading indicator, breakout teams surface 21 to 47 days before a funding round is announced, which is why signal-first sourcing, diligence, and monitoring consistently beat waiting for the announcement databases. The methodology page documents every limitation, including the bot filter, the two-period confirmation rule, and the sectors where coverage is thinnest, so a skeptic can check each claim rather than take it on faith.