Skip to content

Author

Sunny Singh

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Sep 2026

Predicting Startup Outcomes Using Explainable Machine Learning and Y Combinator-Inspired Feature Engineering

Early-stage startup success is notoriously difficult to predict, yet the stakes for getting it right - whether for investors, accelerators, or founders themselves - are substantial. Most existing machine learning approaches lean heavily on generic features from Crunchbase or PitchBook, and in doing so tend to miss a category of arguably more informative signals: the domain-specific structural properties of elite accelerator programs like Y Combinator (YC). This paper introduces a YC-Inspired Feature Engineering (YIFE) framework that incorporates batch cohort timing, founder technical depth (proxied via public GitHub activity), team composition, industry category, and geographic cluster alongside conventional funding and operational variables. We evaluate five classifiers - Logistic Regression, Random Forest, XGBoost, Support Vector Machine, and a Multilayer Perceptron - on a curated dataset of 4,323 YC-funded companies spanning 2005-2024 using a temporal train-test design. XGBoost with YIFE achieves the strongest performance, with an F1-score of 0.85 and an area under the receiver operating characteristic curve of 0.91 on the held-out W21-S24 cohort, representing a 19-percentage-point F1 improvement over a generic Crunchbase baseline (0.85 vs. 0.66) and an 8-percentage-point improvement over a replicated literature baseline (0.85 vs. 0.77). SHAP (SHapley Additive exPlanations)-based interpretation identifies funding-related variables, batch-year context, and team size among the most influential model features, while prior FAANG (Facebook (now Meta), Amazon, Apple, Netflix, and Google) experience contributes comparatively little within the available founder-profile sample. Because several funding-related predictors may only be observable after the initial accelerator stage and overlap conceptually with the outcome definition, the framework is best interpreted as a domain-contextualized retrospective classification approach for the YC ecosystem rather than a strict ex-ante forecasting tool. These findings suggest the potential value of accelerator-specific feature engineering while motivating future validation under fixed prediction horizons and across non-YC accelerator ecosystems. The analytical code and pipeline are made publicly available to support this and other future work in computational entrepreneurship and venture analytics.

Siddharth Gupta, Pratham Namdev, Shubham Nagar et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.