Stats Toolkit: free statistical analysis tools for research and teaching
Abstract
StatsToolkit is a free, privacy-first web application that helps students, researchers, and educators make sound statistical decisions without installing any software. It runs entirely in the browser: any data a user provides is processed locally on their own device and is never uploaded or stored. It provides five tools: (1) a statistical test advisor that recommends the appropriate test with plain-English reasoning, assumption checks, and step-by-step instructions for JASP, SPSS, R, and Python; (2) an a priori sample-size / power calculator for two-group, paired, correlation, one-way ANOVA, chi-square, and repeated-measures / mixed (split-plot) ANOVA designs, using exact noncentral-F computations for the ANOVA designs, with partial eta-squared to Cohen's f conversion, Greenhouse-Geisser sphericity adjustment, an attrition buffer, and a ready-to-paste methods sentence; (3) a reliability and agreement tool, in which a guided questionnaire leads to the appropriate statistic for the study design (the intraclass correlation coefficient in all six McGraw & Wong forms, standard error of measurement, minimal detectable change, typical error as a coefficient of variation, Cohen's kappa (unweighted or weighted), Fleiss' kappa, or Bland-Altman limits of agreement for method comparison) and turns each answer into a written justification, with software instructions, a data-layout template, and an in-browser calculator for Excel or CSV files; (4) data-science helpers that recommend a chart, a first-choice model, and an evaluation metric with runnable Python snippets; and (5) a teaching data simulator that pools published summary statistics (Cochrane method) into a synthetic practice dataset. Effect-size and power calculations follow Cohen (1988) and standard methods: exact noncentral-F for the ANOVA designs, and normal approximations with Guenther's small-sample correction for the t-tests and Fisher's z for correlation. Reliability statistics follow McGraw & Wong (1996), Weir (2005), Hopkins (2000), Bland & Altman (1986, 1999) and Fleiss (1971), and were checked against published worked examples (Shrout & Fleiss, 1979; Fleiss, 1971) and independent implementations (pingouin, statsmodels). The tool is intended as an educational aid; users should confirm outputs against established software. Changes in version 1.1: added the reliability and agreement tool. Corrected the version 1.0 description of the power calculations: the t-test and correlation calculations use normal approximations rather than exact noncentral-t computations, and they had not been cross-validated against G*Power. The calculations themselves are unchanged. Live application: https://statstoolkit.com