raphaelvallat/pingouin: v0.7.0
Abstract
This is a major release with many bugfixes, several of which silently returned incorrect results. We strongly recommend all users upgrade. It also brings new features, large speed improvements, and new minimum versions for Python and all dependencies. Bugfixes — incorrect results mixed_anova: the Greenhouse-Geisser corrected p-value of the within factor did not match the F-value and degrees of freedom on the same row, and epsilon and Mauchly's test were computed from the total covariance matrix instead of the pooled within-group covariance. All values now match R, SPSS and JASP, and the interaction also gets a corrected p-value. Reported eps, W_spher, p_spher and p_GG_corr change whenever group means differ across levels of the within factor. (#525) epsilon, sphericity and rm_anova: fixed the epsilon of the interaction in two-way repeated measures designs where both factors have more than 2 levels. Mauchly's test for the interaction is now supported and is reported for all effects in two-way rm_anova. (#527, #528) sphericity: the John, Nagao and Sugiura (method="jns") test statistic was inverted, so sphericity was rejected in 100% of simulations under the null hypothesis (now ~5%). (#531) linear_regression: standard errors and p-values depended on the scale of the predictors. For example, multiplying a predictor by 1e-8 changed its p-value from 0.35 to 0. The same issue was fixed in logistic_regression, as well as in the LMG relative importance (relimp=True), which no longer summed to the model's R² for small-scale predictors. (#520, #523, #540) rcorr: with padjust, the diagonal and lower triangle of the matrix were included as fictitious zero p-values in the multiple comparison correction, making FDR correction severely liberal. (#521) ptests: the padjust argument was ignored, and the Bonferroni correction was always applied. (#536) anova and ancova: unused levels of a categorical factor silently corrupted the SS, DF and F-values. anova now also raises a ValueError when a combination of the between-subject factors has no observation, which previously returned a negative SS and F for the interaction. (#529, #531) corr and pairwise_corr: one-sided p-values of the Kendall correlation used the Pearson t-approximation instead of the exact test. (#534) compute_effsize: one-sample effect sizes (scalar y) ignored eftype and always returned Cohen's d. (#534) power_anova and power_rm_anova: solving for the effect size returned NaN whenever eta-squared was above 0.5. (#534) pairwise_tests: when a global rounding option was set, all pairs but the first were rounded before the multiple comparison correction, giving wrong p_corr. More generally, rounding options no longer leak into internal computations in any function. (#536) mediation_analysis: a binary mediator was fitted with a linear model whenever another mediator was continuous. (#531) chi2_independence: rows with missing values were counted in the sample size (wrong Cramer's V and power), and Yates' correction over-corrected cells close to the expected count. (#537) circ_rayleigh and circ_vtest: missing angles were counted in the sample size. (#537) logistic_regression: estimates were not fully converged with the default solver (off in the 3rd-4th decimal). They now match R glm and statsmodels to about 6 digits. (#526) welch_anova did not drop missing values (slightly wrong np2), and homoscedasticity returned NaN when a sample had missing values. (#529, #531) qqplot: the data were not standardized when only one of loc or scale differed from the default. (#513) Bugfixes — security, crashes and edge cases ancova, anova (unbalanced or 3+ factors) and plot_rm_corr: column names were inserted in a patsy formula and evaluated as Python code. Column names with quotes, commas, parentheses, or integer names now also work. (#537) distance_corr: permutation p-values depended on the platform. (#526) Fixed crashes in ttest with a non-bool paired (e.g. np.True_), logistic_regression with fit_intercept=False, homoscedasticity (Bartlett) with integer data, chi2_mcnemar when every subject switched, and pairwise_corr with mixed-type column labels. NumPy scalars and unsigned integer data are now accepted everywhere. (#526, #529, #531, #536, #537) plot_paired: the boxplot transparency was silently ignored with recent versions of matplotlib. (#513) New features pairwise_tukey and pairwise_gameshowell now accept a list of between-subject factors, in which case all pairs of cells of the interaction are compared. This is equivalent to R's TukeyHSD(aov(dv ~ A * B), which = "A:B"). (#539) rm_anova now reports Mauchly's test of sphericity (sphericity, W_spher, p_spher) for two-way designs, and sphericity now supports two within-subject factors with more than 2 levels each. (#527, #528) plot_blandaltman: new percentage parameter to express the differences as a percentage of the mean, and new symmetric_ylim parameter. qqplot: new line_kwargs and ci_kwargs parameters to customize the fit line and confidence envelope. (#513) pairwise_gameshowell can now be used as a pandas.DataFrame method. (#536) Improvements Many functions are now substantially faster, with identical outputs. Timings below compare v0.6.1 and v0.7.0 on an Apple M1 Max: linear_regression: 209 ms → 128 ms with n = 1,000,000 and 10 predictors. With relimp=True: 6.09 s → 0.11 s (55x) with 12 predictors. With weights: 16 ms → 1 ms with n = 5,000, and a dense (n, n) matrix is no longer allocated (3.2 GB at n = 20,000). (#531, #540) compute_effsize with eftype="cles": 738 ms → 4.2 ms (175x) with 20,000 observations per group, and ~3 GB less memory. (#534) bayesfactor_binom uses the exact beta-binomial distribution instead of numerical integration: 100x faster. (#537) rcorr: 176 ms → 13 ms (13x) with 60 columns. ptests: 61 ms → 4 ms (15x) with 20 columns. (#534, #536) compute_bootci: statistics that accept an axis argument (e.g. numpy.mean) are now vectorized across all bootstrap samples: 54 ms → 19 ms with func=np.mean and 10,000 bootstrap samples. (#538) The formatting of the output dataframe, which runs at the end of every function, is about 20x faster on a 500 × 14 table. (#537) All docstring examples are now tested in the CI. (#526) Breaking changes linear_regression no longer removes duplicate, all-zero or extra constant columns from X. The output keeps one row per input column, a rank-deficiency warning is emitted, and the collinear coefficients are the minimum-norm solution, as in statsmodels. (#540) logistic_regression with fit_intercept=False no longer removes the first non-zero constant column of X, which is the only intercept of the model in that case. (#541) The deprecated gzscore function has been removed. Use scipy.stats.gzscore instead. (#541) Circular functions now raise a ValueError when the angles are not all in [-π, π] or all in [0, 2π], i.e. when they are likely expressed in degrees. (#537) convert_effsize and compute_effsize_from_t now raise a ValueError for 'cohen_dz' and 'cles', which previously returned incorrect values. (#534) Some input checks in the plotting functions now raise ValueError or TypeError instead of AssertionError. (#513) Dependency requirements This version drops support for Python 3.10 and NumPy 1.x (#535). It requires Python >= 3.11 (Python 3.11-3.14 are supported) and: NumPy >= 2.2.2 SciPy >= 1.15.0 Pandas >= 2.3.0 Statsmodels >= 0.14.5 Scikit-learn >= 1.6.1 Matplotlib >= 3.10.1 Seaborn >= 0.13.2 See the documentation changelog for previous releases.