The Normality module evaluates whether continuous quantitative measurements follow a Gaussian normal distribution using 10 statistical significance tests, skewness and kurtosis diagnostics, interactive diagnostic plots, and integrated automatic data transformation.
The Normality & Distribution Analysis module provides comprehensive statistical validation of distributional assumptions across scientific disciplines (biology, chemistry, physics, medicine, psychology, engineering, and environmental science).
Most standard parametric statistical procedures—such as Analysis of Variance (ANOVA), Student's t-tests, Pearson correlation, and linear regression—require the underlying population error terms to be normally distributed. Running parametric tests on severely non-normal data can result in inflated Type I error rates or reduced statistical power. The Normality module tests continuous variables across experimental groups, provides skewness and kurtosis metrics, generates diagnostic plots (Q-Q plots, histograms, density curves, box plots, violin plots), and offers automated data transformation to achieve normality.
When to use this module:
The sidebar control panel and top action bar provide settings for test selection, alpha significance thresholds, decimal precision, and automated transformations:
| Control / Option | What it does | Why it is used | When to use / select |
|---|---|---|---|
| Test Selection Dropdown | Selects a specific statistical normality test (e.g., Shapiro-Wilk, Anderson-Darling, Kolmogorov-Smirnov) or runs All Tests simultaneously. |
Chooses the hypothesis testing algorithm used to calculate test statistics and p-values. | Select All Tests for comprehensive multi-test confirmation, or pick a specific test suited to your sample size. |
| Alpha Level Selector | Toggles the significance threshold between 5% (0.05) and 1% (0.01). |
Sets the critical alpha cutoff for rejecting the null hypothesis of normality. | Select 5% for standard scientific studies; select 1% for stringent clinical or industrial quality bounds. |
| Decimal Rounding | Sets output numerical precision selector (0, 1, 2, 3, or 4 decimal places). | Formats calculated test statistics, p-values, skewness, and kurtosis in result tables. | Located in the top header bar; adjust to match target publication standards. |
| Transformation Toggle | Toggles automated transformation mode On or Off and opens the transformation config modal. |
Automatically applies mathematical transformations if data fails normality at the chosen alpha level. | Turn On when preparing non-normal datasets for downstream parametric models. |
| Variables (Traits) Selection | Pill selectors choosing quantitative continuous columns to analyze for normality. | Specifies target continuous measurement variables for distribution testing. | Select one or more continuous numeric variables from your uploaded dataset. |
The module accepts tabular spreadsheets (Excel .xlsx, .xls, or CSV .csv) with the following structure:
Group, Condition, Site) must contain text labels; quantitative variables must contain numeric values.Below is a representative dataset containing generalized grouping factors (Group, Condition) and quantitative measurement variables (Response_Value, Concentration):
| Group | Condition | Replicate | Response_Value | Concentration |
|---|---|---|---|---|
| Group-01 | Control | R1 | 45.20 | 12.40 |
| Group-01 | Control | R2 | 46.80 | 13.10 |
| Group-01 | Control | R3 | 44.90 | 12.80 |
| Group-01 | Treated | R1 | 58.40 | 25.60 |
| Group-01 | Treated | R2 | 61.20 | 28.10 |
| Group-01 | Treated | R3 | 59.10 | 26.40 |
| Group-02 | Control | R1 | 41.50 | 10.20 |
| Group-02 | Control | R2 | 43.10 | 11.00 |
| Group-02 | Control | R3 | 42.00 | 10.80 |
The Normality module provides 10 standard statistical test algorithms to evaluate distributional symmetry and tail behavior:
| Statistical Normality Test | Optimal Sample Size Range | Null Hypothesis (H0) & Decision Rule | Recommended Science Application |
|---|---|---|---|
| Shapiro-Wilk | Small to Medium (n = 3 to 50) | H0: Data is normally distributed. Reject H0 if p-value < Alpha. | Gold standard test for biological, agricultural, and small-sample experimental data. |
| Anderson-Darling | All sample sizes (n ≥ 5) | H0: Data follows normal distribution. Gives higher weight to distribution tails. | Ideal when tail behavior and extreme values are critical (environmental, engineering trials). |
| Kolmogorov-Smirnov | Large samples (n > 50) | H0: Sample distribution matches continuous normal baseline. | General goodness-of-fit comparison against theoretical cumulative distribution. |
| Lilliefors | Medium to Large samples | Modifies K-S test when population mean and variance are unknown. | Useful for continuous physical and chemical measurements with estimated parameters. |
| Jarque-Bera | Large samples (n > 100) | H0: Sample skewness and excess kurtosis match a normal distribution (skewness = 0, kurtosis = 3). | Commonly applied in financial modeling, econometrics, and large sensor streams. |
| D'Agostino-Pearson | Medium to Large (n ≥ 20) | Combines skewness and kurtosis tests into an omnibus chi-square statistic. | Versatile test for clinical, physiological, and industrial quality control data. |
| Cramér-von Mises | Medium to Large samples | Evaluates distance between empirical and theoretical cumulative distributions. | High statistical power for symmetric, continuous physical measurements. |
| Pearson Chi-Square | Large samples binned into classes | H0: Frequency counts across binned intervals match expected normal counts. | Applied to binned discrete or continuous industrial measurements. |
| Shapiro-Francia | Medium to Large samples (n = 5 to 5000) | Modification of Shapiro-Wilk optimized for platykurtic and leptokurtic data. | Effective for heavy-tailed biological and ecological measurement series. |
| Ryan-Joiner | Small to Medium samples | Calculates correlation between observed values and normal scores (similar to Shapiro-Wilk). | Widely used alternative in engineering quality control and quality assurance studies. |
Upon clicking RUN ANALYSIS, the module displays a multi-tab analytical suite comprising main statistical summary tables, interactive visualization plots, and automated text interpretation.
The main summary table reports sample size, mean, median, standard deviation, skewness, kurtosis, test statistics, p-values, and automated conclusion badges:
| Variable | Group | Sample Size (n) | Mean | Std Dev | Skewness | Kurtosis | Test Statistic | p-Value | Normality Status |
|---|---|---|---|---|---|---|---|---|---|
| Response_Value | Group-01 | 12 | 45.850 | 2.140 | 0.124 | -0.342 | 0.968 | 0.7420 | Pass (Normal) |
| Response_Value | Group-02 | 12 | 42.100 | 1.980 | 0.085 | -0.510 | 0.974 | 0.8250 | Pass (Normal) |
| Concentration | Group-01 | 12 | 18.400 | 6.850 | 1.452 | 2.180 | 0.812 | 0.0084 | Fail (Non-Normal) |
The Plots tab generates 5 interactive visual diagnostic plots for each variable:
.xlsx or .csv spreadsheet using the sidebar upload control.Shapiro-Wilk or All Tests) and set the significance level (5% or 1%) in the top bar..xlsx) or DOCX, or download plot figures as publication-ready images or PowerPoint presentations.Unlike standard hypothesis testing where p-value < Alpha indicates a positive finding, in normality testing the null hypothesis (H0) states that the data is normal. Therefore, a p-value > Alpha (e.g., > 0.05) means you fail to reject normality (Data is Normal), whereas a p-value < Alpha indicates significant deviation from normality.
If you use the DATES Normality & Distribution Analysis module for statistical validation in published scientific work, please cite it as follows: