Chi-Square Test Analysis User Guide

Comprehensive step-by-step documentation for performing Chi-Square Goodness-of-Fit and Test of Independence on categorical datasets in DATES.

1. INTRODUCTION

The Chi-Square Test module in DATES evaluates categorical distributions and cross-tabulated contingency tables to assess whether observed frequency counts conform to expected theoretical proportions or whether two categorical variables exhibit statistically significant association.

Supported Chi-Square Modes:

2. AVAILABLE OPTIONS & SETTINGS

The top toolbar header and sidebar panel allow you to configure test mode, categorical variable selection, significance level ($\alpha$), effect size threshold, and numerical rounding:

Control / Parameter Description Why it is used When to select / set
Upload Data Uploads your .csv, .xlsx, or .xls dataset into memory. Loads raw observational data and populates categorical column selectors. At the start of every analysis session.
Sheet Selector Selects the active worksheet from multi-sheet Excel workbooks. Ensures calculations run on the correct data sheet. When uploading multi-sheet workbooks.
Analysis Mode Toggles between Test of Independence (two-way cross-tabulation) and Goodness of Fit (single variable). Determines whether frequency expectations are computed across two categorical variables or against a single target vector. Select Independence for two-variable contingency tables; select Goodness of Fit for single categorical distribution checks.
Row Variable (Factor A) Selects the categorical variable representing rows in the contingency table (Independence mode). Defines row categories for frequency grouping. Select categorical column for Row Factor A.
Column Variable (Factor B) Selects the categorical variable representing columns in the contingency table (Independence mode). Defines column categories for frequency grouping. Select categorical column for Column Factor B.
Single Category Variable Selects the categorical variable for Goodness-of-Fit testing. Provides observed category frequency counts for single-variable testing. Select categorical column in Goodness-of-Fit mode.
Significance Level (Alpha) Significance threshold (e.g., 0.05 for 5%, 0.01 for 1%). Establishes the critical rejection region boundary for p-values. Set to 0.05 for standard research or 0.01 for strict confidence requirements.
Effect Size Threshold Sets Cramér's V effect size threshold (Small: w = 0.10, Medium: w = 0.30, Large: w = 0.50). Categorizes the magnitude of association between categorical factors. Select threshold based on scientific domain reporting standards.
Decimals Controls rounding precision (1 to 6 decimal places) in output tables. Formats frequency tables and test statistics for journal reporting. Adjust based on required numerical precision.

3. INPUT DATA FORMAT REQUIREMENT

DATES accepts dataset spreadsheets in standard .xlsx, .xls, or .csv formats. Data should contain categorical text labels or discrete coded factor columns:

Categorical_Dataset.xlsx — Sheet1 Format: Categorical Factor Columns
Subject_ID Treatment_Factor Response_Outcome Severity_Grade
Obs-001Method_APositiveModerate
Obs-002Method_APositiveMild
Obs-003Method_BNegativeSevere
Obs-004Method_BPositiveModerate
Obs-005Method_ANegativeMild
Obs-006Method_BPositiveModerate

4. MATHEMATICAL FOUNDATIONS & FORMULAS

The Chi-Square test compares observed frequency counts (O) in each cell against expected frequency counts (E) under the null hypothesis of independence or equal distribution. Below are the plain text definitions:

Chi-Square Test Statistic Calculation

Formula Description:

Chi-Square = Sum of ((Observed Frequency - Expected Frequency)^2 / Expected Frequency) across all table cells.

Expected Frequency (Test of Independence)

Formula Description:

Expected Cell Frequency = (Row Total * Column Total) / Grand Total

Degrees of Freedom: df = (Number of Rows - 1) * (Number of Columns - 1)

Expected Frequency (Goodness-of-Fit)

Formula Description:

Expected Category Frequency = Grand Total / Number of Categories (under equal distribution null)

Degrees of Freedom: df = Number of Categories - 1

Cramér's V Effect Size

Formula Description:

Cramér's V = Square Root of (Chi-Square / (Grand Total * Minimum of (Rows - 1, Columns - 1)))

Measures the strength of association on a scale from 0.0 (no association) to 1.0 (perfect association).

5. STEP-BY-STEP WORKFLOW

  1. Upload Dataset: Click the Upload Spreadsheet panel in the sidebar to upload your .csv or .xlsx file.
  2. Select Worksheet: If using a multi-tab workbook, select your active data sheet from the dropdown menu.
  3. Choose Analysis Mode: Select Test of Independence for two-variable contingency tables or Goodness of Fit for single-variable distributions.
  4. Map Categorical Variables:
    • For Independence: Assign Row Variable and Column Variable.
    • For Goodness of Fit: Assign the target Category Variable.
  5. Set Alpha & Effect Size: In the top header toolbar, set your Alpha level (5% or 1%) and Cramér's V effect size threshold (Small, Medium, Large).
  6. Run Analysis: Click the bold RUN ANALYSIS button in the sidebar.
  7. Inspect Contingency Tables & Visuals: View observed frequency tables, expected frequency matrices, residual heatmaps, bar charts, and text interpretation.
  8. Export Reports: Download results as Excel workbooks (.xlsx), Word documents (.docx), PowerPoint presentations (.pptx), or publication-grade PNG images.

6. SAMPLE RESULTS & INTERPRETATION

Below is an example of an output summary table generated for a Chi-Square Test of Independence between two categorical factors:

Chi-Square Test Summary Alpha = 0.05 | Contingency Matrix
Comparison / Variables Chi-Square Statistic df p-Value Cramér's V Effect Magnitude Association Status
Treatment_Factor vs Response_Outcome 12.450 2 0.0020 0.342 Medium Association Reject H0 (Significant Association)
Treatment_Factor vs Severity_Grade 1.180 2 0.5543 0.082 Negligible Fail to Reject H0 (Independent)

How to Read the Output:

7. IMPORTANT NOTES & BEST PRACTICES

Minimum Expected Cell Frequency Assumption

Chi-Square tests require that no more than 20% of expected cell frequencies are less than 5, and all expected cell frequencies exceed 1. If small sample sizes lead to low expected frequencies, consider collapsing adjacent categories or running Fisher's Exact Test.

Categorical Data Integrity

Ensure that all observational units in your dataset are independent (each subject contributes to exactly one cell in the contingency table). For repeated categorical measurements, use McNemar's test instead of standard Chi-Square.

Cite DATES in Research Papers

If you use the DATES Chi-Square Test module in published scientific research, please cite it as follows:

@software{dates_app_2026, author = {DATES Development Team}, title = {DATES: Data Analysis and Trial Evaluation System}, year = {2026}, url = {https://dates-app.org}, note = {Basic Statistics — Chi-Square Test Module} }