The Chi-Square Test of Independence is used for categorical data only.
It allows you to determine whether two categorical variables are associated with one another.
In this context, association means that the distribution of one categorical variable differs across the categories of another categorical variable.
The research design is observational. Therefore, the variables are not manipulated by the researcher.
| Component | General Structure | Example |
|---|---|---|
| Research Question | Is there an association between Variable 1 and Variable 2? | Is there an association between gender and voting behavior? |
| Null Hypothesis | There is no association between the two categorical variables. | There is no association between gender and voting behavior. |
| Alternative Hypothesis | There is an association between the two categorical variables. | There is an association between gender and voting behavior. |
The Chi-Square Test of Independence compares the observed frequencies in a contingency table to the frequencies that would be expected if the two variables were completely unrelated.
If the observed frequencies differ substantially from the expected frequencies, the test may indicate that an association exists between the variables.
The Chi-Square Test of Independence does not require a traditional independent variable (IV) and dependent variable (DV).
Because the design is observational, both variables are treated as categorical variables and the goal is to determine whether they are associated.
Use the Inferential Test Selector to determine which inferential procedure is appropriate for your study.
Below is an example of what your code will look like. Your code will be different because the dataset name, variable name, expected proportions, results, and interpretations will depend on your assigned dataset.
library(readxl)
data2026 <- read_excel(
"C:/Users/John/OneDrive/Documents/AA5221/Datasets/political_data.xlsx"
)
polit_table <- table(data2026$gender, data2026$voting)
polit_table
barplot(
polit_table,
beside = TRUE,
col = rainbow(nrow(polit_table)),
legend = rownames(polit_table)
)
chi_result <- chisq.test(polit_table)
chi_result
chi_result$expected
cramer_v <- rcompanion::cramerV(polit_table)
cramer_v
# A Chi-Square Test of Independence was conducted to determine if there was an association between gender and voting behavior.
# The results showed that there was an association between the two variables, χ²(1) = 7.84, p = .005.
# The association was moderate (Cramer's V = .32).
Packages add additional functionality to R.
| Package | Purpose |
|---|---|
readxl |
Import Excel datasets |
Copy-and-paste the following code into your R Script file.
Run the code once.
install.packages("readxl")
You may see warning messages appear when you install a package. These messages do not necessarily mean that the installation failed. Packages only need to be installed once.
Copy-and-paste the code below into your R Script.
Keep these lines of code in your R Script.
library(readxl)
No output will appear. This simply opens the packages so you can use their functions in RStudio.
Although datasets can be imported entirely through code, many students experience difficulty locating file paths. Therefore, this course uses the point-and-click import method to generate the necessary code automatically.
Once imported, the dataset will appear in a new tab and will also appear in the Environment pane.
RStudio also automatically generates the code used to import the dataset. This code appears in the Console window. Copy the import code and paste it into your R Script. This allows the dataset to be automatically imported the next time you run your script.
The generated code will vary depending on your file location and file name.
DatasetName <- read_excel("filepath")
DatasetName <- read_excel(
"C:/Users/John/OneDrive/Documents/AA5221/Datasets/DatasetName.xlsx"
)
A frequency table summarizes how many observations belong to each category of a categorical variable.
dataset with the name of your dataset.
variable1 and variable2 with the names of the two categorical variables you are analyzing.
The table() function creates a two-way frequency table showing the number of observations in each combination of the two categorical variables.
table(dataset$variable1, dataset$variable2)
Dataset name:
data2026
Variable 1 name:
gender
Variable 2 name:
voting
table(data2026$gender, data2026$voting)
Instead of displaying the frequency table only once, it is helpful to save it as an object that can be reused later in your analysis.
Add the <- before the table function to name the object.
The arrow tells RStudio to name the object, and the word it points to is the name you give it.
Here, we have intentionally named the table mytable, because we will use it in the Chi-Square code later.
By writing the word mytable again on a new line, we are "calling" the object and asking RStudio to show it to us in the Console pane.
mytable <- table(dataset$variable1, dataset$variable2)
mytable
polit_table <- table(data2026$gender, data2026$voting)
polit_table
Voting
Gender No Yes
Female 42 158
Male 55 145
This output indicates the number of participants in each combination of gender and voting behavior.
Important:
The order of the categories displayed in the frequency table is important.
Later, when interpreting the Chi-Square test, pay attention to the row and column categories shown in your table.
A bar chart provides a visual summary of the frequencies contained in your frequency table. It allows you to quickly compare the number of observations across categories.
The bar chart provides a quick visual comparison of the category counts before conducting the Chi-Square test.
Bar charts are appropriate for categorical variables because each category represents a distinct group.
Copy and paste the code below into your R Script.
barplot(
mytable,
beside = TRUE,
col = rainbow(nrow(mytable)),
legend = rownames(mytable)
)
barplot(
polit_table,
beside = TRUE,
col = rainbow(nrow(polit_table)),
legend = rownames(polit_table)
)
The chart displays the observed frequencies from your dataset as bars. Higher bars represent categories with more observations, while shorter bars represent categories with fewer observations.
The first line performs the Chi-Square Test of Independence. The second line displays the results.
chi_result <- chisq.test(mytable)
chi_result
chi_result <- chisq.test(polit_table)
chi_result
Pearson's Chi-squared test
data: polit_table
X-squared = 7.84, df = 1, p-value = 0.005
| P-Value | Decision | Conclusion |
|---|---|---|
| p < .05 | Statistically Significant | There is evidence of an association between the two categorical variables. |
| p > .05 | Not Statistically Significant | There is not sufficient evidence of an association between the two categorical variables. |
If you need help interpreting your p-value, use the interactive p-value interpreter.
The Chi-Square Test of Independence is only trustworthy when the expected count in every cell is 5 or greater. Run the code below to display the expected frequencies.
chi_result$expected
Review the output and confirm that every expected frequency is at least 5. If one or more expected frequencies are below 5, the results may not be trustworthy.
For Chi-Square Tests of Independence, the effect size is Cramér's V.
Cramér's V describes the strength of the association between the two categorical variables.
Copy-and-paste the following code into your R Script.
The first line calculates Cramér's V. The second line displays the result.
cramer_v <- rcompanion::cramerV(polit_table)
cramer_v
0.40
Use the table below to classify the strength of the association.
| Cramér's V | Interpretation |
|---|---|
| Less than 0.10 | Negligible Association |
| 0.10 to less than 0.30 | Small Association |
| 0.30 to less than 0.50 | Moderate Association |
| 0.50 or Higher | Large Association |
After the analyses, report your findings in a few clear sentences in your R Script.
Copy the template and replace the highlighted portions with your output results. There is a standardized method of reporting. DO NOT be creative. Use the provided reporting format.
P-Value Reporting
| p-value | How to Report |
|---|---|
| p < .001 | Report p < .001 |
| .001 < p < .05 | Report the exact p-value to three decimals (example: p = .003) |
| p > .05 | Report p > .05 |
| All other values | Report two decimal places (example: .1252 → .13) |
Copy-and-paste the template below into your RScript. Replace the bracketed text with information from your analysis.
# A Chi-Square Test of Independence was conducted to determine if there was an association between [Variable 1] and [Variable 2].
# The results showed that there [was / was not] an association between the two variables, χ²(df) = xx.xx, p = .xxx.
# The association was [small / moderate / large] (Cramer's V = .xx).
# A Chi-Square Test of Independence was conducted to determine if there was an association between gender and voting behavior.
# The results showed that there was an association between the two variables, χ²(1) = 7.84, p = .005.
# The association was moderate (Cramer's V = .32).
After completing your R Script, the next step is to convert it into an R Markdown document.
An R Markdown document combines your code, output, and written explanation into a single HTML report. This allows others to reproduce your analysis and view both the code and results in one document.
Before proceeding, review the R Markdown lesson.
You do not need to complete any of the optional fields. The default settings are sufficient for this assignment.
R code inside an R Markdown file must be placed inside a code chunk.
Add the following line at the very beginning of your document:
```{r}
Then add the following line at the very end of your document:
```
All of your code should now be located between the two lines.
```{r}
paste all your code here
```
When the code chunk has been created correctly, the code region will typically appear shaded or highlighted within RStudio.
Once the code chunk has been created, generate the HTML report.
The report will display your code, output, tables, charts, and statistical results in a web-friendly format.
After creating and knitting your R Markdown document, the final step is to publish your report to RPubs.
RPubs allows you to share your analysis as a web page that can be viewed in any browser without requiring RStudio.
Before beginning, review the RPubs lesson.
After knitting, the completed HTML report should appear in the Viewer pane or open in a web browser.
In addition to publishing the report online, save a copy of the HTML file for your records.