Correlation is used to examine the relationship between two quantitative variables.
It allows you to determine whether two variables are related and, if so, the direction and strength of that relationship.
In this context, a relationship means that changes in one variable are associated with changes in the other variable.
The research design is observational. Therefore, the variables are not manipulated by the researcher, and correlation does not establish a cause-and-effect relationship.
| Component | General Structure | Example |
|---|---|---|
| Research Question | Is there a relationship between Variable 1 and Variable 2? | Is there a relationship between income and age (years)? |
| Null Hypothesis | There is no relationship between the two variables. | There is no relationship between income and age (years). |
| Alternative Hypothesis | There is a relationship between the two variables. | There is a relationship between income and age (years). |
Correlation produces a correlation coefficient that describes the relationship between the two variables.
The correlation coefficient ranges from -1.00 to +1.00. The sign indicates the direction of the relationship, while the absolute value indicates the strength of the relationship.
Correlation does not require a traditional independent variable (IV) and dependent variable (DV).
Because the design is observational, both variables are treated as quantitative variables, and the goal is to determine whether they are related.
Although one variable may be described as a predictor and the other as an outcome, correlation alone does not demonstrate that one variable causes changes in the other.
Use the Inferential Test Selector to determine which inferential procedure is appropriate for your study.
Below is an example of what your code may look like. Your code will be different because the dataset name, variable names, normality results, correlation test, and interpretations will depend on your assigned dataset.
library(readxl)
library(ggpubr)
data2026 <- read_excel("C:/Users/John/OneDrive/Documents/AA5221/Datasets/correlation_data.xlsx")
ggscatter(
data2026,
x = "Age",
y = "Income",
add = "reg.line",
xlab = "Age",
ylab = "Income"
)
# The relationship is linear.
# The relationship is positive.
# There are no outliers.
mean(data2026$Age)
sd(data2026$Age)
median(data2026$Age)
mean(data2026$Income)
sd(data2026$Income)
median(data2026$Income)
hist(data2026$Age)
hist(data2026$Income)
# The variable Age is normally distributed.
# The variable Income is normally distributed.
shapiro.test(data2026$Age)
shapiro.test(data2026$Income)
# The variable Age was normally distributed (p = .754).
# The variable Income was normally distributed (p = .865).
cor.test(
data2026$Age,
data2026$Income,
method = "pearson"
)
# A Pearson correlation was conducted to test the relationship between age and income.
# There was a statistically significant relationship between the variables, r(98) = .42, p = .003.
# The relationship was positive and moderate.
# As age increased, income increased.
Packages add additional functionality to R.
| Package | Purpose |
|---|---|
readxl |
Import Excel datasets |
ggpubr |
Create scatterplots |
Copy-and-paste the following code into your R Script file.
Run the code once.
install.packages("readxl")
install.packages("ggpubr")
You may see warning messages appear when you install a package. These messages do not necessarily mean that the installation failed. Packages only need to be installed once.
Copy-and-paste the code below into your R Script.
Keep these lines of code in your R Script.
library(readxl)
library(ggpubr)
No output will appear. This simply opens the packages so you can use their functions in RStudio.
Although datasets can be imported entirely through code, many students experience difficulty locating file paths. Therefore, this course uses the point-and-click import method to generate the necessary code automatically.
Once imported, the dataset will appear in a new tab and will also appear in the Environment pane.
RStudio also automatically generates the code used to import the dataset. This code appears in the Console window. Copy the import code and paste it into your R Script. This allows the dataset to be automatically imported the next time you run your script.
The generated code will vary depending on your file location and file name.
DatasetName <- read_excel("filepath")
DatasetName <- read_excel(
"C:/Users/John/OneDrive/Documents/AA5221/Datasets/DatasetName.xlsx"
)
Scatterplots allow you to visually inspect the relationship between two variables.
Read: Scatterplots
In a scatterplot:
Copy and paste the following code into your R Script:
ggscatter(
DatasetName,
x = "IndependentVariable",
y = "DependentVariable",
add = "reg.line"
)
Replace:
DatasetNameIndependentVariableDependentVariablewith the names used in your dataset.
ggscatter(
DatasetZ,
x = "Age",
y = "USD",
add = "reg.line"
)
What will you see?
A scatterplot showing the relationship between the two variables.
The line of best fit makes the relationship easier to interpret.
After creating the scatterplot, summarize what you see as comments in your R Script.
Begin each sentence with a hashtag (#).
Determine whether the relationship is:
A linear relationship is required for a Pearson Correlation. If the relationship appears curved, you will likely use a Spearman Correlation instead.
Determine whether the relationship is:
A positive relationship means both variables increase together. A negative relationship means one variable increases while the other decreases.
Look for observations that are far away from the main cluster of points. Determine whether meaningful outliers are present.
# The relationship is linear / curved.
# The relationship is positive / negative / does not exist.
# There are / are no outliers.
Calculate the mean, standard deviation, and median for both variables. These values will later be included in your results summary.
Replace:
DatasetName with the name of your dataset.IndependentVariable with your independent variable.DependentVariable with your dependent variable.mean(DatasetName$Variable1)
sd(DatasetName$Variable1)
median(DatasetName$Variable1)
mean(DatasetName$Variable2)
sd(DatasetName$Variable2)
median(DatasetName$Variable2)
mean(DatasetZ$Age)
sd(DatasetZ$Age)
median(DatasetZ$Age)
mean(DatasetZ$USD)
sd(DatasetZ$USD)
median(DatasetZ$USD)
mean(DatasetZ$Age)
1 40.12
sd(DatasetZ$Age)
1 2.19
median(DatasetZ$Age)
1 40
mean(DatasetZ$USD)
1 41000.11
sd(DatasetZ$USD)
1 1002.09
median(DatasetZ$USD)
1 41000
Checking normality is extremely important when determining whether a Pearson or Spearman Correlation should be used.
Therefore, we check normality in multiple ways. The first method is a visual inspection using histograms.
A histogram displays the distribution of a variable and allows you to look for skewness, symmetry, and the presence of a bell curve.
Create a histogram for each variable.
A histogram is a graph that shows how the values of a continuous variable are distributed. It groups values into ranges and uses bars to show how many observations fall within each range.
hist(DatasetName$Variable1,
breaks = 15,
col = "skyblue",
border = "white")
hist(DatasetName$Variable2,
breaks = 15,
col = "firebrick",
border = "white")
hist(DatasetZ$Age,
breaks = 15,
col = "skyblue",
border = "white")
hist(DatasetZ$USD,
breaks = 15,
col = "firebrick",
border = "white")
Histograms appear in the Plots pane.
In your R Script, report whether you think the histogram is normally or abnormally distributed based on a visual assessment. Specifically, look for approximate symmetry and a shape that is reasonably consistent with a bell-shaped distribution.
You can calculate skewness and kurtosis statistically in RStudio to determine if they are normal or abnormal. However, we are keeping things simple for this class.
In order for data to be considered normal, it must have normal skewness AND kurtosis. Review the Data Visualization and Normality lesson (linked below) for what is considered normal skewness and kurtosis.
Data Visualization and Normality
#Data for Variable1 appears abnormally/ normally distributed.
#Data for Variable2 appears abnormally/ normally distributed.
#Data for Age appears normally distributed.
#Data for USD appears normally distributed.
The Shapiro-Wilk test allows you to check whether your data are consistent with a normal distribution. A p-value of less than .05 indicates that the data significantly differ from a normal distribution. A p-value greater than .05 indicates that there is not enough evidence to conclude that the data differ from a normal distribution.
shapiro.test(DatasetName$Variable1)
shapiro.test(DatasetName$Variable2)
shapiro.test(DatasetZ$Age)
shapiro.test(DatasetZ$USD)
Shapiro-Wilk normality test
data: DatasetZ$Age
W = 0.959
p-value = 0.754
Shapiro-Wilk normality test
data: DatasetZ$USD
W = 0.970
p-value = 0.865
In your R Script, report if the data are normal or abnormal according to the Shapiro-Wilk test.
| p-value | Interpretation | Decision |
|---|---|---|
| p > .05 | There is no significant evidence that the data differ from a normal distribution. | Treat the data as normal |
| p < .05 | There is significant evidence that the data differ from a normal distribution. | Treat the data as abnormal |
#Variable1 is normally / abnormally distributed, (p = .xxx).
#Variable2 is normally / abnormally distributed, (p = .xxx).
#Age was normally distributed, W = 0.959, p = .754.
#USD was normally distributed, W = 0.970, p = .865.
NA
Were any of your histograms or Shapiro-Wilk tests abnormal? If any of them were abnormal, this would require an investigation.
In more advanced data analytics classes, we investigate potential causes of abnormal data (if we had expected it to be normal). Potential causes could be data-entry errors or even major methodological errors. Sometimes abnormal data is due to errors that can be fixed.
Since we are keeping our class simple, just look at your Shapiro-Wilk tests. Were they both normal, or was one or both abnormal?
If both variables are normally distributed and the scatterplot shows a linear relationship, choose the Pearson Correlation.
If one or both variables are not normally distributed, choose the Spearman Correlation.
If both variables are normally distributed and the scatterplot shows a linear relationship, conduct a Pearson Correlation.
The Pearson Correlation measures the strength and direction of a linear relationship between two variables.
Replace the dataset and variable names with the names used in your dataset.
cor.test(
DatasetName$IndependentVariable,
DatasetName$DependentVariable,
method = "pearson"
)
cor.test(
DatasetZ$Age,
DatasetZ$USD,
method = "pearson"
)
Pearson's product-moment correlation
data: DatasetZ$Age and DatasetZ$USD
t = 4.275
df = 98
p-value = 0.003
sample estimates:
cor
0.417
If you need help interpreting your p-value, use the interactive p-value interpreter.
For correlations, the effect size is the correlation coefficient. For Pearson's correlation, it is also known as the r-value.
Unlike tests such as chi-square tests where a separate effect size needs to be calculated, correlation analysis directly produces a measure of effect size.
The correlation coefficient ranges from -1.00 to +1.00. The sign tells you the direction of the relationship, while the number tells you the strength of the relationship.
A positive number means that as Variable 1 increases, Variable 2 increases.
A negative number means that as Variable 1 increases, Variable 2 decreases.
| Correlation | Interpretation |
|---|---|
| ±0.00–0.19 | Very weak relationship |
| ±0.20–0.39 | Weak relationship |
| ±0.40–0.59 | Moderate relationship |
| ±0.60–0.79 | Strong relationship |
| ±0.80–1.00 | Very strong relationship |
After the analyses, report your findings in a few clear sentences in your R Script. Copy the template and replace the highlighted portions with your output results. There is a standardized method of reporting. DO NOT be creative. Use the provided reporting format.
P-Value Reporting
| p-value | How to Report |
|---|---|
| p < .001 | Report p < .001 |
| .001 < p < .05 | Report the exact p-value to three decimals (example: p = .003) |
| p > .05 | Report p > .05 |
| All other values | Report two decimal places (example: .1252 → .13) |
# A Pearson correlation was conducted to test the relationship between Variable1 (M = xx.xx, SD = xx.xx) and Variable2 (M = xx.xx, SD = xx.xx).
# There was / was not a statistically significant relationship between the two variables, r(df) = .xx, p = .xxx.
# The relationship was positive / negative and weak / moderate / strong.
# As Variable1 increased, Variable2 increased / decreased.
# A Pearson correlation was conducted to test the relationship between age (M = 40.12, SD = 2.19) and income (M = 41000.11, SD = 1002.09).
# There was a statistically significant relationship between the two variables, r(98) = .42, p = .003.
# The relationship was positive and moderate.
# As age increased, income increased.
If one or both variables are not normally distributed, conduct a Spearman Correlation.
The Spearman Correlation measures the strength and direction of a monotonic relationship between two variables.
cor.test(
DatasetName$IndependentVariable,
DatasetName$DependentVariable,
method = "spearman"
)
cor.test(
DatasetZ$Age,
DatasetZ$USD,
method = "spearman"
)
Spearman's rank correlation rho
data: DatasetZ$Age and DatasetZ$USD
S = 1620
p-value = 0.003
sample estimates:
rho
0.42
If you need help interpreting your p-value, use the interactive p-value interpreter.
For correlations, the effect size is the correlation coefficient itself (ρ). Unlike tests such as chi-square tests where a separate effect size may be calculated, correlation analysis directly produces a measure of effect size.
The value of ρ tells us both the direction and strength of the relationship between two variables, so no additional effect size calculation is needed.
The correlation coefficient ranges from -1.00 to +1.00. The sign tells you the direction of the relationship, while the number tells you the strength of the relationship.
| Correlation | Interpretation |
|---|---|
| ±0.00–0.19 | Very weak relationship |
| ±0.20–0.39 | Weak relationship |
| ±0.40–0.59 | Moderate relationship |
| ±0.60–0.79 | Strong relationship |
| ±0.80–1.00 | Very strong relationship |
After the analyses, report your findings in a few clear sentences in your R Script.
Copy the template and replace the highlighted portions with your output results. There is a standardized method of reporting. DO NOT be creative. Use the provided reporting format.
P-Value Reporting
| p-value | How to Report |
|---|---|
| p < .001 | Report p < .001 |
| .001 < p < .05 | Report the exact p-value to three decimals (example: p = .003) |
| p > .05 | Report p > .05 |
| All other values | Report two decimal places (example: .1252 → .13) |
# A Spearman correlation was conducted to test the relationship between Variable 1 (Mdn = xx.xx) and Variable 2 (Mdn = xx.xx).
# There was / was not a statistically significant relationship between the two variables, ρ = .xx, p = .xxx.
# The relationship was positive / negative and weak / moderate / strong.
# As Variable1 increased, Variable2 increased / decreased.
# A Spearman correlation was conducted to test the relationship between age (Mdn = 40) and income (Mdn = 41000).
# There was a statistically significant relationship between the two variables, ρ = .42, p = .003.
# The relationship was positive and moderate.
# As age increased, income increased.
After completing your R Script, the next step is to convert it into an R Markdown document.
An R Markdown document combines your code, output, and written explanation into a single HTML report. This allows others to reproduce your analysis and view both the code and results in one document.
Before proceeding, review the R Markdown lesson.
You do not need to complete any of the optional fields. The default settings are sufficient for this assignment.
R code inside an R Markdown file must be placed inside a code chunk.
Add the following line at the very beginning of your document:
```{r}
Then add the following line at the very end of your document:
```
All of your code should now be located between the two lines.
```{r}
paste all your code here
```
When the code chunk has been created correctly, the code region will typically appear shaded or highlighted within RStudio.
Once the code chunk has been created, generate the HTML report.
The report will display your code, output, tables, charts, and statistical results in a web-friendly format.
After creating and knitting your R Markdown document, the final step is to publish your report to RPubs.
RPubs allows you to share your analysis as a web page that can be viewed in any browser without requiring RStudio.
Before beginning, review the RPubs lesson.
After knitting, the completed HTML report should appear in the Viewer pane or open in a web browser.
In addition to publishing the report online, save a copy of the HTML file for your records.