Pearson Correlation & Spearman Correlation

Prepare
Overview of the Test

Correlation is used to examine the relationship between two quantitative variables.

It allows you to determine whether two variables are related and, if so, the direction and strength of that relationship.

In this context, a relationship means that changes in one variable are associated with changes in the other variable.

The research design is observational. Therefore, the variables are not manipulated by the researcher, and correlation does not establish a cause-and-effect relationship.

Component General Structure Example
Research Question Is there a relationship between Variable 1 and Variable 2? Is there a relationship between income and age (years)?
Null Hypothesis There is no relationship between the two variables. There is no relationship between income and age (years).
Alternative Hypothesis There is a relationship between the two variables. There is a relationship between income and age (years).

Correlation produces a correlation coefficient that describes the relationship between the two variables.

The correlation coefficient ranges from -1.00 to +1.00. The sign indicates the direction of the relationship, while the absolute value indicates the strength of the relationship.

Independent Variable and Dependent Variable

Correlation does not require a traditional independent variable (IV) and dependent variable (DV).

Because the design is observational, both variables are treated as quantitative variables, and the goal is to determine whether they are related.

Although one variable may be described as a predictor and the other as an outcome, correlation alone does not demonstrate that one variable causes changes in the other.

Assumptions

  • Both variables are quantitative.
  • Each observation is independent (each participant appears only once).
  • The relationship between the variables should be approximately linear for Pearson correlation.
  • The variables should be approximately normally distributed for Pearson correlation.
  • If the data do not meet the assumptions for Pearson correlation, Spearman correlation may be appropriate.

Interactive Study Tools

Use the Inferential Test Selector to determine which inferential procedure is appropriate for your study.

Inferential Test Selector

Prepare
R Script Code Template

Below is an example of what your code may look like. Your code will be different because the dataset name, variable names, normality results, correlation test, and interpretations will depend on your assigned dataset.

library(readxl)
library(ggpubr)

data2026 <- read_excel("C:/Users/John/OneDrive/Documents/AA5221/Datasets/correlation_data.xlsx")

ggscatter(
data2026,
x = "Age",
y = "Income",
add = "reg.line",
xlab = "Age",
ylab = "Income"
)

# The relationship is linear.

# The relationship is positive.

# There are no outliers.

mean(data2026$Age)
sd(data2026$Age)
median(data2026$Age)

mean(data2026$Income)
sd(data2026$Income)
median(data2026$Income)

hist(data2026$Age)
hist(data2026$Income)

# The variable Age is normally distributed.

# The variable Income is normally distributed.

shapiro.test(data2026$Age)
shapiro.test(data2026$Income)

# The variable Age was normally distributed (p = .754).

# The variable Income was normally distributed (p = .865).

cor.test(
data2026$Age,
data2026$Income,
method = "pearson"
)

# A Pearson correlation was conducted to test the relationship between age and income.

# There was a statistically significant relationship between the variables, r(98) = .42, p = .003.

# The relationship was positive and moderate.

# As age increased, income increased.

Step 1
Install & Open the Packages

Packages add additional functionality to R.

Package Purpose
readxl Import Excel datasets
ggpubr Create scatterplots

Install the Packages

Copy-and-paste the following code into your R Script file.
Run the code once.

install.packages("readxl")
install.packages("ggpubr")

Output

You may see warning messages appear when you install a package. These messages do not necessarily mean that the installation failed. Packages only need to be installed once.


Load the Packages

Copy-and-paste the code below into your R Script.
Keep these lines of code in your R Script.

library(readxl)
library(ggpubr)

Output

No output will appear. This simply opens the packages so you can use their functions in RStudio.

Step 2
Import & Name the Dataset

Although datasets can be imported entirely through code, many students experience difficulty locating file paths. Therefore, this course uses the point-and-click import method to generate the necessary code automatically.

Import the Excel File

  1. Select: File → Import Dataset → From Excel
  2. A new import window will appear. Select: Browse
  3. Locate your assigned Excel dataset.
  4. Select the dataset and choose: Open
  5. Return to the import window and select: Import

Output

Once imported, the dataset will appear in a new tab and will also appear in the Environment pane.

RStudio also automatically generates the code used to import the dataset. This code appears in the Console window. Copy the import code and paste it into your R Script. This allows the dataset to be automatically imported the next time you run your script.

The generated code will vary depending on your file location and file name.

Code Template

DatasetName <- read_excel("filepath")

Example

DatasetName <- read_excel(
"C:/Users/John/OneDrive/Documents/AA5221/Datasets/DatasetName.xlsx"
)
Step 4
View the Relationship with a Scatterplot

Scatterplots allow you to visually inspect the relationship between two variables.

Read: Scatterplots

In a scatterplot:

  • Each dot represents one observation (one person).
  • The x-axis contains the independent (predictor) variable.
  • The y-axis contains the dependent (outcome) variable.

Copy and paste the following code into your R Script:

ggscatter(
    DatasetName,
    x = "IndependentVariable",
    y = "DependentVariable",
    add = "reg.line"
)

Replace:

  • DatasetName
  • IndependentVariable
  • DependentVariable

with the names used in your dataset.

Example

ggscatter(
    DatasetZ,
    x = "Age",
    y = "USD",
    add = "reg.line"
)

What will you see?
A scatterplot showing the relationship between the two variables. The line of best fit makes the relationship easier to interpret.

After creating the scatterplot, summarize what you see as comments in your R Script. Begin each sentence with a hashtag (#).

Relationship Type

Determine whether the relationship is:

  • Linear
  • Curved (non-linear)

A linear relationship is required for a Pearson Correlation. If the relationship appears curved, you will likely use a Spearman Correlation instead.

Relationship Direction

Determine whether the relationship is:

  • Positive
  • Negative
  • No relationship

A positive relationship means both variables increase together. A negative relationship means one variable increases while the other decreases.

Outliers

Look for observations that are far away from the main cluster of points. Determine whether meaningful outliers are present.

Reporting Template

# The relationship is linear / curved.
# The relationship is positive / negative / does not exist.
# There are / are no outliers.
Step 6
Calculate Descriptive Statistics

Calculate the mean, standard deviation, and median for both variables. These values will later be included in your results summary.

Replace:

  • DatasetName with the name of your dataset.
  • IndependentVariable with your independent variable.
  • DependentVariable with your dependent variable.

Code Template

mean(DatasetName$Variable1)
sd(DatasetName$Variable1)
median(DatasetName$Variable1)

mean(DatasetName$Variable2)
sd(DatasetName$Variable2)
median(DatasetName$Variable2) 

Example

mean(DatasetZ$Age)
sd(DatasetZ$Age)
median(DatasetZ$Age)

mean(DatasetZ$USD)
sd(DatasetZ$USD)
median(DatasetZ$USD)

Output

mean(DatasetZ$Age)
1 40.12

sd(DatasetZ$Age)
1 2.19

median(DatasetZ$Age)
1 40

mean(DatasetZ$USD)
1 41000.11

sd(DatasetZ$USD)
1 1002.09

median(DatasetZ$USD)
1 41000
Step 7
Check Normality Visually: Histograms

Checking normality is extremely important when determining whether a Pearson or Spearman Correlation should be used.

Therefore, we check normality in multiple ways. The first method is a visual inspection using histograms.

A histogram displays the distribution of a variable and allows you to look for skewness, symmetry, and the presence of a bell curve.

Create Histograms

Create a histogram for each variable.

A histogram is a graph that shows how the values of a continuous variable are distributed. It groups values into ranges and uses bars to show how many observations fall within each range.

Code Template

hist(DatasetName$Variable1,
     breaks = 15,
     col = "skyblue",
     border = "white")

hist(DatasetName$Variable2,
breaks = 15,
col = "firebrick",
border = "white") 

Example

hist(DatasetZ$Age,
     breaks = 15,
     col = "skyblue",
     border = "white")

hist(DatasetZ$USD,
     breaks = 15,
     col = "firebrick",
     border = "white")

Output

Histograms appear in the Plots pane.


Interpret the Histograms

In your R Script, report whether you think the histogram is normally or abnormally distributed based on a visual assessment. Specifically, look for approximate symmetry and a shape that is reasonably consistent with a bell-shaped distribution.

You can calculate skewness and kurtosis statistically in RStudio to determine if they are normal or abnormal. However, we are keeping things simple for this class.

In order for data to be considered normal, it must have normal skewness AND kurtosis. Review the Data Visualization and Normality lesson (linked below) for what is considered normal skewness and kurtosis.

Data Visualization and Normality

Reporting Template

#Data for Variable1 appears abnormally/ normally distributed.
#Data for Variable2 appears abnormally/ normally distributed.

Example

#Data for Age appears normally distributed.
#Data for USD appears normally distributed.
Step 9
Check Normality Statistically: Shapiro-Wilk Test

The Shapiro-Wilk test allows you to check whether your data are consistent with a normal distribution. A p-value of less than .05 indicates that the data significantly differ from a normal distribution. A p-value greater than .05 indicates that there is not enough evidence to conclude that the data differ from a normal distribution.

Code Template

shapiro.test(DatasetName$Variable1)
shapiro.test(DatasetName$Variable2)

Example

shapiro.test(DatasetZ$Age)
shapiro.test(DatasetZ$USD)

Output

Shapiro-Wilk normality test

data: DatasetZ$Age
W = 0.959
p-value = 0.754

Shapiro-Wilk normality test

data: DatasetZ$USD
W = 0.970
p-value = 0.865

Interpret the Shapiro-Wilk Test

In your R Script, report if the data are normal or abnormal according to the Shapiro-Wilk test.

p-value Interpretation Decision
p > .05 There is no significant evidence that the data differ from a normal distribution. Treat the data as normal
p < .05 There is significant evidence that the data differ from a normal distribution. Treat the data as abnormal

Code Template

#Variable1 is normally / abnormally distributed, (p = .xxx).
#Variable2 is normally / abnormally distributed, (p = .xxx).

Example

#Age was normally distributed, W = 0.959, p = .754.
#USD was normally distributed, W = 0.970, p = .865.

Output

NA

Step 11
Determine Which Correlation to Use

Were any of your histograms or Shapiro-Wilk tests abnormal? If any of them were abnormal, this would require an investigation.

In more advanced data analytics classes, we investigate potential causes of abnormal data (if we had expected it to be normal). Potential causes could be data-entry errors or even major methodological errors. Sometimes abnormal data is due to errors that can be fixed.

Since we are keeping our class simple, just look at your Shapiro-Wilk tests. Were they both normal, or was one or both abnormal?

If both variables are normally distributed and the scatterplot shows a linear relationship, choose the Pearson Correlation.

If one or both variables are not normally distributed, choose the Spearman Correlation.

Option A: Pearson Correlation

Step 12a
Conduct the Pearson Correlation

If both variables are normally distributed and the scatterplot shows a linear relationship, conduct a Pearson Correlation.

The Pearson Correlation measures the strength and direction of a linear relationship between two variables.

Replace the dataset and variable names with the names used in your dataset.

Code Template

cor.test(
    DatasetName$IndependentVariable,
    DatasetName$DependentVariable,
    method = "pearson"
)

Example

cor.test(
    DatasetZ$Age,
    DatasetZ$USD,
    method = "pearson"
)

Output

Pearson's product-moment correlation

data: DatasetZ$Age and DatasetZ$USD
t = 4.275
df = 98
p-value = 0.003

sample estimates:
cor
0.417
  • cor = The correlation coefficient (effect size)
  • ```
  • df = Degrees of freedom. Used to calculate the p-value. For Pearson's correlation, it is the sample size minus 2.
  • p-value = Used to determine whether the correlation is statistically significant.
  • ```

If you need help interpreting your p-value, use the interactive p-value interpreter.

P-Value Interpreter


Effect Size (Correlation Coefficient)

For correlations, the effect size is the correlation coefficient. For Pearson's correlation, it is also known as the r-value.

Unlike tests such as chi-square tests where a separate effect size needs to be calculated, correlation analysis directly produces a measure of effect size.

The correlation coefficient ranges from -1.00 to +1.00. The sign tells you the direction of the relationship, while the number tells you the strength of the relationship.
A positive number means that as Variable 1 increases, Variable 2 increases.
A negative number means that as Variable 1 increases, Variable 2 decreases.

Correlation Interpretation
±0.00–0.19 Very weak relationship
±0.20–0.39 Weak relationship
±0.40–0.59 Moderate relationship
±0.60–0.79 Strong relationship
±0.80–1.00 Very strong relationship
Step 13a
Report the Pearson Correlation

After the analyses, report your findings in a few clear sentences in your R Script. Copy the template and replace the highlighted portions with your output results. There is a standardized method of reporting. DO NOT be creative. Use the provided reporting format.

P-Value Reporting

p-value How to Report
p < .001 Report p < .001
.001 < p < .05 Report the exact p-value to three decimals (example: p = .003)
p > .05 Report p > .05
All other values Report two decimal places (example: .1252 → .13)

Pearson Reporting Template

# A Pearson correlation was conducted to test the relationship between Variable1 (M = xx.xx, SD = xx.xx) and Variable2 (M = xx.xx, SD = xx.xx).

# There was / was not a statistically significant relationship between the two variables, r(df) = .xx, p = .xxx.

# The relationship was positive / negative and weak / moderate / strong.

# As Variable1 increased, Variable2 increased / decreased.

Pearson Example

# A Pearson correlation was conducted to test the relationship between age (M = 40.12, SD = 2.19) and income (M = 41000.11, SD = 1002.09).

# There was a statistically significant relationship between the two variables, r(98) = .42, p = .003.

# The relationship was positive and moderate.

# As age increased, income increased.

Option B: Spearman Correlation

Step 12b
Conduct the Spearman Correlation

If one or both variables are not normally distributed, conduct a Spearman Correlation.

The Spearman Correlation measures the strength and direction of a monotonic relationship between two variables.

Code Template

cor.test(
    DatasetName$IndependentVariable,
    DatasetName$DependentVariable,
    method = "spearman"
)

Example

cor.test(
    DatasetZ$Age,
    DatasetZ$USD,
    method = "spearman"
)

Output

Spearman's rank correlation rho

data: DatasetZ$Age and DatasetZ$USD
S = 1620
p-value = 0.003

sample estimates:
rho
0.42
  • Spearman's rho (ρ) = The correlation coefficient. Values closer to −1 or +1 indicate a stronger relationship between the two variables.
  • Degrees of Freedom (df) = Not reported for Spearman's correlation.
  • p-value = Used to determine whether the correlation is statistically significant.

If you need help interpreting your p-value, use the interactive p-value interpreter.

P-Value Interpreter

For correlations, the effect size is the correlation coefficient itself (ρ). Unlike tests such as chi-square tests where a separate effect size may be calculated, correlation analysis directly produces a measure of effect size.

The value of ρ tells us both the direction and strength of the relationship between two variables, so no additional effect size calculation is needed.

The correlation coefficient ranges from -1.00 to +1.00. The sign tells you the direction of the relationship, while the number tells you the strength of the relationship.

Correlation Interpretation
±0.00–0.19 Very weak relationship
±0.20–0.39 Weak relationship
±0.40–0.59 Moderate relationship
±0.60–0.79 Strong relationship
±0.80–1.00 Very strong relationship
Step 13b
Report the Spearman Correlation

After the analyses, report your findings in a few clear sentences in your R Script.

Copy the template and replace the highlighted portions with your output results. There is a standardized method of reporting. DO NOT be creative. Use the provided reporting format.

P-Value Reporting

p-value How to Report
p < .001 Report p < .001
.001 < p < .05 Report the exact p-value to three decimals (example: p = .003)
p > .05 Report p > .05
All other values Report two decimal places (example: .1252 → .13)

Spearman Reporting Template

# A Spearman correlation was conducted to test the relationship between Variable 1 (Mdn = xx.xx) and Variable 2 (Mdn = xx.xx).

# There was / was not a statistically significant relationship between the two variables, ρ = .xx, p = .xxx.

# The relationship was positive / negative and weak / moderate / strong.

# As Variable1 increased, Variable2 increased / decreased.

Example

# A Spearman correlation was conducted to test the relationship between age (Mdn = 40) and income (Mdn = 41000).

# There was a statistically significant relationship between the two variables, ρ = .42, p = .003.

# The relationship was positive and moderate.

# As age increased, income increased.
Prepare to Publish
Create an RMarkdown File

After completing your R Script, the next step is to convert it into an R Markdown document.

An R Markdown document combines your code, output, and written explanation into a single HTML report. This allows others to reproduce your analysis and view both the code and results in one document.

Before proceeding, review the R Markdown lesson.

R Markdown Files


Create a New R Markdown File

  1. Open RStudio.
  2. Select: File → New File → R Markdown
  3. Ensure that HTML is selected as the output format.
  4. Select OK.

You do not need to complete any of the optional fields. The default settings are sufficient for this assignment.


Save the R Markdown File

  1. Select File → Save As.
  2. Save the file to your desktop.
  3. Delete all of the prewritten sample text generated by RStudio.
  4. Open your completed R Script.
  5. Copy all of your code.
  6. Paste the code into the empty R Markdown document.

Create a Code Chunk

R code inside an R Markdown file must be placed inside a code chunk.

Add the following line at the very beginning of your document:

```{r}

Then add the following line at the very end of your document:

```

All of your code should now be located between the two lines.

Example Structure

```{r}

paste all your code here

```

When the code chunk has been created correctly, the code region will typically appear shaded or highlighted within RStudio.


Step 5: Knit the Document

Once the code chunk has been created, generate the HTML report.

  1. Select the Knit button at the top of the R Markdown document.
  2. The Knit button resembles a ball of yarn with a knitting needle.
  3. Wait for the document to compile.
  4. A new HTML report should open automatically.

The report will display your code, output, tables, charts, and statistical results in a web-friendly format.

Publish
Create an RPubs File

After creating and knitting your R Markdown document, the final step is to publish your report to RPubs.

RPubs allows you to share your analysis as a web page that can be viewed in any browser without requiring RStudio.

Before beginning, review the RPubs lesson.

RPubs


Open Your R Markdown File

  1. Open your completed R Markdown file.
  2. Verify that the document knits successfully and displays all output correctly.
  3. Select the Knit button if you have not already generated the HTML report.

After knitting, the completed HTML report should appear in the Viewer pane or open in a web browser.


Publish to RPubs

  1. Locate the Publish button in the Viewer pane.
  2. Select Publish to RPubs.
  3. Log in to your RPubs account.
  4. If you do not already have an account, create a free RPubs account.
  5. Enter a title for your report.
  6. Optionally enter a description.
  7. Select Publish.

Save the RPubs HTML File

In addition to publishing the report online, save a copy of the HTML file for your records.

  1. Open your published RPubs report in a web browser.
  2. Press Ctrl + S (Windows) or Command + S (Mac).
  3. Choose a location on your computer.
  4. If available, select Webpage, HTML Only.
  5. Save the file.