What Is Convergent Validity? Definition and Examples

Assessment developers are tasked not only with creating tools to measure knowledge and understanding, but also with ensuring those tools measure what they’re intended to measure. In other words, they need validity evidence to support the interpretation and use of their scores. But understanding the different sources of validity evidence and how to gather them can get complicated quickly.

Convergent validity is one source of validity evidence within this broader process. This article explores what convergent validity is, how to evaluate it, and what its results can indicate. 

What Is Convergent Validity and Why Does It Matter?

Convergent validity examines whether assessments designed to measure the same or a closely related concept, ability, or objective—known as a “construct”—produce scores that are related as expected. The assessments may use different question types, formats, or approaches, but their scores should generally show a meaningful relationship if they are measuring the same construct.

Developed jointly by the American Educational Research Association (AERA), the American Psychological Association (APA), and the National Council on Measurement in Education (NCME), The Standards for Educational and Psychological Testing identify relationships between measures of the same or similar constructs as an important source of validity evidence. This means evidence of convergence can contribute to the case for interpreting and using assessment scores as intended. 

This evidence matters because assessment scores are often used to make important decisions. In education, they may inform instructional planning; in certification and professional testing, they may determine whether someone has demonstrated the knowledge or skills required for a credential.

How Is Convergent Validity Evaluated?

To evaluate convergent validity, developers first need to select an established assessment that measures the same, or a closely related, construct. The credibility of the comparison assessment is vital because it provides a reliable basis for evaluating how well the new assessment measures the intended construct.

Then, the same group of test takers should complete both assessments, allowing developers to examine how scores on the new assessment relate to scores on the established measure. 

Developers can quantify this relationship using correlation coefficients, which describe the strength and direction of the relationship between 2 sets of scores. The calculation considers each test taker’s scores on both assessments, then examines whether higher or lower scores on one tend to correspond with higher or lower scores on the other. This captures the overall pattern of scores instead of simply comparing the average score on each assessment. 

Correlation coefficients range from -1 to +1. A value close to +1 means the two sets of scores tend to rise and fall together more consistently, indicating a stronger positive relationship and generally providing stronger convergent evidence. 

However, no single correlation coefficient automatically establishes convergent validity. Developers should interpret the strength of the relationship in the context of the assessments, the construct being measured, and the level of correlation they would reasonably expect. If the relationship is weaker than expected, it does not necessarily mean the new assessment is invalid, but it does provide a reason to investigate why the results differ and whether additional validity evidence is needed.

Practical Examples of Convergent Validity

You can apply the process for examining convergent validity across any assessment setting, provided the comparison measure assesses the same or a closely related construct. The table below illustrates 3 different settings and how they could use this process.

Setting Construct Assessment Being Validated Comparison Measure
Education Reading comprehension Newly developed reading comprehension test Gates-MacGinitie Reading Tests
Certification Food safety knowledge Newly developed food safety certification exam ServSafe Food Protection Manager Certification Exam
Workforce Critical thinking ability Newly developed workplace critical-thinking assessment Watson-Glaser Critical Thinking Appraisal

In the education example, a strong relationship with the Gates-MacGinitie Reading Tests would provide evidence that the new assessment is measuring reading comprehension as intended, rather than another factor like general test-taking ability. 

For a new food safety certification exam, convergence with the ServSafe Food Protection Manager Certification Exam would support the intended focus on food safety knowledge and skills. A weaker relationship could prompt developers to examine whether the assessment adequately covers the content test takers are expected to know, or whether some items allow them to rely on general knowledge rather than the intended skills.

Similarly, a strong relationship between a workplace critical-thinking assessment and the Watson-Glaser Critical Thinking Appraisal would support the conclusion that the new assessment is measuring critical-thinking ability. However, if the scores don’t relate as expected, developers may examine whether the assessment places too much emphasis on job-specific knowledge or another skill unrelated to critical thinking.

Convergent Validity vs. Discriminant Validity

While convergent validity looks for stronger relationships between measures of the same or related constructs, discriminant validity (also called divergent validity) looks for weaker relationships between measures of different constructs. 

In the certification example above, a newly developed certification exam was expected to show a positive relationship with the ServSafe Food Protection Manager Certification Exam. For example, if the correlation coefficient between those was .75, that could provide evidence supporting convergent validity, assuming the strength of the relationship is consistent with what developers would reasonably expect.

To evaluate discriminant validity with the same assessment, developers could also compare scores with an established measure of reading comprehension. Reading ability may affect how easily test takers understand exam questions, but it’s distinct from the food safety knowledge the certification exam is intended to measure. If the correlation coefficient between the two assessments was .20, that weaker relationship could provide evidence supporting discriminant validity, assuming such a relationship was expected.

Together, convergent and discriminant validity provide stronger evidence about what the assessment is actually measuring. A strong relationship with another measure of food safety knowledge, combined with a weaker relationship with reading comprehension, helps developers distinguish measurement of the intended construct from broader or unrelated factors that could influence scores. 

Common Misconceptions About Convergent Validity

Even when assessment developers understand the purpose of convergent validity, there are common misconceptions about what this evidence can—and cannot—tell them.

  • Misconception #1: Convergent validity is enough to prove an assessment is valid.
    Convergence does not, on its own, establish that an assessment is valid. It’s only one source of validity evidence, so developers also need to consider other evidence, such as whether the test content adequately represents the construct it is intended to measure.
  • Misconception #2: Convergent validity and reliability are the same thing.
    Reliability is about consistency, while convergent validity is about whether assessment scores relate to other measures as expected. An assessment could consistently produce the same results, making it reliable, and still be measuring the wrong thing. For example, a reading assessment could consistently measure vocabulary knowledge even though it was designed to measure reading comprehension.
  • Misconception #3: Convergent validity only matters during initial assessment development.
    Developers don’t evaluate convergent validity once and then move on. As an assessment changes or is used with new populations or in new settings, developers may need to gather new evidence to make sure its scores continue to behave as expected.

How Convergent Validity Fits Into the Assessment Validation Process

No single source of validity evidence can give developers a complete picture of an assessment. Developers may also consider evidence from test content, response processes, internal structure, relationships with other variables, and the consequences of testing. The appropriate sources will depend on the assessment, its intended use, and the claims developers need to support. 

Validation should therefore be treated as an ongoing process rather than a one-time check. During assessment design, developers can identify the construct they intend to measure, consider how scores will be interpreted and used, and plan what evidence will be needed to support those uses.

Once an assessment exists, developers can consider convergent validity evidence alongside other sources of evidence to build and refine the overall validity argument. If the evidence raises questions about whether the assessment is measuring what developers intended, they may need to examine whether administration or content could have affected performance, or whether the established assessment was an appropriate comparator.

The process should also continue into the long term. Changes to assessment content, administration, test-taking populations, or intended uses may affect the evidence supporting score interpretations. Developers may therefore need to gather and evaluate new evidence as an assessment evolves. 

Digital assessment platforms like TAO can support this ongoing work through consistent assessment delivery and structured access to assessment items and performance data. This makes it easier for developers to track changes, review results, and gather further evidence as an assessment evolves.

Build Better Assessments With the Right Evidence

Convergent validity gives developers one way to examine whether assessment scores behave as expected and identify areas that may need further improvement. When considered alongside other sources of validity evidence, it can help developers build a stronger case that an assessment measures its intended construct and supports the decisions made from its scores.