What is PISA and Why It Matters Beyond Rankings

Every PISA cycle produces the same news story. Within hours of publication a league table circulates, a few countries’ systems are declared to be rising or falling, and the assessment behind the numbers goes largely unexamined. Initial results from PISA 2025 are due on 8 September 2026, and the pattern will probably repeat.

Nevertheless, there is far more to PISA than the ranking. Behind it sits a program designed to ensure that a science task written once can reach 15-year-olds in more than 90 education systems, in 54 languages, on widely varying hardware, while still producing scores that carry the same meaning. None of that happens by default. That comparability has to be built into every stage, from sampling and question design to translation, delivery, and response capture.

This article looks beyond the rankings to explain what PISA measures, how its methodology makes international comparisons possible and what has changed as a result of digital delivery. It also examines what PISA 2025 can tell us—and how its results can be interpreted when they are released.

What Is PISA?

The Programme for International Student Assessment is run by the Organization for Economic Co-operation and Development (OECD). It has run every 3 years since 2000, and PISA 2025 is its ninth cycle

PISA assesses 15-year-olds, an age chosen because it falls near the end of compulsory schooling in most participating systems. It’s a sample survey rather than a census, so each country tests a statistically selected sample designed to represent its wider population of 15-year-olds instead of the whole cohort.

Delivery is not handled by the OECD alone. A governing board of participating countries sets priorities, while an international consortium of contractors builds and administers the assessment against detailed technical standards. That division matters later, because most of the machinery protecting comparability sits with the consortium.

Why the OECD Created It

By the late 1990s, governments could compare educational inputs without much difficulty. Spending per student, class sizes, and teaching hours were all reasonably well documented. What was harder to compare, however, was educational outcomes.

National examinations measure how well students have learned a national curriculum, which makes them well suited to domestic accountability but less useful for international comparison. A rising pass rate in one country reveals little about how its students perform against those in another. PISA was created to address this gap, providing outcome evidence that could legitimately be compared across education systems. 

What PISA Actually Measures

PISA isn’t designed to test whether 15-year-olds have mastered a particular national curriculum. Instead, it assesses whether they can apply knowledge and skills in unfamiliar, real-world contexts, rather than simply recall what they have been taught. This is why the framework describes what it measures as literacy rather than attainment.

Scientific literacy, for example, is not a matter of recalling definitions. It asks whether a student can interpret evidence, evaluate a claim, and reason about a situation they have not met in a textbook. By focusing on broadly applicable knowledge and skills rather than individual national curricula, PISA provides a fairer common basis for comparing students across education systems, since no single national curriculum could serve as an appropriate yardstick for all participating countries.

Each cycle covers reading, mathematics, and science, with one domain treated in depth. PISA 2025 makes science the focus, with reading and mathematics as minor domains. Two additions distinguish this cycle:

Students, school leaders and, in some systems, teachers and parents also complete questionnaires. These provide contextual information about students’ backgrounds, attitudes, and learning environments, making the scores more interpretable and allowing researchers and policymakers to explore differences between education systems in greater depth.

How Governments Use the Results

Ministries take PISA seriously, as it provides evidence they cannot get from national assessments alone. However, it’s worth being precise about what the evidence can support. PISA does not tell a government which policies to adopt. It is a cross-sectional study, so it identifies associations rather than causes, and education systems differ in ways that no questionnaire can fully capture.

What PISA does provide is genuinely useful:

  • Internationally comparable evidence on student outcomes, measured the same way in every participating system
  • Trend data across cycles, which is often more informative than a single ranking position
  • Contextual information on equity, student wellbeing, attitudes, and school conditions
  • Signals about where a system might look more closely using its own national data

The most defensible use of PISA is therefore as a prompt for further investigation rather than a diagnosis in itself. A decline in one domain, or a widening gap between advantaged and disadvantaged students, can prompt policymakers to examine national evidence to identify what may be contributing to the pattern. 

Why the Methodology Matters More Than the Ranking

A ranking is only meaningful if the results behind it are genuinely comparable. If samples are drawn differently in 2 countries, or if a translated task turns out to be harder than its source, the gap between those countries could be measuring how the assessment was delivered rather than the students’ performance. 

For that reason, PISA builds comparability into its methodology at every stage of the assessment, from design to delivery and interpretation. 

Sampling: Who actually sits the assessment

PISA defines its target population tightly. Eligible students are between 15 years and 3 months and 16 years and 2 months at the start of testing and enrolled in grade 7 or above, in public or private schools. Countries submit a sampling frame, which the international sampling contractor validates before any student is selected.

Selection then runs in 2 stages. Schools are drawn with probability proportional to their number of eligible students, and a set number of students is sampled within each participating school. Participation thresholds are explicit: At least 85% of initially selected schools must agree to take part, alongside a minimum student participation rate. Replacement schools can be used only once a lower participation threshold has been reached. 

These standards are enforced rather than aspirational. Countries that fall short are identified in the published results and their data annotated, which is worth checking before comparing any 2 systems. 

Language: One assessment, many national versions

Producing a Finnish or Korean version of a task is not translation in the ordinary sense. The consortium develops parallel source versions in English and French, and the translation guidelines require 2 independent translations, one from each source, which are then merged by a third-party reconciler into a single national version.

Independent verifiers then check that version against the source versions, while field trial data is used to assess whether linguistic equivalence has genuinely been achieved. Adaptation applies even where no translation is needed, because a task set in an unfamiliar cultural context can be harder for reasons unrelated to science or mathematics.

Validity and consistency in practice

Validity here means something specific: that the inferences drawn from scores are justified for the use being made of them. For international comparisons, that depends partly on whether the assessment works consistently across countries. If a task functions differently for students of equal ability in 2 countries, any comparison between them is compromised. PISA therefore uses item-level analysis after the field trial to identify tasks that behave differently and, where necessary, remove them from scaling.

Consistency covers how the assessment is administered and scored. Testing takes place within a defined window, and administration follows a common script. Coding of open responses is also governed by shared rules, with reliability checks across coders. The PISA 2025 technical standards set these requirements out across sampling, translation, security, platform testing, and data submission. 

What Digital Delivery Has Changed

PISA moved to computer-based delivery in 2015, and the shift did more than just replace paper. Digital delivery has expanded what the assessment can measure, while introducing a new set of factors that must be controlled.

Adaptive design and interactive tasks

PISA 2025 extends multi-stage adaptive testing across all 3 core literacies. Rather than giving every student an identical form, the assessment routes students through blocks of tasks based on their earlier performance. This allows PISA to measure ability more precisely across the performance range, particularly at the upper and lower ends, without requiring every student to complete the same set of questions.

Interactive tasks further extend what can be measured. In the Learning in the Digital World units, students work within a simulated environment, allowing the platform to capture aspects of how they approach a problem as well as whether they reach the correct solution. This type of process evidence would not be possible to gather through a traditional paper-based assessment. 

Accessibility as a comparability requirement

Accessibility is often filed under compliance. In an international assessment, it’s also a measurement issue: If a student cannot use the interface effectively, their score may reflect the barrier they encountered rather than their competence. Every such score weakens the comparison the program exists to support. 

The practical requirements for accessibility include compatibility with assistive technology, keyboard navigation, sufficient color contrast, resizable text, and dependable audio playback for the listening and speaking components introduced in 2025. Accommodation and exclusion rules matter too, since who is included or left out of a sample changes what the resulting average describes. Platforms built to recognized accessibility standards, such as TAO, make these decisions easier to apply consistently across very different school environments. 

Technical consistency

Digital delivery is comparable only when the technical environment behaves the same way everywhere. In practice, that covers 4 things:

  • Consistent delivery, so a task looks and behaves identically regardless of where or how it is administered
  • Predictable platform behavior across the range of devices, browsers, and operating systems schools actually have
  • Controlled data handling, so responses, timing, and interaction records return intact and attached to the right student and session
  • Interoperability through open standards such as the Question and Test Interoperability (QTI) standard, which keeps assessment content portable rather than tied to one vendor’s format

What PISA 2025 Signals About International Assessment

PISA 2025 was delivered through a single centrally managed digital platform supporting both online and offline sessions, a first for the program. The consortium, led by the Australian Council for Educational Research (ACER) with the TAO platform from Open Assessment Technologies, covered 91 countries and economies and 54 languages across more than 900,000 assessment sessions. Roughly 100,000 of those ran offline.

The scale of offline delivery is significant. Treating intermittent connectivity as a designed delivery route rather than an emergency fallback meant schools with limited infrastructure could still participate in the same assessment program. In that sense, offline capability is not simply a technical feature; it is part of how PISA maintains broad and consistent participation. TAO has published a longer account of what delivery at that scale involved.

The direction of travel is clear, however. PISA 2029 will focus on reading and add media and AI literacy as its innovative domain, which will ask more of the assessment platform, not less. 

Reading the Next Set of Results

When the 2025 figures arrive, the ranking will be the easiest part to report and the least informative part to read. The trend across cycles, the spread between a system’s strongest and weakest students, the contextual data behind the average, and the participation standards a country did or did not meet all carry more information than a position in a table.

None of that evidence exists without the methodology underneath it. Comparability across 91 systems and 54 languages is not a by-product of running the same test everywhere—it is the result of sampling rules, translation procedures, technical standards, and delivery infrastructure all working as intended.

Explore PISA 2025 Beyond the Rankings

Follow TAO’s PISA 2025 webinar series for expert analysis of the results, regional perspectives and practical implications for education systems. Explore upcoming webinars and book your place.

 

Explore What's Possible with TAO