Proficiency Scales Were Supposed to Change Everything. Here's Why They Often Don't.


Marzano's research shows proficiency scales can produce effect sizes of up to 1.0 standard deviations in student achievement. That's enormous.

So why aren't we seeing those results everywhere?

The Reality Check: The research assumes scales are built well. Most of them aren't.

Here's what the research actually requires for scales to work — and where most schools fall short.

1

The scale has to accurately reflect the depth of the standard.

Not just the language of the standard — the depth. A 3.0 target that paraphrases a standard without unpacking what mastery actually requires isn't a proficiency scale. It's a restatement.

Teachers need to understand the standard well enough to describe what a student who has truly met it can do, say, write, and demonstrate. Most scale creation processes don't build that understanding first.

Anatomy of an Effective Scale
1.0 With Help
2.0 The Entry Point
(Instructional Bridge)
3.0 Target Standard Depth
4.0 Cognitive Transfer

The 2.0 level is where instruction happens, yet it is often the thinnest descriptor in practice.

2

The levels have to mean something instructionally — especially the 2.0.

The space between a 1.0 and a 2.0 is where the real teaching happens. For students who are below grade level, the 2.0 is the entry point — the prerequisite skills and foundational knowledge that make the grade-level standard accessible.

In most proficiency scales, the 2.0 descriptor is too thin to be instructionally useful. It tells teachers where a student is. It doesn't tell them nearly enough about how to move them.

3

The scale has to be calibrated to something external.

This is the piece most schools skip entirely. When teachers build scales independently — even when they do it well — there's no guarantee that a 3.0 in one classroom means the same thing as a 3.0 in the classroom next door, or on the state assessment.

Proficiency that isn't calibrated to an external standard is just a local opinion. And local opinions don't translate to state assessment results.

 

Marzano's research works when scales are built with precision, calibrated to the standard at the right depth, and connected to what proficiency actually looks like on the assessments students face.

Most scale creation processes don't get there. Not because teachers aren't capable — but because they're being asked to do technically complex work without the training, time, or tools to do it well.

The research is sound. The implementation is where we're falling short. And that gap is costing students.

Join the Conversation

What does your district's scale creation process look like? What's working — and what isn't? Comment below!

Previous
Previous

I Stood in That Room Wishing I Hadn’t Done the Training

Next
Next

The Problem with Proficiency Scales