AZELLA Scores Explained: How Arizona Decides Who Exits EL Services

Telo AI helps school districts improve speaking outcomes for English Learners and support bilingual education programs through conversational AI and practical tools for educators.

Two educators reviewing AZELLA score reports together in a bright Arizona classroom

Last updated:

Estimated reading time: 14 minutes

AZELLA scores are reported as two different things at once: a single Total Proficiency Scale Score, and four separate domain scale scores for Listening, Speaking, Reading and Writing. Arizona then applies a rule that most district reports never spell out, which is that both have to clear their bar independently. A student can finish above the Proficient cut on the total and still be classified Intermediate because one domain fell short. This guide covers the four levels, the cut scores in force for Spring 2026, the domain gate, and why Arizona’s proficiency rate moved this year.

Table of contents

Executive Summary

AZELLA reports four proficiency levels: Pre-Emergent/Emergent, Basic, Intermediate and Proficient. Arizona’s own English Language Proficiency Standards define only three of them as instructional levels, because the fourth is not a place to teach a student, it is the exit from EL services. To reach it a student needs a Total Proficiency Scale Score at or above a grade-specific cut and a domain level of Intermediate or Proficient in all four domains. That second condition is the one that decides most borderline cases. Speaking carries as little as nineteen percent of the total score depending on grade band, and holds a complete veto over the result. For Spring 2026 the Arizona Department of Education lowered every cut score after a field test with native English speakers, and the projected proficiency rate moved from thirteen percent to twenty-three percent, which means a district comparing this year’s reclassification rate to last year’s is comparing two different scales.

Key Takeaways

  • Four reported levels, three instructional ones. Proficient is an exit, not a teaching band.
  • Two numbers, not one: a Total Proficiency Scale Score and four domain scale scores.
  • The gate: Proficient overall requires the total cut and Intermediate or better in all four domains.
  • Cut scores are now grade-specific. Proficient ranges from 944 in kindergarten to 965 in grade 5.
  • Every cut was lowered for Spring 2026 using the lower point of the conditional standard error of measurement.
  • The projected proficiency rate went from 13% to 23%. Year-over-year comparisons across that line are invalid.
  • The hardest domain changes by grade: Speaking in kindergarten, Writing in high school.
  • Proficient starts a two-year monitoring period, not the four years California uses.

The Four AZELLA Proficiency Levels

Quick answer: Pre-Emergent/Emergent, Basic, Intermediate, Proficient. The same four apply to the overall result and to each individual domain.

LevelWhat it describesWhat it triggers
Pre-Emergent/EmergentThe earliest stage, reported as one combined levelEL services, most intensive support
BasicDeveloping control, still well short of grade-level accessEL services continue
IntermediateFunctional but incomplete academic EnglishEL services continue, retest next February
ProficientThe exit criterion, not an instructional stageReclassification and two years of monitoring

Here is the distinction almost nobody explains. Arizona’s 2019 English Language Proficiency Standards define three instructional proficiency levels: Pre-Emergent/Emergent, Basic and Intermediate. The test reports four. The fourth exists because the assessment has to name the point at which a student stops being an English Learner, and that point is a policy line rather than a curriculum stage. There are no Proficient-level standards to teach toward, because a student who reaches Proficient is meant to be in the mainstream classroom.

One warning about older paperwork. The EL70 student test history report shows codes such as PrE/E/B, Emerging, B/I, Progressing, No PL and NAT. Those come from earlier AZELLA scales and appear because the report spans several years of a student’s history. They are not the current levels, and a district that builds a longitudinal dashboard on them will produce a chart of scale changes rather than a chart of student progress.

How AZELLA Scores Are Built: Two Numbers, Not One

Quick answer: every student receives one Total Proficiency Scale Score plus four domain scale scores, each on a 100 to 400 range, and each domain carries its own proficiency level.

The domain scores are not equally weighted into the total, and the weights vary by grade band. Reading carries roughly 25 to 33 percent, Writing 24 to 32 percent, Speaking 19 to 30 percent and Listening 16 to 21 percent. The Reading Foundations standards are assessed in kindergarten through grade 5 only, which is part of why the elementary weights differ from the secondary ones.

That weighting is worth holding next to the rule in the following section, because the two together produce the result districts find hardest to explain to families.

The Rule That Decides Most Cases

Arizona states it plainly in the Overall Proficiency Determination: a determination of Proficient requires the Total Proficiency Scale Score to be at or above the grade’s Proficient cut and an Intermediate or Proficient domain proficiency level in all four domains. Miss one domain and the overall result is reported as Intermediate, no matter how high the total.

What that looks like in practice. Take a grade 3 student in Spring 2026. The Proficient cut for grade 3 is a total of 963. Suppose the student scores 975, comfortably above it. Suppose the Speaking domain comes back at 220. The Intermediate cut for grade 3 Speaking is 223. The student is three points short in one domain, is reported Intermediate overall, remains an English Learner, and is retested the following February.

Now put the weighting beside it. Speaking can account for as little as nineteen percent of the total score, and it holds a complete veto over the outcome. No other element of the assessment behaves that way. It is the single most consequential fact about AZELLA scores, and it is why a district reviewing its results should sort by domain rather than by total.

AZELLA Cut Scores for Spring 2026, by Grade

Quick answer: the total-score cuts are now grade-specific at every level. Under the determination Arizona published in November 2023, the Proficient cut was a flat 1000 and the Intermediate cut a flat 920 for every grade.

GradeBasic cutIntermediate cutProficient cutChange to Proficient cut
Kindergarten794870944−56
Grade 1764873947−53
Grade 2799882961−39
Grade 3799887963−37
Grade 4746882958−42
Grade 5775889965−35
Grade 6751883956−44
Grade 7752884955−45
Grade 8758886957−43
Grades 9-12789887962−38

The change column is measured against that November 2023 determination. The average reduction is about 43 scale score points, and it is not evenly distributed: kindergarten and grade 1 fell the furthest, by 56 and 53 points, while grades 3 and 5 fell by 37 and 35. The youngest students saw the biggest movement, which is consistent with the reason the review happened at all.

Domain Cut Scores: The Hardest Domain Changes by Grade

Domain cuts used to be flat too. Under the earlier determination, every domain at every grade used the same bands: Intermediate began at 230 and Proficient at 250. For Spring 2026 they are set per grade and per domain, and the spread between them is where the useful information sits.

Grade bandListeningSpeakingReadingWriting
Kindergarten, Intermediate cut211224217218
Kindergarten, Proficient cut230241236237
Grade 3, Intermediate cut222223223219
Grade 3, Proficient cut242242242237
Grades 9-12, Intermediate cut220223219225
Grades 9-12, Proficient cut238242239243

Read the rows against each other and the pattern is clear. In kindergarten, Speaking is the hard gate: its Proficient cut sits at 241 against 230 for Listening, an eleven-point gap, the widest anywhere in the table. By high school the hard gate has moved to Writing at 243, the highest single cut in the whole scheme. At grade 3, Listening, Speaking and Reading tie at 242 and Writing is the most forgiving at 237.

This matters operationally because intervention is usually planned by grade level and delivered as a single block. The scores say the binding constraint is not the same one in a kindergarten room as in a ninth grade room, and a district running one uniform English language development strategy across both is optimising for the wrong domain in at least one of them.

What Changed in 2026, and Why It Matters

Arizona did not lower these cuts arbitrarily. In spring 2025 the Department ran a Native Speaker Field Test, administering AZELLA to roughly 600 students identified through Home Language Survey data as native English speakers with no prior AZELLA history. The design intent was to collect data from 500; oversampling brought it near 600.

Two findings came back. Native speakers outperformed EL students, which is what a valid test should show. But the data also showed a higher difficulty rate than anticipated, meaning the assessment was harder than the standard setting had assumed, including for students whose English was never in question. That triggered a psychometric review, and the Department reset the cuts using the lower point of the conditional standard error of measurement rather than the midpoint, which is the technical mechanism behind every number in the two tables above. The performance level descriptors were revised to match.

The projected effect is not small. Arizona’s own adjustment documentation puts the current EL proficiency rate at 13 percent and the proposed rate under the new cuts at 23 percent. The Department attributes the increase to four things together: a test design changed on the field test data, growing familiarity with the assessment, the cut score changes themselves, and several years of professional development.

For a district the practical consequence is a reporting one. A 2026 reclassification rate cannot be compared to a 2025 one. A district whose rate roughly doubles has not necessarily taught better, and one whose rate stays flat may in fact have gone backwards against the old scale. Any board presentation covering both years needs to say so explicitly, or the change will be read as a programme result when it is a scale change.

Reading an AZELLA Score Report

Arizona delivers English Language Proficiency reports through ADEConnect. Three matter for most roles.

ReportWhat it holdsWho needs it
EL70, ELP Student Test HistoryA student’s AZELLA history across yearsCase decisions, placement disputes
EL72, ELP Test RosterResults by roster for a testing groupSite leaders, ELD teachers
EL73, EL Student NeedStudents with an identified EL needScheduling, staffing, compliance

When reading any of them, take the domain levels before the overall level. The overall level tells you the policy outcome. The domain levels tell you why, and in a system where one domain can veto the result, why is the only actionable half.

What an AZELLA Score Actually Decides

Reaching Proficient overall ends EL services and begins a two-year monitoring period. That number catches districts out, because the better-known Californian equivalent monitors reclassified students for four years. A team that learned the process elsewhere will apply the wrong window. Arizona’s is two.

Until a student reaches Proficient, they are reassessed every year in the Spring Reassessment window, which runs roughly from the first of February to the middle of March. Placement testing, by contrast, happens year-round as students enroll. The AZELLA practice test guide covers the difference between the two administrations and where the official materials sit.

Funding Options

Score interpretation costs nothing. What costs money is the instruction that moves the domain holding a cohort below the line, and in Arizona that is usually the same domain every year.

Funding sourceEligible usesHow to access
Title III, Part ASupplemental EL technology, tutoring, professional developmentFormula grant through the Arizona Department of Education
Title I, Part AAcademic support technology in high-poverty schoolsFormula grant through the state
Structured English Immersion fundingDirect costs of the mandated ELD timeState allocation, tied to EL counts
Purchasing cooperativesPre-approved purchasing of qualifying toolsCooperative contract, often without a separate RFP

The constraint to check first is supplement not supplant, set out in the complete guide to Title III funding. Arizona’s mandated ELD block is a state obligation, so Title III cannot fund the block itself. It can fund what runs alongside it.

District Benchmark: The Domain That Holds the Gate

Take an Arizona district with 1,200 English Learners. Suppose 30 percent clear the total-score cut in a given February, which is 360 students. If the four domains behaved independently and each carried a 90 percent chance of also clearing Intermediate, the share clearing all four would be 0.9 to the fourth power, about 66 percent, and roughly 122 students would be blocked by a single domain despite a passing total.

Domains are not independent in reality, so treat that as an illustration of the shape rather than a forecast. But the shape is the point: a four-condition gate fails far more often than any one condition does, and the failure concentrates in whichever domain is weakest across the cohort. Arizona districts consistently report that this is Speaking.

The Speaking Time Gap, Where the Veto Sits

The Speaking Time Gap is the shortfall between the responsive spoken practice an English Learner needs and what a single teacher can supply across a full class. It is a structural constraint rather than a teaching failure, and the AZELLA scoring rule is what converts it from an instructional concern into a reclassification outcome.

The arithmetic. Arizona mandates a minimum of 120 minutes a day of English language development in grades K-5 and 100 minutes in grades 6-12, more protected language time than almost any other state requires. Take that 120-minute block with 25 English Learners in the room, and give an ambitious third of it, 40 minutes, to spoken interaction with the teacher. Divided across 25 students, that is 96 seconds each per day, or under five hours of individually attended speech across a 180-day year.

Doubling the block to 240 minutes takes 96 seconds to just over three minutes. The arithmetic barely improves, because the divisor is the number of students, not the length of the period. Reading and writing scale with the length of a block, since a class can read simultaneously. Speaking is produced one student at a time.

Now recall the weighting. Speaking is worth as little as nineteen percent of the total score, and it can block a Proficient determination on its own. Arizona bought the most protected language time in the country and still watches the one domain that does not divide well decide who exits EL services.

Common Mistakes District Leaders Make

  1. Reading the total score and stopping. The domain levels decide the outcome as often as the total does.
  2. Comparing 2026 reclassification rates to 2025. Different cut scores, different scale, not a programme result.
  3. Treating the EL70 legacy codes as current levels. They come from earlier scales.
  4. Assuming four years of monitoring. Arizona monitors for two.
  5. Assuming the flat 1000 and 250 cuts still apply. Both were replaced by grade-specific and domain-specific values.
  6. Planning one ELD strategy across all grades. The binding domain is Speaking in kindergarten and Writing in high school.
  7. Calling Proficient an instructional level. The standards define three; the fourth is an exit.

Immediate (this month): Pull last February’s results and count how many students cleared the total cut but were held at Intermediate by a single domain. That number is the size of the opportunity, and most districts have never calculated it.

Medium-term (this year): Add a footnote to every board or state report covering both 2025 and 2026 explaining the cut score change. Do it before someone else reads the jump as a programme result.

Long-term (strategy): Resource the blocking domain at a volume that does not divide by class size, and measure the district on domain-level movement rather than on overall proficiency percentage.

Questions District Leaders Should Ask

  • How many of our students missed Proficient last year on exactly one domain, and which domain was it?
  • Does our year-over-year reporting disclose the Spring 2026 cut score change?
  • Are our dashboards built on current level names or on legacy EL70 codes?
  • Do our kindergarten and high school ELD plans target different domains, as the cut scores imply they should?
  • How many minutes of individual spoken production does each English Learner actually receive per day?

Frequently Asked Questions

What are the AZELLA proficiency levels?

Four: Pre-Emergent/Emergent, Basic, Intermediate and Proficient. The same four apply to the overall determination and to each of the four domains. Arizona’s 2019 standards define only the first three as instructional levels, because Proficient marks the exit from EL services rather than a stage of teaching.

What AZELLA score is passing?

There is no pass mark, but there is an exit criterion. A student must reach the grade’s Proficient cut on the Total Proficiency Scale Score and hold a level of Intermediate or Proficient in all four domains. For Spring 2026 the total cut ranges from 944 in kindergarten to 965 in grade 5.

Can a student be Proficient overall with one weak domain?

No. Any domain below Intermediate caps the overall determination at Intermediate regardless of the total score. This is the rule that decides most borderline cases, and speaking is the domain that most often triggers it.

Why did AZELLA cut scores change in 2026?

A Native Speaker Field Test run in spring 2025 with roughly 600 native English speakers showed a higher difficulty rate than anticipated. Arizona reviewed the standard setting and reset every cut using the lower point of the conditional standard error of measurement, revising the performance level descriptors to match.

Will more students be reclassified under the new cut scores?

Arizona’s own documentation projects a move from a 13 percent proficiency rate to 23 percent, attributing it to the cut score change together with test design changes, familiarity and several years of professional development. Districts should expect their reclassification numbers to rise for reasons that are partly external to their instruction.

What happens after a student scores Proficient?

EL services end and a two-year monitoring period begins. Arizona monitors for two years, not the four used in California, and misapplying the longer window is a common source of reporting error.

Why does the score report show levels I do not recognise?

The EL70 student test history report spans multiple years and therefore multiple scales. Codes such as Emerging, B/I, Progressing, No PL and NAT are historical, not current. Only Pre-Emergent/Emergent, Basic, Intermediate and Proficient are in use now.

Conclusion

AZELLA scores are simpler than they look and stricter than they read. Four levels, one total, four domains, and a gate that requires all of them to clear at once. The cut scores that define that gate moved for Spring 2026 and moved most for the youngest students, which changes what a reclassification rate means for at least one reporting cycle. The districts that will read their results correctly this year are the ones that sort by domain, footnote the scale change, and know which single domain is holding their cohort at Intermediate.

Sources: ADE, AZELLA Overall Proficiency Determination, Spring 2026; ADE, AZELLA Reassessment Adjustments Spring 2026; ADE, AZELLA Overall Proficiency Determination, November 2023; ADE, AZELLA Assessment; ADE, English Language Proficiency Standards.

Nineteen Percent, and a Full Veto

There is something almost unfair in the structure. A domain worth less than a fifth of the score can hold a student in EL services for another year while every other measure says they are ready. Arizona did not design that by accident; it designed it because a student who cannot speak the language has not acquired it, whatever the reading score says. The rule is right.

What the rule exposes is a capacity problem. Reading and writing improve when a block gets longer, because thirty children can read at once. Speaking does not, because one child speaks at a time and the rest wait. Arizona demonstrated that more convincingly than any other state, by legislating two hours a day and watching the speaking domain keep deciding the outcome anyway.

So some districts are adding spoken practice that does not queue for a teacher’s attention, running alongside the mandated block rather than competing with it for minutes. Telo AI is one example of that approach, offering adaptive conversation in English plus Spanish and French for dual language programmes.

See how districts add individual speaking practice inside a mandated ELD block: https://mytelo.ai/how-telo-works/

The cut scores tell you where the line is. They do not tell you how many children can speak at once, and that is the number that decides who crosses it.

Found this guide useful?

Preferred sources show up more often in your Google results, including AI Overviews.

Add Telo AI as a preferred source in Google

See how Telo works in the classroom

Learn how Telo helps English Learners practice speaking at their own level while giving teachers real-time insights.