TELPAS Calibration: Rater Training, the 70 Percent Standard, and When to Run It

Telo AI helps school districts improve speaking outcomes for English Learners and support bilingual education programs through conversational AI and practical tools for educators.

Group of K-12 teachers comparing holistic ratings during a TELPAS calibration training session

Estimated reading time: 10 minutes

TELPAS calibration is the process by which a Texas teacher proves they can assign holistic proficiency ratings that match the state’s standard before rating their own Emergent Bilingual students. It is not paperwork. In the grades and domains that are rated holistically, the teacher is the instrument, and calibration is the only thing that makes one teacher’s Advanced mean the same as another’s. This guide covers what the required training actually involves, the 70 percent standard, where districts lose the year, and why the domain hardest to rate is also the one the schedule under-supplies.

Table of contents

Executive Summary

Texas requires every educator who assigns TELPAS holistic ratings to complete four things before the window opens: test administrator training, grade-cluster-specific online holistic rating training, monitored calibration sessions, and a signed Oath of Test Security and Confidentiality. Calibration itself uses two sets, takes most raters one to two hours per set, and requires 70 percent or better to demonstrate sufficient calibration. New raters must calibrate successfully before rating students; returning raters are expected to calibrate annually but may rate at district discretion if unsuccessful. Successful completion earns two professional development hours. The operational failure is almost never the training itself, which districts complete. It is that the evidence a rater is supposed to judge has to be collected across the whole year, and districts that treat TELPAS as a February event have already lost the data by the time they train.

Key Takeaways

  • Four required activities: TA training, holistic rating training, calibration, security oath.
  • Two calibration sets, one to two hours each for most raters.
  • 70 percent or better demonstrates sufficient calibration.
  • New raters must pass before rating; returning raters calibrate annually, with district discretion.
  • Two professional development hours are earned on successful completion.
  • Training is not the bottleneck. Year-long evidence collection is.

What TELPAS Calibration Is

Quick answer: calibration is a supervised exercise in which a rater scores pre-rated student samples and compares their judgments to the state’s, establishing that they apply the proficiency level descriptors the same way Texas does.

The logic is worth stating plainly, because it explains why the requirement exists at all. Most state assessments are scored by a machine or a contracted scorer, so consistency is an engineering problem. TELPAS holistic ratings are assigned by the classroom teacher, which makes consistency a human problem: without a shared reference, one campus’s Intermediate is another campus’s Advanced, and the district’s data becomes incomparable with itself. Calibration is the mechanism that keeps a rating meaningful across 40 campuses.

That is also why the 70 percent standard is not arbitrary. It is not a test of teaching ability. It is a check that the rater’s internal scale matches the published one closely enough that the ratings they produce can be aggregated, compared year over year, and used in reclassification decisions that affect a child’s placement.

The Four Required Activities

The TELPAS Rater Manual and TEA’s coordinator resources set out four distinct obligations. Districts commonly complete three and assume the fourth is covered.

ActivityWhat it coversWho must complete it
Test administrator trainingTest security, administration procedures, protocols for the cycleEveryone involved in administration
Online holistic rating trainingApplying the holistic rating rubrics, which are the PLDs from the ELPSEvery holistic rater, by grade cluster
Monitored calibration sessionsRating pre-scored samples under supervisionEvery holistic rater
Oath of Test Security and ConfidentialitySigned acknowledgment before handling secure materialsEveryone handling secure materials

Two details matter operationally. The holistic rating training is grade-cluster specific, so a teacher who moves from grade 1 to grade 3 has not already done it. And the rubrics being trained are not a separate TELPAS document: they are the proficiency level descriptors from the English Language Proficiency Standards, which is why the ELPS and the assessment stay locked together, and why a change to the standards is automatically a change to what raters must learn.

How Calibration Works in Practice

  1. Complete the online holistic rating training for the correct grade cluster. This precedes calibration and is not interchangeable with it.
  2. Work through calibration set one. Most raters need one to two hours. The rater scores samples and receives feedback against the state’s ratings.
  3. Check the result against the 70 percent standard. Seventy percent or better demonstrates sufficient calibration.
  4. Use set two if needed. Two sets are available, which gives a rater who misses on the first attempt a genuine second path rather than a retake of identical material.
  5. Confirm the rater’s status before the window. New raters must calibrate successfully before rating students. Returning raters should calibrate annually, and districts retain discretion where a returning rater is unsuccessful.
  6. Record the two professional development hours earned on successful completion, and map them against the Bilingual Education Allotment training requirement.

That last step is the one districts skip and should not. It converts a compliance obligation into funded professional development, which is covered below.

Funding Rater Training

Calibration is TEA-provided and free, but the staff time it consumes is real and fundable. The Bilingual Education Allotment is the natural source and the one most often left on the table.

Funding sourceEligible usesHow to access
Bilingual Education AllotmentAt least 10 percent must fund professional development supporting EB instructionState formula, Texas Education Code section 48.105
Title III, Part ASupplemental EB professional development and technologyFormula grant through TEA
Title II, Part AEducator effectiveness and trainingFormula grant through TEA

The Bilingual Education Allotment carries a spending floor of 10 percent on EB professional development. Rater training is EB professional development. A district that funds calibration release time from general funds is paying twice: once from the general fund and once by failing to spend money it is obligated to spend. See Title III funding in Texas for the federal side.

District Benchmark: The Training Load Nobody Schedules

Calibration looks small on paper and large on a master calendar. Work it through.

Take a district with 2,000 Emergent Bilinguals. Kindergarten and grade 1 are rated holistically across all four domains, and in a typical distribution roughly 300 of those students sit in K-1. Because Emergent Bilinguals are spread across classrooms rather than concentrated, those 300 students are likely distributed among 60 to 80 teachers. Every one of those teachers is a rater.

At one to two hours per calibration set, plus the online holistic rating training and the administrator training, a conservative figure is four hours per rater. Seventy raters at four hours is 280 hours of teacher time, or roughly 47 instructional days’ worth of coverage, before a single student has been rated. Add the writing collections in grades 2-12 and the number grows again.

That is the real cost of TELPAS, and it is invisible in most budgets because it is paid in teacher hours rather than invoices. Districts that plan it in September absorb it. Districts that discover it in January pay for it in substitutes and in ratings assigned under time pressure.

Where Districts Actually Lose the Year

Training completion rates are usually good. The failure sits one layer down, in what the trained rater has to work from.

A holistic rating is a judgment about a student’s typical language use, made against the descriptors, from evidence gathered over time. The rater is not scoring a performance on a given day; they are summarising a year. If the evidence for that summary was not collected, the training does not help, because a perfectly calibrated rater with a thin sample still produces a thin rating.

This is the argument for moving the whole rater workflow to September. Calibration in the autumn tells a teacher what to look for while there is still a year in which to look for it. Calibration in February tells them what they should have been collecting since August.

The Speaking Time Gap: Rating a Domain the Day Barely Produces

One domain is systematically harder to rate than the other three, and the reason is not the rubric. It is the Speaking Time Gap: the shortfall between the responsive spoken practice an Emergent Bilingual needs and what one teacher can supply across a full class.

The math, counted in evidence rather than instruction. A rater assigning a speaking rating needs a representative sample of how a student uses spoken English. In a class of 25 with a 45-minute language block, there are 1,125 student-minutes available, and spoken production is the only domain that cannot be run in parallel: while everyone reads simultaneously, only one child speaks at a time. Even a well-run block yields a couple of minutes of individual oral production per student. Across a year that is a genuine sample for the confident volunteers and a very thin one for the quiet students, who are precisely the students whose rating carries the most consequence.

The result is a systematic bias that no amount of calibration corrects. Calibration aligns how a rater interprets evidence. It cannot manufacture evidence that the schedule never produced. A district serious about rating validity in the speaking domain has to increase the amount of spoken English its Emergent Bilinguals actually produce, which is an instructional decision made in September, not a rating decision made in February.

Metrics Worth Tracking

  • Rater training completion date, measured against September rather than the window.
  • First-attempt calibration pass rate, by campus, as an early warning on rating quality.
  • Number of raters carried on district discretion after an unsuccessful calibration.
  • Documented language samples per Emergent Bilingual collected before January.
  • Rating distribution variance between campuses, which surfaces uncalibrated judgment after the fact.

Common Mistakes District Leaders Make

  1. Scheduling calibration in February. It should tell teachers what to collect, not what they missed.
  2. Assuming a returning rater is calibrated. Annual calibration exists because judgment drifts.
  3. Treating grade-cluster training as transferable. A teacher who changes grade band needs the right cluster.
  4. Funding release time from the general fund while the Bilingual Education Allotment training floor goes unspent.
  5. Auditing training completion but never rating variance. Completion is an input; variance is the outcome.

Immediate (this month): Pull last year’s rater roster and calibration results, and identify how many raters were carried on district discretion.

Medium-term (this year): Move all rater training and calibration into September, and start documented language-sample collection at the same time. Fund the release time from the Bilingual Education Allotment professional development requirement.

Long-term (strategy): Raise the volume of individual oral production Emergent Bilinguals generate, so that speaking ratings rest on a real sample rather than a handful of volunteered answers.

Questions District Leaders Should Ask

  • What is our first-attempt calibration pass rate, and does it vary by campus?
  • How many raters rated students last year without successful calibration?
  • When in the year did our raters complete training?
  • How much documented speaking evidence exists per Emergent Bilingual before January?
  • Does our rating distribution differ between campuses in ways student demographics do not explain?

Frequently Asked Questions

What is TELPAS calibration?

It is a monitored exercise in which a rater scores pre-rated student samples and compares their judgments against the state’s, demonstrating that they apply the proficiency level descriptors consistently before rating their own students.

What score do you need to pass TELPAS calibration?

Seventy percent or better demonstrates sufficient calibration. Two calibration sets are available, and most raters need one to two hours per set.

Who has to complete TELPAS rater training?

Every educator who assigns holistic ratings, plus everyone involved in administration. The required activities are test administrator training, grade-cluster-specific online holistic rating training, monitored calibration, and a signed Oath of Test Security and Confidentiality.

Do returning raters have to calibrate again every year?

Returning raters should complete calibration annually. New raters must complete it successfully before rating students, while districts retain discretion for a returning rater who is unsuccessful.

Does TELPAS calibration count as professional development?

Yes. Raters earn two professional development hours on successful completion, and because the training supports Emergent Bilingual instruction it maps onto the Bilingual Education Allotment requirement that at least 10 percent of the allotment fund EB professional development.

What is a holistic rating?

A rating a trained teacher assigns from ongoing classroom observation, judged against the proficiency level descriptors, rather than from a single sit-down test. It summarises a student’s typical language use over time, which is why the evidence has to be gathered across the year.

When should a district run rater training?

September. Calibration completed in the autumn tells a teacher what evidence to collect while the year is still ahead of them. Calibration completed in February tells them what they should have collected since August.

Conclusion

TELPAS calibration exists because in the holistically rated grades and domains, the teacher is the measuring instrument, and instruments drift. The requirements are modest and the materials are free, so compliance is rarely the problem. The problem is timing: a district that calibrates in February has trained its raters to recognise evidence it no longer has time to collect. Move the whole workflow to September, fund it from the allotment money already earmarked for training, and the ratings that come back in the spring describe students rather than describing how well the district guessed.

Sources: TEA, TELPAS Rater Manual, Holistic Administrations; Texas Education Agency, TELPAS; TEA, TELPAS Proficiency Standards.

When the Rater Is Ready and the Evidence Is Not

A well-run Texas district finishes calibration on time, files the oaths, and puts a trained rater in front of every Emergent Bilingual. Then February arrives and the same teachers say the same thing about the same domain: they are confident about reading, and they are guessing about speaking.

It is not a training failure. A rating summarises a year of language use, and in a full classroom the quiet students simply do not generate enough spoken English for anyone to summarise. The rubric is fine. The sample is thin.

So districts are working on the input rather than the judgment, adding daily speaking practice that does not queue for a teacher’s attention. Telo AI is one example, offering adaptive conversation in English plus Spanish and French for bilingual programmes.

See how districts build the speaking evidence a rating depends on: https://mytelo.ai/how-telo-works/

Calibration makes a rater consistent. It cannot make a student audible. The districts whose speaking ratings hold up are the ones that spent the year making sure there was something to hear.

Found this guide useful?

Preferred sources show up more often in your Google results, including AI Overviews.

Add Telo AI as a preferred source in Google

See how Telo works in the classroom

Learn how Telo helps English Learners practice speaking at their own level while giving teachers real-time insights.