Articles · Schools and testing
A machine reads it first. How Texas scores written STAAR
An automated engine scores the longer written answers on English language STAAR tests first. People read a quarter or more, and any answer a district asks to have rescored.
Read the story








On STAAR tests not given in Spanish, the Texas Education Agency scores the constructed responses with an automated scoring engine first. At least 25 percent of them then go to human scorers. In the agency's own chart the engine assigns the score of record on approximately 75% of responses. In December 2023, the agency started moving from fully human scoring to this hybrid model.
The agency gave its reason in a presentation on the change. Keeping full human scoring would have cost $15 million to $20 million more a year, because the redesigned test carries several times more written answers to grade. The engine handles the longer written answers, the short and extended constructed responses. On STAAR tests not given in Spanish it is the first reader of those answers.
How the engine learns, and what it is held to
For each question the engine is programmed on a sample of about 3,000 answers that people scored during field testing. It is trained on those answers and the scores people gave them. The agency requires it to agree with human scorers at the same rate human scorers agree with one another. Both the engine's scores and the human scores are monitored daily through the scoring window.
Which answers go to people
Part of the share that goes to people is a random sample. The rest is chosen by the engine itself. It flags an answer that is blank, too short or mostly repeated. It flags one written in another language or leaning on the reading passage. It flags one worded unlike the answers it was trained on, or one that reads as off topic. It also sends answers it scores with low confidence, which are often the ones on the border between two score points.
An answer routed to people is read by two of them, one giving the score of record and one auditing it. Where a person scores an answer, the person's score is the one that counts. Written answers on the Spanish language STAAR never reach the engine. They are 100 percent human scored.
The fall review, and what it did not touch
Texas released its preliminary school ratings on August 14th, 2026. In the review, content experts reread short typed answers, the kind where a student enters a number, a word or a phrase. The agency's example is a student who typed ferst when the answer was first, and who may now get credit.
Agency officials said about 27,200 students received additional credit, and that about 1.7% of more than 1.6 million exams reviewed saw scores improve. Nine campuses and five districts saw their A-F ratings rise as a result, and the agency said it would update those ratings without an appeal. The agency said no score will be lowered as a result of the review.
The agency said none of the regraded answers had been scored by its automated scoring engine. The engine handles the longer written answers. So the correction that moved school ratings this fall came from people reading answers the machine never scored. It says nothing either way about how often the machine's own scores would change if a person read them.
What a family can do
A rescore can start with a parent or guardian who has concerns about a student's score. The district's testing staff can request that the answer be rescored, for a fee. A rescore is sent to people, not to the engine. If the score on that question changes, the fee is waived. The request has to go in within a specific window, and the agency points to its Calendar of Events for the dates.
The agency's scoring document says a rescore costs a fee and names no amount, and it leaves the dates to that calendar. The record does not yet hold either. A family that wants a person to read a written answer starts with the district's testing staff.
Sources for this story
- TEA, Scoring Process for STAAR Constructed Responses, December 2023
- TEA, Hybrid Scoring Key Questions, March 2024
- The Texas Tribune, Megan Menchaca, September 4th
Claim-by-claim verification · 35 claims
Under TEA's hybrid model an automated scoring engine scores the constructed responses on tests not given in Spanish first. At least 25 percent of them then go to human scorers.
This means that the responses are initially scored by an automated scoring engine (ASE) and at least 25 percent of student responses are then routed to human scorers.
Constructed responses on the Spanish-language STAAR are 100 percent human scored and are never scored by the engine.
Student responses to SCR and ECR questions for STAAR Spanish assessments are 100 percent human scored and do not get scored by the ASE.
When a response is sent to a human scorer, the human's score becomes the score of record.
Any student responses that are routed for human scoring maintain the score assigned by the human scorer as the score of record.
The engine is trained on student responses from field tests and on the scores humans gave those responses.
The ASE is trained on student responses and human scores from the field-test data.
TEA requires the engine to agree with human scorers at the same rate human scorers agree with one another.
human scorers at the same rate human scorers agree with one another and that the distribution of ASE scores is similar to the distribution of human scores.
The engine gives a condition code to a response that is blank, too short or mostly duplicated text. It does the same for one in another language or mostly stimulus material. It also codes one worded unlike its training responses, or one off topic or off task.
Condition codes indicate that a response is blank, uses too few words, uses mostly duplicated text, is written in another language, consists primarily of stimulus material, uses vocabulary that does not overlap with the vocabulary in the subset of responses used to train the ASE, or uses language patterns that are reflective of off-topic or off-task responses.
Responses the engine marks as low confidence are often ones that fall on the border between two score points.
The low confidence responses are often those responses that are on the border between
Both engine and human scoring results are monitored daily for the whole scoring window.
Throughout the scoring window, both ASE and human scoring results are monitored daily
District testing personnel can ask for a student's written response to be rescored, for a fee.
district testing personnel can request that the student’s response be rescored for a fee.
A requested rescore goes to human scorers.
Rescore requests are sent for human scoring.
If the rescore changes the score on the question, the fee is waived.
If the score on the constructed-response question changes, the fee is waived.
A rescore request can start with concerns about a student's score raised by district personnel or by a parent or guardian.
If district personnel or a parent or guardian has concerns about a student’s score
TEA's March 2024 presentation says the engine assigns the score of record for approximately 75% of responses.
Scenario 1: Auto scoring engine assigns score of record Approx. 75% responses
TEA said that keeping full human scoring would have cost $15-20M more per year.
maintaining full human scoring would have cost $15-20M more per year
TEA said the STAAR redesign left it with 6-7x more constructed responses to grade each year.
With 6-7x more constructed responses to grade annually for STAAR
The engine is programmed on a sample of about 3,000 human scored responses from the field test for each item it scores.
The engine uses a sample of ~3,000 human scored responses from the field test for programming.
Hybrid scoring began in December 2023, when student responses moved from fully human scoring to hybrid scoring.
The hybrid scoring transition started in December 2023, and student responses went from fully human scoring to hybrid scoring.
A human scorer who does not keep at least 65% exact agreement and 95% adjacent agreement during the scoring window can't remain a scorer for that administration.
If a scorer does not maintain at least at 65% exact agreement and 95% adjacent agreement during the scoring window, they cannot remain as a scorer for that admin.
TEA officials said about 27,200 students received additional credit for text responses to some questions on the reading portion of STAAR.
Texas Education Agency officials said Friday about 27,200 students received additional credit for their text responses to some questions on the reading portion of the State of Texas Assessments of Academic Readiness.
TEA spokesperson Jake Kobersky said none of the regraded responses had been graded by the agency's automated scoring engine.
However, Kobersky said none of the regraded responses were graded by its “automated scoring engine.”
Content experts rescored text entry responses, in which students type brief text such as a number, word or phrase.
Content experts rescored “text entry” responses, where students type brief text — such as a number, word, or phrase.
The automated model handles short and extended constructed-response questions, which ask students for longer answers.
The automated model reviews short and extended constructed-response questions, which ask students to provide longer answers.
Kobersky said content experts may decide to credit a misspelled answer, and gave the example of a student who entered ferst when the correct answer was first.
For example: If the correct answer to a question was ‘first,’ content experts may decide to award credit for a student that entered the response ‘ferst.’
Agency officials said about 1.7% of more than 1.6 million exams reviewed saw their scores improve.
About 1.7% of more than 1.6 million exams reviewed saw scores improve
Officials said nine campuses and five districts saw their A-F ratings rise as a result of the rescore.
Nine campuses and five districts saw their A-F rating increases as a result
Kobersky named the affected districts as Chester, Crandall, Daingerfield-Lone Star, Floydada Collegiate, Forney, Fredericksburg, Galveston, Jim Hogg County, Odem-Edroy, Southwest and Walnut Bend ISDs.
The affected districts consist of Chester, Crandall, Daingerfield-Lone Star, Floydada Collegiate, Forney, Fredericksburg, Galveston, Jim Hogg County, Odem-Edroy, Southwest and Walnut Bend ISDs
Texas released preliminary accountability ratings on August 14th, 2026.
Texas released the preliminary accountability ratings on Aug. 14.
Districts could appeal their preliminary ratings until September 11th, 2026.
Districts can file an appeal of their preliminary ratings until Sept. 11.
TEA officials will update ratings for the affected schools and districts automatically, without requiring an appeal.
TEA officials will automatically update ratings for the impacted schools and districts without requiring an appeal.
The agency said no student's score will be lowered as a result of the review.
No students’ scores will be lowered as a result of the review.
When hybrid scoring sends an answer to people, two human scorers read it, one giving the score of record and one auditing.
All scores go through auto scoring engine; 25% are double human scored (one score of record, one for auditing)
Part of the share that goes to people is a random sample, scored by two people.
Scenario 3: Random sample of responses for double human scoring
Answers the engine gives a condition code or marks as low confidence are routed to trained human scorers.
Student responses that the ASE assigns condition codes (e.g., use words not seen in the responses used to train the ASE) or that are identified as “low confidence” are routed to
The agency's presentation shows the engine flagging answers for double human scoring by condition code or low confidence.
Scenario 2: Engine flags responses for double human scoring (assignment of condition codes or low confidence*)
District testing staff must submit a rescore request within a specific window, and the agency's Calendar of Events gives the dates.
The rescore request must be submitted in the Test Information Distribution Engine (TIDE) by district testing personnel within a specific window. Refer to the Calendar of Events for specific dates.