Texas AI Docket

Articles · Schools and testing

A machine reads it first. How Texas scores written STAAR

An automated engine scores the longer written answers on English language STAAR tests first. People read a quarter or more, and any answer a district asks to have rescored.

Read the story

On STAAR tests not given in Spanish, the Texas Education Agency scores the constructed responses with an automated scoring engine first. At least 25 percent of them then go to human scorers. In the agency's own chart the engine assigns the score of record on approximately 75% of responses. In December 2023, the agency started moving from fully human scoring to this hybrid model.

The agency gave its reason in a presentation on the change. Keeping full human scoring would have cost $15 million to $20 million more a year, because the redesigned test carries several times more written answers to grade. The engine handles the longer written answers, the short and extended constructed responses. On STAAR tests not given in Spanish it is the first reader of those answers.

How the engine learns, and what it is held to

For each question the engine is programmed on a sample of about 3,000 answers that people scored during field testing. It is trained on those answers and the scores people gave them. The agency requires it to agree with human scorers at the same rate human scorers agree with one another. Both the engine's scores and the human scores are monitored daily through the scoring window.

The human scorers are held to a standard of their own. A scorer who falls below 65% exact agreement and 95% adjacent agreement during the window can't keep scoring for that administration.

Which answers go to people

Part of the share that goes to people is a random sample. The rest is chosen by the engine itself. It flags an answer that is blank, too short or mostly repeated. It flags one written in another language or leaning on the reading passage. It flags one worded unlike the answers it was trained on, or one that reads as off topic. It also sends answers it scores with low confidence, which are often the ones on the border between two score points.

An answer routed to people is read by two of them, one giving the score of record and one auditing it. Where a person scores an answer, the person's score is the one that counts. Written answers on the Spanish language STAAR never reach the engine. They are 100 percent human scored.

The fall review, and what it did not touch

Texas released its preliminary school ratings on August 14th, 2026. In the review, content experts reread short typed answers, the kind where a student enters a number, a word or a phrase. The agency's example is a student who typed ferst when the answer was first, and who may now get credit.

Agency officials said about 27,200 students received additional credit, and that about 1.7% of more than 1.6 million exams reviewed saw scores improve. Nine campuses and five districts saw their A-F ratings rise as a result, and the agency said it would update those ratings without an appeal. The agency said no score will be lowered as a result of the review.

The agency said none of the regraded answers had been scored by its automated scoring engine. The engine handles the longer written answers. So the correction that moved school ratings this fall came from people reading answers the machine never scored. It says nothing either way about how often the machine's own scores would change if a person read them.

What a family can do

A rescore can start with a parent or guardian who has concerns about a student's score. The district's testing staff can request that the answer be rescored, for a fee. A rescore is sent to people, not to the engine. If the score on that question changes, the fee is waived. The request has to go in within a specific window, and the agency points to its Calendar of Events for the dates.

The agency's scoring document says a rescore costs a fee and names no amount, and it leaves the dates to that calendar. The record does not yet hold either. A family that wants a person to read a written answer starts with the district's testing staff.

Sources for this story

Claim-by-claim verification · 35 claims
  1. Under TEA's hybrid model an automated scoring engine scores the constructed responses on tests not given in Spanish first. At least 25 percent of them then go to human scorers.

    This means that the responses are initially scored by an automated scoring engine (ASE) and at least 25 percent of student responses are then routed to human scorers.

    TEA, Scoring Process for STAAR Constructed Responses, December 2023 · Checked September 30th, 2026

  2. Constructed responses on the Spanish-language STAAR are 100 percent human scored and are never scored by the engine.

    Student responses to SCR and ECR questions for STAAR Spanish assessments are 100 percent human scored and do not get scored by the ASE.

    TEA, Scoring Process for STAAR Constructed Responses, December 2023 · Checked September 30th, 2026

  3. When a response is sent to a human scorer, the human's score becomes the score of record.

    Any student responses that are routed for human scoring maintain the score assigned by the human scorer as the score of record.

    TEA, Scoring Process for STAAR Constructed Responses, December 2023 · Checked September 30th, 2026

  4. The engine is trained on student responses from field tests and on the scores humans gave those responses.

    The ASE is trained on student responses and human scores from the field-test data.

    TEA, Scoring Process for STAAR Constructed Responses, December 2023 · Checked September 30th, 2026

  5. TEA requires the engine to agree with human scorers at the same rate human scorers agree with one another.

    human scorers at the same rate human scorers agree with one another and that the distribution of ASE scores is similar to the distribution of human scores.

    TEA, Scoring Process for STAAR Constructed Responses, December 2023 · Checked September 30th, 2026

  6. The engine gives a condition code to a response that is blank, too short or mostly duplicated text. It does the same for one in another language or mostly stimulus material. It also codes one worded unlike its training responses, or one off topic or off task.

    Condition codes indicate that a response is blank, uses too few words, uses mostly duplicated text, is written in another language, consists primarily of stimulus material, uses vocabulary that does not overlap with the vocabulary in the subset of responses used to train the ASE, or uses language patterns that are reflective of off-topic or off-task responses.

    TEA, Scoring Process for STAAR Constructed Responses, December 2023 · Checked September 30th, 2026

  7. Responses the engine marks as low confidence are often ones that fall on the border between two score points.

    The low confidence responses are often those responses that are on the border between

    TEA, Scoring Process for STAAR Constructed Responses, December 2023 · Checked September 30th, 2026

  8. Both engine and human scoring results are monitored daily for the whole scoring window.

    Throughout the scoring window, both ASE and human scoring results are monitored daily

    TEA, Scoring Process for STAAR Constructed Responses, December 2023 · Checked September 30th, 2026

  9. District testing personnel can ask for a student's written response to be rescored, for a fee.

    district testing personnel can request that the student’s response be rescored for a fee.

    TEA, Scoring Process for STAAR Constructed Responses, December 2023 · Checked September 30th, 2026

  10. A requested rescore goes to human scorers.

    Rescore requests are sent for human scoring.

    TEA, Scoring Process for STAAR Constructed Responses, December 2023 · Checked September 30th, 2026

  11. If the rescore changes the score on the question, the fee is waived.

    If the score on the constructed-response question changes, the fee is waived.

    TEA, Scoring Process for STAAR Constructed Responses, December 2023 · Checked September 30th, 2026

  12. A rescore request can start with concerns about a student's score raised by district personnel or by a parent or guardian.

    If district personnel or a parent or guardian has concerns about a student’s score

    TEA, Scoring Process for STAAR Constructed Responses, December 2023 · Checked September 30th, 2026

  13. TEA's March 2024 presentation says the engine assigns the score of record for approximately 75% of responses.

    Scenario 1: Auto scoring engine assigns score of record Approx. 75% responses

    TEA, Hybrid Scoring Key Questions, March 2024 · Checked September 30th, 2026

  14. TEA said that keeping full human scoring would have cost $15-20M more per year.

    maintaining full human scoring would have cost $15-20M more per year

    TEA, Hybrid Scoring Key Questions, March 2024 · Checked September 30th, 2026

  15. TEA said the STAAR redesign left it with 6-7x more constructed responses to grade each year.

    With 6-7x more constructed responses to grade annually for STAAR

    TEA, Hybrid Scoring Key Questions, March 2024 · Checked September 30th, 2026

  16. The engine is programmed on a sample of about 3,000 human scored responses from the field test for each item it scores.

    The engine uses a sample of ~3,000 human scored responses from the field test for programming.

    TEA, Hybrid Scoring Key Questions, March 2024 · Checked September 30th, 2026

  17. Hybrid scoring began in December 2023, when student responses moved from fully human scoring to hybrid scoring.

    The hybrid scoring transition started in December 2023, and student responses went from fully human scoring to hybrid scoring.

    TEA, Hybrid Scoring Key Questions, March 2024 · Checked September 30th, 2026

  18. A human scorer who does not keep at least 65% exact agreement and 95% adjacent agreement during the scoring window can't remain a scorer for that administration.

    If a scorer does not maintain at least at 65% exact agreement and 95% adjacent agreement during the scoring window, they cannot remain as a scorer for that admin.

    TEA, Hybrid Scoring Key Questions, March 2024 · Checked September 30th, 2026

  19. TEA officials said about 27,200 students received additional credit for text responses to some questions on the reading portion of STAAR.

    Texas Education Agency officials said Friday about 27,200 students received additional credit for their text responses to some questions on the reading portion of the State of Texas Assessments of Academic Readiness.

    The Texas Tribune, Megan Menchaca, September 4th · Checked September 30th, 2026

  20. TEA spokesperson Jake Kobersky said none of the regraded responses had been graded by the agency's automated scoring engine.

    However, Kobersky said none of the regraded responses were graded by its “automated scoring engine.”

    The Texas Tribune, Megan Menchaca, September 4th · Checked September 30th, 2026

  21. Content experts rescored text entry responses, in which students type brief text such as a number, word or phrase.

    Content experts rescored “text entry” responses, where students type brief text — such as a number, word, or phrase.

    The Texas Tribune, Megan Menchaca, September 4th · Checked September 30th, 2026

  22. The automated model handles short and extended constructed-response questions, which ask students for longer answers.

    The automated model reviews short and extended constructed-response questions, which ask students to provide longer answers.

    The Texas Tribune, Megan Menchaca, September 4th · Checked September 30th, 2026

  23. Kobersky said content experts may decide to credit a misspelled answer, and gave the example of a student who entered ferst when the correct answer was first.

    For example: If the correct answer to a question was ‘first,’ content experts may decide to award credit for a student that entered the response ‘ferst.’

    The Texas Tribune, Megan Menchaca, September 4th · Checked September 30th, 2026

  24. Agency officials said about 1.7% of more than 1.6 million exams reviewed saw their scores improve.

    About 1.7% of more than 1.6 million exams reviewed saw scores improve

    The Texas Tribune, Megan Menchaca, September 4th · Checked September 30th, 2026

  25. Officials said nine campuses and five districts saw their A-F ratings rise as a result of the rescore.

    Nine campuses and five districts saw their A-F rating increases as a result

    The Texas Tribune, Megan Menchaca, September 4th · Checked September 30th, 2026

  26. Kobersky named the affected districts as Chester, Crandall, Daingerfield-Lone Star, Floydada Collegiate, Forney, Fredericksburg, Galveston, Jim Hogg County, Odem-Edroy, Southwest and Walnut Bend ISDs.

    The affected districts consist of Chester, Crandall, Daingerfield-Lone Star, Floydada Collegiate, Forney, Fredericksburg, Galveston, Jim Hogg County, Odem-Edroy, Southwest and Walnut Bend ISDs

    The Texas Tribune, Megan Menchaca, September 4th · Checked September 30th, 2026

  27. Texas released preliminary accountability ratings on August 14th, 2026.

    Texas released the preliminary accountability ratings on Aug. 14.

    The Texas Tribune, Megan Menchaca, September 4th · Checked September 30th, 2026

  28. Districts could appeal their preliminary ratings until September 11th, 2026.

    Districts can file an appeal of their preliminary ratings until Sept. 11.

    The Texas Tribune, Megan Menchaca, September 4th · Checked September 30th, 2026

  29. TEA officials will update ratings for the affected schools and districts automatically, without requiring an appeal.

    TEA officials will automatically update ratings for the impacted schools and districts without requiring an appeal.

    The Texas Tribune, Megan Menchaca, September 4th · Checked September 30th, 2026

  30. The agency said no student's score will be lowered as a result of the review.

    No students’ scores will be lowered as a result of the review.

    The Texas Tribune, Megan Menchaca, September 4th · Checked September 30th, 2026

  31. When hybrid scoring sends an answer to people, two human scorers read it, one giving the score of record and one auditing.

    All scores go through auto scoring engine; 25% are double human scored (one score of record, one for auditing)

    TEA, Hybrid Scoring Key Questions, March 2024 · Checked September 30th, 2026

  32. Part of the share that goes to people is a random sample, scored by two people.

    Scenario 3: Random sample of responses for double human scoring

    TEA, Hybrid Scoring Key Questions, March 2024 · Checked September 30th, 2026

  33. Answers the engine gives a condition code or marks as low confidence are routed to trained human scorers.

    Student responses that the ASE assigns condition codes (e.g., use words not seen in the responses used to train the ASE) or that are identified as “low confidence” are routed to

    TEA, Scoring Process for STAAR Constructed Responses, December 2023 · Checked September 30th, 2026

  34. The agency's presentation shows the engine flagging answers for double human scoring by condition code or low confidence.

    Scenario 2: Engine flags responses for double human scoring (assignment of condition codes or low confidence*)

    TEA, Hybrid Scoring Key Questions, March 2024 · Checked September 30th, 2026

  35. District testing staff must submit a rescore request within a specific window, and the agency's Calendar of Events gives the dates.

    The rescore request must be submitted in the Test Information Distribution Engine (TIDE) by district testing personnel within a specific window. Refer to the Calendar of Events for specific dates.

    TEA, Scoring Process for STAAR Constructed Responses, December 2023 · Checked September 30th, 2026