Researchers gave roughly a thousand high school students access to GPT-4 while they worked through mathematics practice problems. During practice, the students with the assistant scored 48 percent higher than those without it. Then the researchers took the tool away and tested them.
The same students scored 17 percent lower than classmates who had never touched it.
That finding, published in the Proceedings of the National Academy of Sciences by Hamsa Bastani, Osbert Bastani and colleagues, is the most direct evidence available on a question North Texas parents have been arguing about since ChatGPT arrived. It is worth sitting with, because the shape of it is not what either side of the argument usually expects.
The part that should get a parent's attention
The damage did not look like damage. It looked like progress.
Practice scores went up. Homework quality went up. Every signal a school routinely collects moved in the reassuring direction. The deficit only appeared later, on a test taken without the tool present.
That is a measurement problem before it is a teaching problem. A district tracking assignment completion and homework grades through an AI saturated year would see improvement, and would be measuring the software rather than the student. The study's authors describe the students as using the model as a crutch.
The finding that got less attention
There was a third group, and it changes the story.
Those students used the identical GPT-4 model, with one difference: it had been instructed to give incremental hints and never the full answer. They gained more during practice than anyone, 127 percent above the control group. On the exam afterward, they showed no significant harm at all.
Same technology. Same students. Same curriculum. The only variable was how the tool had been told to behave. Which means the question is not really whether AI helps or hurts children. It is who is configuring it, and on what evidence.
Texas already put a machine on the other side of the desk
While the debate has focused on what students do with AI, the state has been quietly using it to grade them.
Since the 2023-24 school year the Texas Education Agency has run an automated scoring engine on open ended STAAR responses in reading, writing, science and social studies. The system scores every constructed response first, with about a quarter rescored by people.
When it has low confidence, or meets a response it does not recognize, such as one heavy with slang or written partly in another language, the answer goes to a human.
The agency built it after the 2023 STAAR redesign multiplied the number of written answers by six or seven. It saves an estimated 15 to 20 million dollars a year. In 2023 the TEA hired about 6,000 temporary scorers; the following year it needed fewer than 2,000.
TEA has pushed back on the artificial intelligence label. "We are way far away from anything that's autonomous or can think on its own," Chris Rozunick, the agency's division director for assessment development, told The Texas Tribune. Jose Rios, TEA's director of student assessment, framed it as a volume problem, saying open ended responses "take an incredible amount of time to score."
Not everyone felt consulted. Kevin Brown, executive director of the Texas Association of School Administrators and a former superintendent, told the Tribune there ought to be some consensus about whether the change is "a good thing, or not a good thing, a fair thing or not a fair thing."
What local districts have settled on
North Texas districts have landed in roughly the same place, and it is closer to the guarded configuration than to a ban.
Plano ISD's published position is specific: "A student shall only use AI tools with teacher permission and shall be expected to produce original work and properly credit sources, including AI tools used in creating the work." The district adds that students who use the tools to harm, bully or harass others face discipline under the student code of conduct.
Its family resources include a guide for parents on using AI for homework help, and its own FAQ takes on the question directly: is AI used to grade student work?






