Back to Frisco

Is AI Helping Our Children? The Research Is Blunter Than the Debate

A randomized trial gave about 1,000 students a chatbot for math practice. Their practice scores jumped 48 percent. Then the tool was taken away, and they scored 17 percent worse than classmates who never had it.

Camille Rourke

August 20, 20267 min read

Bar chart of results from a randomized trial of GPT-4 in high school mathematics. During practice, students using a standard chatbot scored 48 percent higher than classmates with no AI access and students using a hints-only tutor scored 127 percent higher. On an exam with the tool removed, the standard-chatbot group scored 17 percent lower, while the hints-only group showed no significant change. Source: Bastani et al., PNAS, 2025.
Bar chart of results from a randomized trial of GPT-4 in high school mathematics. During practice, students using a standard chatbot scored 48 percent higher than classmates with no AI access and students using a hints-only tutor scored 127 percent higher. On an exam with the tool removed, the standard-chatbot group scored 17 percent lower, while the hints-only group showed no significant change. Source: Bastani et al., PNAS, 2025.

Researchers gave roughly a thousand high school students access to GPT-4 while they worked through mathematics practice problems. During practice, the students with the assistant scored 48 percent higher than those without it. Then the researchers took the tool away and tested them.

The same students scored 17 percent lower than classmates who had never touched it.

That finding, published in the Proceedings of the National Academy of Sciences by Hamsa Bastani, Osbert Bastani and colleagues, is the most direct evidence available on a question North Texas parents have been arguing about since ChatGPT arrived. It is worth sitting with, because the shape of it is not what either side of the argument usually expects.

The part that should get a parent's attention

The damage did not look like damage. It looked like progress.

Practice scores went up. Homework quality went up. Every signal a school routinely collects moved in the reassuring direction. The deficit only appeared later, on a test taken without the tool present.

That is a measurement problem before it is a teaching problem. A district tracking assignment completion and homework grades through an AI saturated year would see improvement, and would be measuring the software rather than the student. The study's authors describe the students as using the model as a crutch.

The finding that got less attention

There was a third group, and it changes the story.

Those students used the identical GPT-4 model, with one difference: it had been instructed to give incremental hints and never the full answer. They gained more during practice than anyone, 127 percent above the control group. On the exam afterward, they showed no significant harm at all.

Same technology. Same students. Same curriculum. The only variable was how the tool had been told to behave. Which means the question is not really whether AI helps or hurts children. It is who is configuring it, and on what evidence.

Texas already put a machine on the other side of the desk

While the debate has focused on what students do with AI, the state has been quietly using it to grade them.

Since the 2023-24 school year the Texas Education Agency has run an automated scoring engine on open ended STAAR responses in reading, writing, science and social studies. The system scores every constructed response first, with about a quarter rescored by people.

When it has low confidence, or meets a response it does not recognize, such as one heavy with slang or written partly in another language, the answer goes to a human.

The agency built it after the 2023 STAAR redesign multiplied the number of written answers by six or seven. It saves an estimated 15 to 20 million dollars a year. In 2023 the TEA hired about 6,000 temporary scorers; the following year it needed fewer than 2,000.

TEA has pushed back on the artificial intelligence label. "We are way far away from anything that's autonomous or can think on its own," Chris Rozunick, the agency's division director for assessment development, told The Texas Tribune. Jose Rios, TEA's director of student assessment, framed it as a volume problem, saying open ended responses "take an incredible amount of time to score."

Not everyone felt consulted. Kevin Brown, executive director of the Texas Association of School Administrators and a former superintendent, told the Tribune there ought to be some consensus about whether the change is "a good thing, or not a good thing, a fair thing or not a fair thing."

What local districts have settled on

North Texas districts have landed in roughly the same place, and it is closer to the guarded configuration than to a ban.

Plano ISD's published position is specific: "A student shall only use AI tools with teacher permission and shall be expected to produce original work and properly credit sources, including AI tools used in creating the work." The district adds that students who use the tools to harm, bully or harass others face discipline under the student code of conduct.

Its family resources include a guide for parents on using AI for homework help, and its own FAQ takes on the question directly: is AI used to grade student work?

The teachers are improvising

National survey data suggests the adults are working this out in real time.

Gallup and the Walton Family Foundation found about six in ten teachers used AI during the past school year, and roughly a third use it weekly. Those weekly users report saving an average of 5.9 hours a week, close to six weeks across a school year, which is a real benefit against a backdrop of burnout.

But only 18 percent said they had received any formal guidance from administrators on how the tools should be used. About a third reported no guidance at all. The staff room acquired a productivity tool well before it acquired a policy.

The evidence pointing the other way

A fair reading has to include the strongest counter-result, which is substantial.

A World Bank randomized trial in Benin City, Nigeria, ran first year senior secondary students through six weeks of after school English sessions using Microsoft Copilot, supervised by teachers. On the full assessment, covering English, knowledge of AI and digital skills, the gain was 0.31 standard deviations. On English alone, the study's main outcome, it was 0.23.

The researchers report that the program outperformed 80 percent of the education interventions in a comparison database of randomized trials in developing countries.

Two details matter. The sessions were teacher supervised and structured, much closer to the hint giving arm of the Turkish study than to a child alone with a chatbot.

And the effects were uneven: the study found the largest gains among female students and among those who arrived with higher initial academic performance, which suggests that even where these tools clearly work, they may widen gaps while lifting averages.

Read together, the studies do not actually disagree. They converge on one variable: whether the student is still doing the cognitive work.

What this does not prove

The Turkish trial covers one subject, in one country, across four sessions covering about 15 percent of the math curriculum taught that semester. Mathematics suits a hint-based tutor unusually well, because the problems have determinate answers and the mistakes students make are well catalogued. Whether the same design holds for essay writing or history is untested.

Every study here measures weeks, not years. None can say whether the exam deficit persists, or whether students who learn alongside these tools develop different strengths over time.

And falling test scores are not evidence against AI. National assessment scores peaked in 2012 and the declines among the lowest performing students set in from there, years before ChatGPT existed. Any account blaming recent score movement on chatbots does not survive the timeline.

The question worth asking at home

The research does not support banning the tools, and it does not support handing them over unsupervised. It supports a narrower and more useful question, one a parent can ask tonight without knowing anything technical:

When your child is stuck on a problem, does the tool give the answer, or does it give the next hint?

That single distinction was worth the entire difference between learning and not learning in the only controlled test available. It is also, at the moment, a choice being made mostly by software vendors rather than by parents, teachers or school boards.

Sources

  • Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakci, O. and Mariman, R. "Generative AI without guardrails can harm learning: Evidence from high school mathematics." Proceedings of the National Academy of Sciences, 2025.
  • Gallup and the Walton Family Foundation, teacher AI surveys, 2025.
  • AcadeResearch, "Practice Up 48%, Exams Down 17%: Does AI Stop Children From Learning?", August 2026.
  • The Texas Tribune, "Texas will use computers to grade written answers on this year's STAAR tests," April 2024.
  • Plano ISD, Generative Artificial Intelligence Position and Guidance.
  • World Bank Education Global Department, "From Chalkboards to Chatbots," Nigeria.
Share

Camille Rourke

Camille Rourke covers community life, events, and neighborhood features around Frisco.

Related Stories

More in Texas