Does AI Really Improve English Language Learning?
Willy A Renandya, 11 Oct 2026
Meta-analyses of research on artificial intelligence (AI) in language learning often report positive results. Students who use AI tools appear to improve their English proficiency, and in many studies, the reported gains are substantial. These findings are encouraging, but I believe we need to look at them more carefully.
My concern is not that AI cannot help students learn English. It certainly has the potential to do so. Rather, I wonder how confidently we can attribute the reported improvements to AI itself. Are students learning more because of AI, or are other factors contributing to their progress?
Improvement Does Not Necessarily Mean AI Is Responsible
To understand this issue, we need to distinguish between two common types of research findings.
The first is within-group improvement. Researchers compare students’ performance before and after an AI-assisted learning programme. If students perform better on the post-test, the researchers conclude that learning has taken place.
However, this improvement does not necessarily mean that AI caused it. Students might have improved because they received more instruction, practised more frequently, received additional feedback, or simply became more familiar with the test. Without a suitable comparison group, it is difficult to determine how much of the improvement can be attributed to AI.
The second is between-group improvement. Researchers compare students who use AI with those who do not. If the AI group performs significantly better, this provides stronger evidence that AI-assisted learning may be beneficial.
However, we need to ask another question: What exactly did the comparison group receive?
Imagine a study in which one group of university students uses an AI chatbot to practise English writing. The chatbot provides immediate feedback, suggests improvements and allows students to practise as often as they wish. The other group receives regular classroom instruction but has fewer opportunities to practise and receives less feedback.
If the AI group performs better, can we conclude that AI is responsible for the difference? Not necessarily. The students may have benefited from the additional practice and feedback rather than from AI itself.
This is an important distinction. Evidence that students benefit from an AI-assisted learning programme is not the same as evidence that AI itself produces the improvement.
What Would Convince Me?
I would be more convinced by studies that compare AI-assisted learning with a well-designed non-AI alternative.
Consider the example of Chinese university students learning English writing. Researchers could divide the students into two groups.
Both groups would receive the same writing instruction, work on the same tasks and have the same amount of practice time. The main difference would be how they receive feedback. One group would use AI, while the other would receive comparable feedback from a teacher or another non-AI source.
Both groups would complete the same writing tests before and after the intervention. Ideally, students would be assigned randomly to the two groups, and their writing would be assessed by independent raters who did not know which group they belonged to.
If the AI group made significantly greater progress, we would have stronger evidence that AI contributed something beyond the benefits of instruction, practice and feedback alone.
We could make the evidence even more convincing by testing students again several weeks or months later. We would also ask them to complete new writing tasks without AI assistance. This would help us determine whether they had actually developed their writing ability or had simply become better at producing texts with AI support.
Of course, even this design would not prove that AI works independently of every other factor. Learning is too complex for such a claim. Nevertheless, it would provide much stronger evidence of AI’s additional contribution.
A Meta-Analysis Cannot Fix Every Research Problem
Meta-analyses are valuable because they bring together findings from multiple studies and provide an overall estimate of the effects of an intervention. They can help us see patterns that individual studies might not reveal.
However, a meta-analysis is only as convincing as the research it brings together.
If the original studies have weak comparison groups, unequal learning conditions or limitations in how they measure learning, combining their findings does not automatically solve these problems. A large number of studies does not necessarily mean that the evidence for a causal relationship is strong.
We also need to examine what researchers mean by improved English proficiency. Does this refer to better writing, more accurate grammar, greater vocabulary knowledge or improved speaking performance? Are students assessed independently, or are they evaluated while using AI? Are the reported gains maintained over time?
These questions matter because improvements in performance do not always translate into lasting learning. A student who produces a better essay with AI may not necessarily become a better writer when working independently.
For this reason, when reading a meta-analysis, I would look beyond the overall effect size. I would examine the quality of the original studies, the nature of their comparison groups, the outcomes measured and the extent to which the findings are consistent across different learning contexts.
Should We Expect AI to Work on Its Own?
There is, however, an important qualification to my argument.
I would not insist that researchers must prove that AI alone is responsible for every improvement. In real classrooms, learning results from the interaction of many factors, including teaching methods, materials, feedback, practice and learner motivation.
Furthermore, some of AI’s potential benefits come precisely from the opportunities it creates. For example, AI may make it possible for students to receive more immediate feedback or to practise English more frequently. These are not necessarily competing explanations for AI’s effectiveness; they may be the very ways in which AI helps students learn.
The more useful question, therefore, is not whether AI works independently of all other factors. It is whether AI makes a meaningful additional contribution when compared with other reasonable ways of supporting learning.
This also means that we need to be clear about what we want to find out. If we want to know whether an AI-supported programme is more effective than ordinary classroom instruction, we should compare those two programmes. If we want to know whether AI offers advantages over equally intensive non-AI support, we need a different comparison.
Both questions are important, but they address different issues.
What Does This Mean for TESOL Educators?
For TESOL educators, the practical implications are straightforward.
First, we should be cautious about claims that AI improves English proficiency simply because a meta-analysis reports a positive and statistically significant effect. We need to understand what the studies actually compared and what the students learned.
Second, we should distinguish between using AI to complete a language task and using AI to develop language ability. These are not the same thing. AI may help students produce a more accurate essay, for example, but we still need evidence that they can write more accurately on their own.
Third, we should evaluate AI in relation to sound language teaching principles. Does it provide useful input? Does it encourage meaningful practice? Does its feedback help students notice and correct their errors? Does it support the development of fluency, accuracy and independent language use? These are more important questions than whether a tool is new or technologically impressive.
Finally, we should compare AI with good teaching, not simply with poor teaching or limited learning opportunities. If AI is to become a valuable part of language education, we need to know where it adds value, for which learners, and under what conditions.
The Question We Should Be Asking
I am cautiously optimistic about the role of AI in English language teaching. It offers opportunities to personalise learning, provide timely feedback and increase access to language practice. However, these possibilities should not be confused with conclusive evidence of learning gains.
The key question is not simply whether students who use AI improve their English. It is whether they learn more effectively than they would have through good teaching and appropriately designed non-AI activities.
Until research addresses this question more convincingly, we should interpret claims about AI’s effectiveness with appropriate caution.
Our task as TESOL educators is not to prove that AI works. It is to find out when, how and for whom it improves learning—and whether those improvements extend beyond the immediate assistance the technology provides.
More readings
