Learning academic English requires far more than memorizing individual words. University students must also learn which words naturally occur together, a challenge that becomes especially demanding when they enter terminology-heavy fields such as economics, engineering, medicine and the social sciences. Native speakers instinctively say “strong wind” rather than “muscular wind,” or “reach a conclusion” rather than “pull a conclusion.” For learners, however, these combinations—known as collocations—can be difficult to predict because grammatical correctness alone does not determine whether a phrase sounds natural. New research from the University of British Columbia’s Sauder School of Business suggests that the most dependable way to teach these patterns combines traditional language databases, known as corpora, with carefully controlled generative artificial intelligence.
The study, led by UBC Sauder lecturer Dr. Déogratias Nizonkiza, examines how data-driven learning, or DDL, can help students discover the vocabulary patterns used in real academic and professional communication. In a DDL classroom, learners do not rely exclusively on lists of phrases selected by textbook authors. Instead, they investigate large collections of authentic language, searching for repeated combinations, examining the contexts in which they appear and comparing how frequently different expressions are used. This approach turns students into language researchers. Rather than simply being told that “control volatility” is a common phrase in finance, for example, they can observe how the expression functions in reports, journal articles and market analyses.
The foundation of this method was established decades before artificial intelligence became part of everyday education. In 1961, researchers at Brown University created the Brown Corpus, a pioneering collection containing approximately one million words from 500 samples of American English. The texts represented a broad range of sources, including newspapers, magazines, fiction, religion and scientific writing. Although modest by modern standards, the corpus gave researchers a systematic way to study language as it was actually used. It allowed them to catalogue vocabulary, measure word frequency and identify recurring grammatical and lexical patterns that intuition alone could easily miss.
As computing technology became more accessible in the 1980s, these collections expanded dramatically, and the concept of data-driven learning emerged. Students could search increasingly large databases containing tens or hundreds of millions of words, looking for evidence of how expressions were used beyond the simplified examples of language textbooks. Technically, corpus-based systems can analyze a word’s “collocational profile”—the words that occur near it, the frequency of those combinations and the grammatical structures in which they appear. Such information helps distinguish between possible phrases that are understandable but unusual and combinations that are conventional in a particular discipline.
A major advance arrived with the Corpus of Contemporary American English, or COCA. Containing roughly one billion words and nearly 500,000 texts produced between 1990 and 2019, COCA allows users to search for vocabulary across different genres and fields. A student studying economics can examine how “volatility” appears in academic papers, newspapers, spoken language or popular magazines. The database can reveal that “market volatility,” “price volatility” and “control volatility” are established combinations, while other grammatically possible phrases may be rare or absent. This ability to compare contexts is particularly valuable in higher education, where the same word can acquire specialized meanings and preferred partnerships in different disciplines.
Despite their power, traditional corpus tools have not always been easy to use. Searches can be slow, interfaces may require specialized training and the results can overwhelm students who are unfamiliar with linguistic data. Nizonkiza notes that many instructors were unaware that such corpora existed, while others treated them as optional classroom supplements rather than central teaching resources. Even when teachers introduced corpus searches, students could become discouraged by the time required to sort through hundreds of examples. These barriers limited the wider adoption of DDL, despite evidence that it can improve learners’ awareness of vocabulary patterns and increase the accuracy of their written and spoken English.
Generative AI now offers a faster and more familiar way to interact with language data. Tools such as ChatGPT and Copilot can produce examples, explain possible meanings and suggest words that commonly appear together. They can also transform raw corpus evidence into exercises suited to a particular group of learners. Yet the convenience of these systems creates a scientific and educational problem: users often cannot determine where an AI-generated example originated. A sentence may sound plausible while containing an uncommon collocation, a subtle grammatical error or a phrase that is appropriate in one field but misleading in another. Large language models generate text by predicting likely sequences of words from patterns in their training data; they do not automatically verify every sentence against a curated, transparent corpus.
For that reason, Nizonkiza’s review of empirical research, including meta-analyses and systematic reviews, supports a hybrid model rather than an AI-only approach. In this model, a teacher first selects relevant vocabulary, perhaps from the Economics Academic Word List, and confirms important collocations through a reliable corpus. Terms such as “control volatility” and “market volatility” can then be incorporated into materials generated or adapted with the help of AI. Students might complete a sentence such as “Portfolio diversification is a key strategy used by fund managers to __ volatility in emerging markets.” Only after making their own decision would they ask an AI system to assess the answer, compare alternatives and explain the difference between a natural collocation and an merely possible phrase.
This sequence is crucial because it preserves the cognitive work that makes language learning durable. If students ask an AI platform to supply every answer, they may receive fluent-looking language without developing the ability to evaluate it. If they first investigate evidence, predict a word combination and justify their choice, AI becomes a verification and reflection tool rather than an automatic replacement for learning. The process also encourages critical digital literacy: students must question the origin of an example, compare it with corpus evidence and recognize that frequency does not always equal appropriateness. A phrase may be common in informal conversation but unsuitable for a research paper, or frequent in journalism but technically imprecise in engineering.
The findings arrive as universities seek practical ways to prepare students for specialized communication in an era of rapidly advancing AI. According to Nizonkiza, the transition into higher education is often when learners encounter vocabulary that was absent from their everyday lives. Mastering those words requires understanding not only what they mean, but also how experts routinely combine them. Corpora provide transparent evidence from real language, while generative AI can make that evidence faster to access and easier to turn into individualized practice. Used together, the technologies could make collocation learning more efficient without sacrificing accuracy, context or student judgment. The research therefore presents AI not as the end of corpus-based language teaching, but as a potentially powerful layer built on top of it.
Subject of Research: Data-driven learning, corpus linguistics, generative artificial intelligence and the teaching of English collocations in higher education.
Article Title: At the Intersection of Corpora and Artificial Intelligence: A Critical Review of DDL Approaches to Teaching Collocations in Higher Education
Web References: https://doi.org/10.7202/1126572ar
References: Nizonkiza, Déogratias. “L’enseignement des collocations en anglais au niveau tertiaire : revue critique des approches DDL entre corpus et intelligence artificielle.” Nouvelles perspectives en sciences sociales. DOI: 10.7202/1126572ar. Article publication date: 30 June 2026.
Keywords: English language learning, collocations, data-driven learning, corpus linguistics, generative AI, ChatGPT, higher education, academic English, language education, educational technology

