MENU
Authors are listed in the order of first author followed by co-authors.
To view the affiliations of the authors, please click the "Details" button.
Please note that a registered account and completed registration payment are required to view the session details page.
Click on "View Abstract" or "View Full Paper" to view the contents.
Abstracts and Full Papers are also available on the session details pages.
All content published on this page is licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0).
Modern Artificial Intelligence breakthroughs - from robust Machine Translation to Large Language Models (LLMs) - do not exist in a vacuum; they stand on the shoulders of over fifty years of foundational, infrastructural work in Computational Linguistics. This talk traces the genealogy of language AI, analyzing how the transition from explicit, hand-crafted structures to implicit, neural representations trained at scale was made possible by the theoretical and empirical foundations of Language Resources (LR) development.
We revisit pivotal milestones, such as the 1980s revolution in automatic acquisition of lexical information from Machine-Readable Dictionaries (MRDs) and the landmark 1986 Grosseto Workshop. It was here that the term "Language Resources" was officially coined and a manifesto was produced, defining LRs as an essential "infrastructural" utility for technological progress - the digital equivalent of roads, aqueducts, and electricity. Through strategic standardisation initiatives like EAGLES, ISLE and ISO, our community created the "DNA of interoperability" that allows massive datasets to be merged and scaled for training today’s AI.
The current "explosion of AI" is the direct payoff of these decades of data intensive international (many EC) projects and strategic initiatives, characterised by an underlying global strategy paying special attention not only to scientific issues but also to essential methodological, policy and infrastructural dimensions around the notion of language. Among these: reusability, strategic sharing, interoperability, openness, to mention a few.
We argue that true advancements happen with (very BIG) data, but only because these data were painstakingly engineered to be computationally tractable. Finally, we address the next frontier: for example, reshaping our infrastructure for the Generative AI era to combat the digital extinction of under-resourced languages and address the critical constraints of AI ethics and data integrity. We conclude that the complexity of modern NLP tasks demands a "broad and open" vision rooted in global teamwork and interdisciplinary collaboration.
Modern Artificial Intelligence breakthroughs - from robust Machine Translation to Large Language Models (LLMs) - do not exist in a vacuum; they stand on the shoulders of over fifty years of foundational, infrastructural work in Computational Linguistics. This talk traces the genealogy of language AI, analyzing how the transition from explicit, hand-crafted structures to implicit, neural representations trained at scale was made possible by the theoretical and empirical foundations of Language Resources (LR) development.
We revisit pivotal milestones, such as the 1980s revolution in automatic acquisition of lexical information from Machine-Readable Dictionaries (MRDs) and the landmark 1986 Grosseto Workshop. It was here that the term "Language Resources" was officially coined and a manifesto was produced, defining LRs as an essential "infrastructural" utility for technological progress - the digital equivalent of roads, aqueducts, and electricity. Through strategic standardisation initiatives like EAGLES, ISLE and ISO, our community created the "DNA of interoperability" that allows massive datasets to be merged and scaled for training today’s AI.
The current "explosion of AI" is the direct payoff of these decades of data intensive international (many EC) projects and strategic initiatives, characterised by an underlying global strategy paying special attention not only to scientific issues but also to essential methodological, policy and infrastructural dimensions around the notion of language. Among these: reusability, strategic sharing, interoperability, openness, to mention a few.
We argue that true advancements happen with (very BIG) data, but only because these data were painstakingly engineered to be computationally tractable. Finally, we address the next frontier: for example, reshaping our infrastructure for the Generative AI era to combat the digital extinction of under-resourced languages and address the critical constraints of AI ethics and data integrity. We conclude that the complexity of modern NLP tasks demands a "broad and open" vision rooted in global teamwork and interdisciplinary collaboration.
At its most powerful, corpus linguistics is a way of making language visible: revealing the patterns, assumptions, routines, silences, and inequalities that shape social and professional life. In the age of generative AI, this role is not diminished; it is intensified. As large language models increasingly mediate how people write, work, learn, decide, and communicate, applied corpus linguistics (ACL) offers essential methods for examining language at scale while remaining grounded in context, interpretation, accountability, and real-world consequence (Friginal, 2025; Thompson & Friginal, 2021).
In this presentation, I reflect on the continuing relevance of ACL through the lens of my own research on professional communication across intercultural contexts and institutional settings. With domains including aviation, customer service, health communication, office talk, and other workplace and public-facing environments, my work has explored how recurrent linguistic patterns connect to professional practice, intercultural understanding, institutional norms, and decision-making. These connections have helped inform both micro-level policies—such as interactional guidelines, training practices, communication protocols, and feedback frameworks—and macro-level policies concerned with organisational communication, safety, service quality, equity, and public trust.
I also situate this discussion within the broader development of the field, drawing on my role as founding Editor-in-Chief of Applied Corpus Linguistics (ACORP), as the journal celebrates its fifth anniversary. The journal was established in recognition of a major shift: corpus linguistics is no longer confined primarily to general language description or to linguistics alone. Corpus methods are now being adopted, adapted, and extended across diverse areas including forensic linguistics, social policy, food studies, anthropology, writing development, translation and interpreting, and corporate and government communication (Reppen, Goulart, & Biber, 2026). This expansion has created a need for spaces where researchers and practitioners can share case studies, develop methods, theorise practice, and communicate findings accessibly to audiences beyond corpus linguistics. I argue that the “applied” in ACL is not merely a descriptor of topic or method. It is a commitment to using corpus resources, tools, and techniques to address real-world questions: to understand communication in specific settings, to support professional reflection, to improve institutional practice, and to contribute to socially responsive policy. In the context of generative AI, ACL is uniquely positioned to scrutinise machine-generated discourse, evaluate claims about language technologies, support ethical and domain-sensitive uses of AI, and preserve the crucial link between textual patterning and situated human experience.
At its most powerful, corpus linguistics is a way of making language visible: revealing the patterns, assumptions, routines, silences, and inequalities that shape social and professional life. In the age of generative AI, this role is not diminished; it is intensified. As large language models increasingly mediate how people write, work, learn, decide, and communicate, applied corpus linguistics (ACL) offers essential methods for examining language at scale while remaining grounded in context, interpretation, accountability, and real-world consequence (Friginal, 2025; Thompson & Friginal, 2021).
In this presentation, I reflect on the continuing relevance of ACL through the lens of my own research on professional communication across intercultural contexts and institutional settings. With domains including aviation, customer service, health communication, office talk, and other workplace and public-facing environments, my work has explored how recurrent linguistic patterns connect to professional practice, intercultural understanding, institutional norms, and decision-making. These connections have helped inform both micro-level policies—such as interactional guidelines, training practices, communication protocols, and feedback frameworks—and macro-level policies concerned with organisational communication, safety, service quality, equity, and public trust.
I also situate this discussion within the broader development of the field, drawing on my role as founding Editor-in-Chief of Applied Corpus Linguistics (ACORP), as the journal celebrates its fifth anniversary. The journal was established in recognition of a major shift: corpus linguistics is no longer confined primarily to general language description or to linguistics alone. Corpus methods are now being adopted, adapted, and extended across diverse areas including forensic linguistics, social policy, food studies, anthropology, writing development, translation and interpreting, and corporate and government communication (Reppen, Goulart, & Biber, 2026). This expansion has created a need for spaces where researchers and practitioners can share case studies, develop methods, theorise practice, and communicate findings accessibly to audiences beyond corpus linguistics. I argue that the “applied” in ACL is not merely a descriptor of topic or method. It is a commitment to using corpus resources, tools, and techniques to address real-world questions: to understand communication in specific settings, to support professional reflection, to improve institutional practice, and to contribute to socially responsive policy. In the context of generative AI, ACL is uniquely positioned to scrutinise machine-generated discourse, evaluate claims about language technologies, support ethical and domain-sensitive uses of AI, and preserve the crucial link between textual patterning and situated human experience.
Corpus linguistics has benefited substantially from increasingly sophisticated statistical methods. Yet choosing an appropriate model requires more than adopting a method that is widely used or technically advanced. Statistical models should also reflect what we know about the nature and generation of the variables being analysed. This talk argues for greater attention to such structural faithfulness in the statistical modelling of corpus-derived measures.
I begin with syntactic complexity measures, many of which are ratios constructed from counts, such as the number of clauses per sentence. Murakami (2025) demonstrates that conventional regression models assuming normally distributed errors can be poorly suited to such measures. They disregard their theoretical bounds and treat observations based on different numbers of linguistic units as equally informative, despite differences in sampling variability. Analyses of learner corpus data and simulations show that this mismatch can produce theoretically impossible prediction intervals, heteroscedasticity, and problematic statistical inference. Alternative approaches that retain information about the underlying counts can mitigate these problems.
I then broaden the discussion beyond syntactic complexity. Corpus research routinely employs derived measures whose statistical properties depend on the observations from which they are calculated. Examples include measures of lexical diversity and sophistication and association measures such as mutual information and ΔP. Murakami (2025) notes that related issues may also arise when count-based measures are used as predictors. Drawing on illustrative examples, I will explore the consequences of overlooking such properties and consider possible directions for more appropriate modelling.
The broader message is simple: Corpus-derived numbers should not be treated as interchangeable continuous measurements. Understanding how a measure is constructed should be an integral part of deciding how it is statistically modelled.
Reference
Murakami, A. (2025). Towards more appropriate modelling of linguistic complexity measures: Beyond traditional regression models. Research Methods in Applied Linguistics, 4(1), 100182. https://doi.org/10.1016/j.rmal.2025.100182
Corpus linguistics has benefited substantially from increasingly sophisticated statistical methods. Yet choosing an appropriate model requires more than adopting a method that is widely used or technically advanced. Statistical models should also reflect what we know about the nature and generation of the variables being analysed. This talk argues for greater attention to such structural faithfulness in the statistical modelling of corpus-derived measures.
I begin with syntactic complexity measures, many of which are ratios constructed from counts, such as the number of clauses per sentence. Murakami (2025) demonstrates that conventional regression models assuming normally distributed errors can be poorly suited to such measures. They disregard their theoretical bounds and treat observations based on different numbers of linguistic units as equally informative, despite differences in sampling variability. Analyses of learner corpus data and simulations show that this mismatch can produce theoretically impossible prediction intervals, heteroscedasticity, and problematic statistical inference. Alternative approaches that retain information about the underlying counts can mitigate these problems.
I then broaden the discussion beyond syntactic complexity. Corpus research routinely employs derived measures whose statistical properties depend on the observations from which they are calculated. Examples include measures of lexical diversity and sophistication and association measures such as mutual information and ΔP. Murakami (2025) notes that related issues may also arise when count-based measures are used as predictors. Drawing on illustrative examples, I will explore the consequences of overlooking such properties and consider possible directions for more appropriate modelling.
The broader message is simple: Corpus-derived numbers should not be treated as interchangeable continuous measurements. Understanding how a measure is constructed should be an integral part of deciding how it is statistically modelled.
Reference
Murakami, A. (2025). Towards more appropriate modelling of linguistic complexity measures: Beyond traditional regression models. Research Methods in Applied Linguistics, 4(1), 100182. https://doi.org/10.1016/j.rmal.2025.100182
Workplace burnout research has, from its inception, established a structural diagnosis, yet the remedy remains persistently individual. As large language models (LLMs) are increasingly positioned as first-line advisors for workplace distress, this study examines whether LLM chatbots reproduce this disjuncture when users disclose workplace burnout. Drawing on critical discourse analysis, it examines how LLM chatbots name participants, represent action, and allocate causality across burnout-related conversations from the LMSYS-Chat-1M corpus. The findings reveal a stable discursive script – acknowledge, reframe, prescribe – that systematically activates the worker as the sole agent of remedy while passivating or suppressing organisational actors. Remedies cluster overwhelmingly at the individual level: self-regulation, boundary-setting, and cognitive reframing predominate, while collective action and institutional redress are structurally absent. This ‘algorithmic responsibilisation’ discursively privatises occupational hazards, framing political problems of labour as technical problems of individual psychological management. Yet, this individualising default is not absolute; it is disrupted when users employ explicit legal or collective-action language, suggesting that access to structural advice is mediated by linguistic capital and that the individualising register functions as a contingent default shaped by alignment processes rather than an inherent property of language modelling. By automating the resilience discourse that critical scholars have long identified in neoliberal management, LLMs risk naturalising workplace harm at scale, making available a ‘resigned common sense’ in which collective resistance is discursively foreclosed. The findings are consistent with the interpretation that alignment protocols (RLHF) stabilise this depoliticised register, though the relative contributions of pre-training data and alignment remain to be disentangled.
Workplace burnout research has, from its inception, established a structural diagnosis, yet the remedy remains persistently individual. As large language models (LLMs) are increasingly positioned as first-line advisors for workplace distress, this study examines whether LLM chatbots reproduce this disjuncture when users disclose workplace burnout. Drawing on critical discourse analysis, it examines how LLM chatbots name participants, represent action, and allocate causality across burnout-related conversations from the LMSYS-Chat-1M corpus. The findings reveal a stable discursive script – acknowledge, reframe, prescribe – that systematically activates the worker as the sole agent of remedy while passivating or suppressing organisational actors. Remedies cluster overwhelmingly at the individual level: self-regulation, boundary-setting, and cognitive reframing predominate, while collective action and institutional redress are structurally absent. This ‘algorithmic responsibilisation’ discursively privatises occupational hazards, framing political problems of labour as technical problems of individual psychological management. Yet, this individualising default is not absolute; it is disrupted when users employ explicit legal or collective-action language, suggesting that access to structural advice is mediated by linguistic capital and that the individualising register functions as a contingent default shaped by alignment processes rather than an inherent property of language modelling. By automating the resilience discourse that critical scholars have long identified in neoliberal management, LLMs risk naturalising workplace harm at scale, making available a ‘resigned common sense’ in which collective resistance is discursively foreclosed. The findings are consistent with the interpretation that alignment protocols (RLHF) stabilise this depoliticised register, though the relative contributions of pre-training data and alignment remain to be disentangled.
Recent developments in multimodal generative AI (GenAI) have led to the rapid expansion of tools and practices in language education. As a result, researchers, teachers, and learners are increasingly questioning how these technologies can be meaningfully integrated with established approaches such as corpus-based methods in pedagogical design. This is especially true in English for Specific Purposes (ESP), where some teachers may feel that GenAI could replace them as a source of guidance on discipline-specific language use.
This paper proposes a pedagogical framework for integrating corpus linguistics and GenAI into language teaching, grounded in established models of course design: needs analysis, language and learning objectives, materials and methods, and evaluation. Within this framework, corpus methods are positioned as a means of grounding pedagogical decisions in empirical evidence of language use, while GenAI tools are used to support the interpretation, accessibility, and interactive exploration of that evidence, as well as the evaluation of learner development.
To illustrate the framework, the paper draws on a series of classroom-informed examples, including corpus-based analysis of learner needs, discipline-specific language modeling using specialized corpora, data-driven learning (DDL) activities supported by AI tools, and corpus-informed evaluation techniques for identifying prototypical and problematic learner texts. Together, these examples demonstrate how corpus methods and AI can be integrated across all stages of course design in a coherent and pedagogically meaningful way.
The paper also considers challenges associated with integrating corpus methods and GenAI, including issues of reliability, bias, and the potential for tool-driven pedagogy. In response, it emphasizes the importance of aligning technological choices with clearly defined learning goals.
The paper concludes by suggesting that the future of corpus-informed pedagogy lies in combining the empirical strengths of corpus linguistics with the interpretive and interactive affordances of GenAI, while maintaining a clear focus on pedagogy, transparency, and responsible use.
Recent developments in multimodal generative AI (GenAI) have led to the rapid expansion of tools and practices in language education. As a result, researchers, teachers, and learners are increasingly questioning how these technologies can be meaningfully integrated with established approaches such as corpus-based methods in pedagogical design. This is especially true in English for Specific Purposes (ESP), where some teachers may feel that GenAI could replace them as a source of guidance on discipline-specific language use.
This paper proposes a pedagogical framework for integrating corpus linguistics and GenAI into language teaching, grounded in established models of course design: needs analysis, language and learning objectives, materials and methods, and evaluation. Within this framework, corpus methods are positioned as a means of grounding pedagogical decisions in empirical evidence of language use, while GenAI tools are used to support the interpretation, accessibility, and interactive exploration of that evidence, as well as the evaluation of learner development.
To illustrate the framework, the paper draws on a series of classroom-informed examples, including corpus-based analysis of learner needs, discipline-specific language modeling using specialized corpora, data-driven learning (DDL) activities supported by AI tools, and corpus-informed evaluation techniques for identifying prototypical and problematic learner texts. Together, these examples demonstrate how corpus methods and AI can be integrated across all stages of course design in a coherent and pedagogically meaningful way.
The paper also considers challenges associated with integrating corpus methods and GenAI, including issues of reliability, bias, and the potential for tool-driven pedagogy. In response, it emphasizes the importance of aligning technological choices with clearly defined learning goals.
The paper concludes by suggesting that the future of corpus-informed pedagogy lies in combining the empirical strengths of corpus linguistics with the interpretive and interactive affordances of GenAI, while maintaining a clear focus on pedagogy, transparency, and responsible use.
Large language models (LLMs) are often assumed to possess register awareness (RA) because they can generate texts that appear to belong to specific registers. However, in LLMs, RA means the ability to generate texts that display the lexicogrammar found in the intended human registers, and this ability has been demonstrated by previous research to be lacking to a large extent (Author, 2024; Goulart et al., 2024; Mizumoto et al., 2024). Specifically, RA is evidenced by the realization of pervasive grammatical features that characterize human registers in comparison to others. Previous corpus-based studies have shown that AI-generated texts diverge from human texts with respect to such features, raising questions about whether RA can be improved through external intervention. To investigate this, our study looks at whether prompting and fine-tuning can induce LLMs to reproduce the pervasive grammatical features that characterize human academic registers, thereby improving their RA. To do that, we resorted to a corpus of human-authored and AI-generated research articles in Applied Linguistics, Biology, and Chemistry, totaling 7,200 texts and 3.6 million words. The AI corpora were generated by GPT and Gemini through baseline generation, grammar-informed prompting, and fine-tuning. To compare the human and AI texts, we applied Key Feature Analysis (KFA) (Egbert & Biber, 2023), which identified the grammatical features that most strongly characterize each research article section in each discipline. The KFA shows that both prompting and fine-tuning increased the realization of human key features, although the effects vary. Some features proved more resistant to modification, particularly modal verbs, complement clauses, personal pronouns, and other grammatical resources associated with authorial stance and interpretation. In contrast, the features most readily induced were nominalizations, passive constructions, technical / abstract nouns, and resources associated with informational and procedural discourse.
References
Author. (2024).
Egbert, J., & Biber, D. (2023). Key feature analysis: A simple, yet powerful method for comparing
text varieties. Corpora, 18(1), 121–133. https://doi.org/10.3366/cor.2023.0271
Goulart, L., Matte, M. L., Mendoza, A., Alvarado, L., & Veloso, I. (2024). AI or student writing? Analyzing the situational and linguistic characteristics of undergraduate student writing and AI-generated assignments. Journal of Second Language Writing, 66(101160), 1–19.
https://doi.org/10.1016/j.jslw.2024.101160
Mizumoto, A., Yasuda, S., & Tamura, Y. (2024). Identifying ChatGPT-generated texts in EFL students’ writing: Through comparative analysis of linguistic fingerprints. Applied Corpus Linguistics, 4(3), 100106.
Large language models (LLMs) are often assumed to possess register awareness (RA) because they can generate texts that appear to belong to specific registers. However, in LLMs, RA means the ability to generate texts that display the lexicogrammar found in the intended human registers, and this ability has been demonstrated by previous research to be lacking to a large extent (Author, 2024; Goulart et al., 2024; Mizumoto et al., 2024). Specifically, RA is evidenced by the realization of pervasive grammatical features that characterize human registers in comparison to others. Previous corpus-based studies have shown that AI-generated texts diverge from human texts with respect to such features, raising questions about whether RA can be improved through external intervention. To investigate this, our study looks at whether prompting and fine-tuning can induce LLMs to reproduce the pervasive grammatical features that characterize human academic registers, thereby improving their RA. To do that, we resorted to a corpus of human-authored and AI-generated research articles in Applied Linguistics, Biology, and Chemistry, totaling 7,200 texts and 3.6 million words. The AI corpora were generated by GPT and Gemini through baseline generation, grammar-informed prompting, and fine-tuning. To compare the human and AI texts, we applied Key Feature Analysis (KFA) (Egbert & Biber, 2023), which identified the grammatical features that most strongly characterize each research article section in each discipline. The KFA shows that both prompting and fine-tuning increased the realization of human key features, although the effects vary. Some features proved more resistant to modification, particularly modal verbs, complement clauses, personal pronouns, and other grammatical resources associated with authorial stance and interpretation. In contrast, the features most readily induced were nominalizations, passive constructions, technical / abstract nouns, and resources associated with informational and procedural discourse.
References
Author. (2024).
Egbert, J., & Biber, D. (2023). Key feature analysis: A simple, yet powerful method for comparing
text varieties. Corpora, 18(1), 121–133. https://doi.org/10.3366/cor.2023.0271
Goulart, L., Matte, M. L., Mendoza, A., Alvarado, L., & Veloso, I. (2024). AI or student writing? Analyzing the situational and linguistic characteristics of undergraduate student writing and AI-generated assignments. Journal of Second Language Writing, 66(101160), 1–19.
https://doi.org/10.1016/j.jslw.2024.101160
Mizumoto, A., Yasuda, S., & Tamura, Y. (2024). Identifying ChatGPT-generated texts in EFL students’ writing: Through comparative analysis of linguistic fingerprints. Applied Corpus Linguistics, 4(3), 100106.
The global environmental crisis has intensified calls for a transition toward a green economy, making public communication about sustainability increasingly significant. In Indonesia, the green economy has become a central component of national development agendas, including commitments toward net-zero emissions by 2060. A successful transition toward a green economy depends not only on policy and technology but also on how the concept is communicated to and understood by the public (Death, 2018); yet, the linguistic construction of this discourse in Indonesian media remains underexplored. This study investigates how the green economy is constructed in Indonesian online news media.
Drawing on a specialised corpus of online news articles published between 2021 and 2025 from five major Indonesian news websites, comprising 623,050 tokens, this study adopts a corpus-assisted ecolinguistic approach (Poole, 2022). This approach combines quantitative corpus methods with qualitative ecolinguistic interpretation to examine how repeated linguistic patterns construct understandings of the relationship between economy, society, and the environment. The analysis was conducted using Sketch Engine. Keyword analysis was used to identify statistically salient words in the green economy corpus by comparing their normalised frequencies with the news genre sub-corpus of the Indonesian Web Corpus 2024 (idTenTen24), a large general Indonesian reference corpus of approximately 7.1 billion words available in Sketch Engine. Keywords were identified using Sketch Engine’s keyness score based on the simple maths method. The keyword results were then examined through collocation and concordance analyses.
Preliminary findings reveal that green economy discourse is dominated by language related to technology, economic growth, and public policy, particularly through keywords such as karbon (carbon), energi (energy), emisi (emissions), transisi (transition), and investasi (investment). The prominence of terms including transformasi (transformation), kolaborasi (collaboration), pembiayaan (financing), and hilirisasi (downstreaming) suggests that environmental issues are primarily depicted through developmental, industrial, and governance narratives rather than ecological justice or community-centred perspectives. The discourse also foregrounds institutional actors and market-based solutions, while ecological degradation and vulnerable communities remain relatively backgrounded. These patterns indicate an ambivalent ecological discourse (Stibbe, 2015), in which sustainability is closely connected to economic growth and national competitiveness.
This study is expected to contribute theoretically to the development of corpus-assisted ecolinguistics in the Indonesian context and practically to critical environmental journalism by demonstrating how media discourse shapes public understandings of sustainability and green transition policies.
Keywords: green economy; corpus-assisted ecolinguistics; news representation; sustainable development awareness
References
Death, C. (2018). Four discourses of the green economy in the global South. In The Green Economy in the Global South (pp. 11-28). Routledge.
Poole, R. (2022). Corpus-assisted ecolinguistics. Bloomsbury Academic.
Stibbe, Arran. (2015). Ecolinguistics: Language, ecology and the stories we live by. London: Routledge.
The global environmental crisis has intensified calls for a transition toward a green economy, making public communication about sustainability increasingly significant. In Indonesia, the green economy has become a central component of national development agendas, including commitments toward net-zero emissions by 2060. A successful transition toward a green economy depends not only on policy and technology but also on how the concept is communicated to and understood by the public (Death, 2018); yet, the linguistic construction of this discourse in Indonesian media remains underexplored. This study investigates how the green economy is constructed in Indonesian online news media.
Drawing on a specialised corpus of online news articles published between 2021 and 2025 from five major Indonesian news websites, comprising 623,050 tokens, this study adopts a corpus-assisted ecolinguistic approach (Poole, 2022). This approach combines quantitative corpus methods with qualitative ecolinguistic interpretation to examine how repeated linguistic patterns construct understandings of the relationship between economy, society, and the environment. The analysis was conducted using Sketch Engine. Keyword analysis was used to identify statistically salient words in the green economy corpus by comparing their normalised frequencies with the news genre sub-corpus of the Indonesian Web Corpus 2024 (idTenTen24), a large general Indonesian reference corpus of approximately 7.1 billion words available in Sketch Engine. Keywords were identified using Sketch Engine’s keyness score based on the simple maths method. The keyword results were then examined through collocation and concordance analyses.
Preliminary findings reveal that green economy discourse is dominated by language related to technology, economic growth, and public policy, particularly through keywords such as karbon (carbon), energi (energy), emisi (emissions), transisi (transition), and investasi (investment). The prominence of terms including transformasi (transformation), kolaborasi (collaboration), pembiayaan (financing), and hilirisasi (downstreaming) suggests that environmental issues are primarily depicted through developmental, industrial, and governance narratives rather than ecological justice or community-centred perspectives. The discourse also foregrounds institutional actors and market-based solutions, while ecological degradation and vulnerable communities remain relatively backgrounded. These patterns indicate an ambivalent ecological discourse (Stibbe, 2015), in which sustainability is closely connected to economic growth and national competitiveness.
This study is expected to contribute theoretically to the development of corpus-assisted ecolinguistics in the Indonesian context and practically to critical environmental journalism by demonstrating how media discourse shapes public understandings of sustainability and green transition policies.
Keywords: green economy; corpus-assisted ecolinguistics; news representation; sustainable development awareness
References
Death, C. (2018). Four discourses of the green economy in the global South. In The Green Economy in the Global South (pp. 11-28). Routledge.
Poole, R. (2022). Corpus-assisted ecolinguistics. Bloomsbury Academic.
Stibbe, Arran. (2015). Ecolinguistics: Language, ecology and the stories we live by. London: Routledge.
This study compares the use of role language in the school slice-of-life manga Skip and Loafer and the action-adventure manga ONE PIECE. It examines how manga genre affects the density, linguistic forms, and functions of role language as character-indexing expressions.
The data consist of 974 utterances from Skip and Loafer and 1,552 utterances from ONE PIECE. Role-language expressions were extracted from each work and classified according to character attributes and linguistic forms. The linguistic forms include sentence-final particles, personal pronouns, dialectal forms, sound changes, imperative expressions, special vocabulary, and character-specific expressions. The character attributes include gender, age/generation, social refinement, region, historical period, occupation/social class, personality, and non-human identity.
The analysis reveals clear differences between the two genres. In Skip and Loafer, 154 instances of role language were identified, corresponding to 15.8 instances per 100 utterances. In contrast, ONE PIECE contained 1,250 instances, corresponding to 80.5 instances per 100 utterances. In Skip and Loafer, role language mainly appears in everyday school-life contexts and subtly differentiates characters and interpersonal relationships through sentence-final particles, personal pronouns, and sound changes. By contrast, in ONE PIECE, role language frequently appears in battle, escape, command, and confrontation scenes. Imperative expressions and character-specific expressions contribute to the construction of conflict, hierarchy, power relations, and non-ordinary character attributes.
These findings suggest that role language performs different functions depending on manga genre. In school slice-of-life manga, it functions as a resource for differentiating characters in everyday interactions. In action-adventure manga, it serves as a central linguistic resource for constructing diverse characters and a stylized fictional world.
This study compares the use of role language in the school slice-of-life manga Skip and Loafer and the action-adventure manga ONE PIECE. It examines how manga genre affects the density, linguistic forms, and functions of role language as character-indexing expressions.
The data consist of 974 utterances from Skip and Loafer and 1,552 utterances from ONE PIECE. Role-language expressions were extracted from each work and classified according to character attributes and linguistic forms. The linguistic forms include sentence-final particles, personal pronouns, dialectal forms, sound changes, imperative expressions, special vocabulary, and character-specific expressions. The character attributes include gender, age/generation, social refinement, region, historical period, occupation/social class, personality, and non-human identity.
The analysis reveals clear differences between the two genres. In Skip and Loafer, 154 instances of role language were identified, corresponding to 15.8 instances per 100 utterances. In contrast, ONE PIECE contained 1,250 instances, corresponding to 80.5 instances per 100 utterances. In Skip and Loafer, role language mainly appears in everyday school-life contexts and subtly differentiates characters and interpersonal relationships through sentence-final particles, personal pronouns, and sound changes. By contrast, in ONE PIECE, role language frequently appears in battle, escape, command, and confrontation scenes. Imperative expressions and character-specific expressions contribute to the construction of conflict, hierarchy, power relations, and non-ordinary character attributes.
These findings suggest that role language performs different functions depending on manga genre. In school slice-of-life manga, it functions as a resource for differentiating characters in everyday interactions. In action-adventure manga, it serves as a central linguistic resource for constructing diverse characters and a stylized fictional world.
Outcome-based writing assessment has become less dependable in the generative-AI era because submitted texts may not fully represent learners’ current knowledge, skills, and abilities. Direct AI correction can also hide the learning process by replacing learners’ search for evidence with an answer. This presentation reports on WATTLE, a learner-oriented system that integrates data-driven learning, corpus evidence, and dynamic assessment to support phraseological revision in L2 writing. The study addressed two research questions: (1) to what extent corpus-mediated interaction supported phraseological revision outcomes, and (2) what dimensions of corpus-evidenced phraseological revision competence emerged from learners’ interaction traces.
The study used a mixed-method exploratory case-study design with 68 Japanese university EFL learners. Learners first wrote an initial draft. WATTLE then identified phraseological candidates, including verb-noun collocations, adjective-noun collocations, noun-preposition patterns, and recurrent academic phrase frames. Rather than correcting the text, the system prompted learners to create corpus queries, inspect concordance lines, revise their expressions, and justify their decisions. Data comprised initial drafts, revised texts, corpus queries, concordance evidence, learner-AI interaction logs, and written justifications. Revision outcomes were examined descriptively by comparing pre- and post-revision accuracy, with reference to a CEFR-informed writing rubric, while interaction traces and justifications were analyzed thematically.
The findings showed that corpus-mediated interaction was associated with more target-like phraseological revisions and identified inquiry patterns that characterized more successful revision processes. The analysis also suggested dimensions of corpus-evidenced phraseological revision competence, including the ability to formulate useful corpus queries, interpret recurrent phraseological patterns, select contextually appropriate alternatives, and justify revisions with evidence. The main contribution is a transparent, process-oriented model of formative assessment in which AI mediates access to corpus evidence instead of producing opaque corrections. The study also offers an empirical basis for phraseological-competence descriptors and rubrics grounded in observable learner engagement with corpus evidence.
Outcome-based writing assessment has become less dependable in the generative-AI era because submitted texts may not fully represent learners’ current knowledge, skills, and abilities. Direct AI correction can also hide the learning process by replacing learners’ search for evidence with an answer. This presentation reports on WATTLE, a learner-oriented system that integrates data-driven learning, corpus evidence, and dynamic assessment to support phraseological revision in L2 writing. The study addressed two research questions: (1) to what extent corpus-mediated interaction supported phraseological revision outcomes, and (2) what dimensions of corpus-evidenced phraseological revision competence emerged from learners’ interaction traces.
The study used a mixed-method exploratory case-study design with 68 Japanese university EFL learners. Learners first wrote an initial draft. WATTLE then identified phraseological candidates, including verb-noun collocations, adjective-noun collocations, noun-preposition patterns, and recurrent academic phrase frames. Rather than correcting the text, the system prompted learners to create corpus queries, inspect concordance lines, revise their expressions, and justify their decisions. Data comprised initial drafts, revised texts, corpus queries, concordance evidence, learner-AI interaction logs, and written justifications. Revision outcomes were examined descriptively by comparing pre- and post-revision accuracy, with reference to a CEFR-informed writing rubric, while interaction traces and justifications were analyzed thematically.
The findings showed that corpus-mediated interaction was associated with more target-like phraseological revisions and identified inquiry patterns that characterized more successful revision processes. The analysis also suggested dimensions of corpus-evidenced phraseological revision competence, including the ability to formulate useful corpus queries, interpret recurrent phraseological patterns, select contextually appropriate alternatives, and justify revisions with evidence. The main contribution is a transparent, process-oriented model of formative assessment in which AI mediates access to corpus evidence instead of producing opaque corrections. The study also offers an empirical basis for phraseological-competence descriptors and rubrics grounded in observable learner engagement with corpus evidence.
Although prior research has shown that different tasks elicit distinct linguistic features, it remains unclear whether task-induced differences remain stable, diverge, or reverse over extended periods of L2 development. This study, therefore, investigates how task type affects longitudinal patterns of linguistic feature use in L2 spoken English.
Data were drawn from 104 Japanese EFL learners aged 16–18 (CEFR A2) who completed five monologic tasks: reasoning, description, present narration, past narration, and argumentation. Speech was collected at eight time points over 23 months, yielding the Longitudinal Corpus of Spoken English (LOCSE; approximately 400,000 tokens; Abe & Kondo, 2019).
Using the Multidimensional Analysis Tagger (Nini, 2019), six features relevant to discourse organization were selected: discourse cohesion (causal ‘because’), informational density (predicative/attributive adjectives, nouns), event elaboration (adverbs), and stance/modality (possibility modals). Mixed-effects zero-inflated negative binomial regression models included task type, months of learning, and their interaction as fixed effects, with random intercepts and slopes per learner (Murakami, 2025; Winter & Buerkner, 2021).
Developmental trends varied across task types. Two contrasts were salient. First, argumentation tasks elicited frequent use of most features except adverbs. Second, a reversed pattern emerged between present narration and argumentation: adverbs predominated in narration (e.g., "Mm I often play the games.") while predicative and attributive adjectives predominated in argumentation (e.g., "But swimming pool is so easy because ... pool has many interesting attractions.").
These contrasting patterns reflect differing discourse demands: adjectives increase informational density in expository registers, while adverbs enrich lexical verbs in narrative discourse (Biber et al., 2021). Narration tasks elicited more adverbs as learners recounted personal events, yielding more verb-centered speech; argumentation tasks prompted adjective use as learners evaluated entities to support their claims. Together, these task-specific longitudinal patterns provide empirical support for principled task sequencing in L2 speaking instruction.
Although prior research has shown that different tasks elicit distinct linguistic features, it remains unclear whether task-induced differences remain stable, diverge, or reverse over extended periods of L2 development. This study, therefore, investigates how task type affects longitudinal patterns of linguistic feature use in L2 spoken English.
Data were drawn from 104 Japanese EFL learners aged 16–18 (CEFR A2) who completed five monologic tasks: reasoning, description, present narration, past narration, and argumentation. Speech was collected at eight time points over 23 months, yielding the Longitudinal Corpus of Spoken English (LOCSE; approximately 400,000 tokens; Abe & Kondo, 2019).
Using the Multidimensional Analysis Tagger (Nini, 2019), six features relevant to discourse organization were selected: discourse cohesion (causal ‘because’), informational density (predicative/attributive adjectives, nouns), event elaboration (adverbs), and stance/modality (possibility modals). Mixed-effects zero-inflated negative binomial regression models included task type, months of learning, and their interaction as fixed effects, with random intercepts and slopes per learner (Murakami, 2025; Winter & Buerkner, 2021).
Developmental trends varied across task types. Two contrasts were salient. First, argumentation tasks elicited frequent use of most features except adverbs. Second, a reversed pattern emerged between present narration and argumentation: adverbs predominated in narration (e.g., "Mm I often play the games.") while predicative and attributive adjectives predominated in argumentation (e.g., "But swimming pool is so easy because ... pool has many interesting attractions.").
These contrasting patterns reflect differing discourse demands: adjectives increase informational density in expository registers, while adverbs enrich lexical verbs in narrative discourse (Biber et al., 2021). Narration tasks elicited more adverbs as learners recounted personal events, yielding more verb-centered speech; argumentation tasks prompted adjective use as learners evaluated entities to support their claims. Together, these task-specific longitudinal patterns provide empirical support for principled task sequencing in L2 speaking instruction.
This study examines how Japanese EFL learners construct the opening and closing of opinion essays — specifically, how thesis sentences and conclusion sentences relate to each other as the same learners advance through school. A longitudinal corpus of 7,552 essays produced by a single cohort of Japanese learners across three consecutive school years in response to 42 opinion prompts was analyzed with a large language model (GPT-5.5) prompted to extract, for each essay, the thesis sentence and the conclusion sentence, and to score their semantic and syntactic similarity on a 0–1 scale. Lexical overlap was independently computed as a lemma-level Jaccard coefficient using spaCy. Frequency lists and Dunning's log-likelihood (G²) keyness were computed over the more than 6,800 thesis–conclusion pairs in which both sentences were present.
Results show that the opening–closing frame is robust across the cohort: over 95% of essays contain an identifiable thesis and over 90% an identifiable conclusion, with the thesis appearing as the first sentence and the conclusion as the last in roughly nine out of ten essays. However, as the same learners progress from Grade 1 to Grade 3, the lexical and structural similarity between the two sentences declines steadily: mean semantic similarity, syntactic similarity, and lemma-level Jaccard overlap all decrease over time, indicating a within-cohort developmental shift from near-verbatim repetition toward genuine paraphrase. Keyness analysis further reveals a functional asymmetry: thesis sentences over-represent topical and stance content (e.g. problem, agree, in order to), whereas conclusions are dominated by metadiscoursal connectors (e.g. for these reasons, therefore, in conclusion). Together, the findings suggest that Japanese EFL learners reliably reproduce a conventional macro-structure from the outset, while their linguistic realization of closure, initially formulaic, becomes steadily more flexible as proficiency develops.
This study examines how Japanese EFL learners construct the opening and closing of opinion essays — specifically, how thesis sentences and conclusion sentences relate to each other as the same learners advance through school. A longitudinal corpus of 7,552 essays produced by a single cohort of Japanese learners across three consecutive school years in response to 42 opinion prompts was analyzed with a large language model (GPT-5.5) prompted to extract, for each essay, the thesis sentence and the conclusion sentence, and to score their semantic and syntactic similarity on a 0–1 scale. Lexical overlap was independently computed as a lemma-level Jaccard coefficient using spaCy. Frequency lists and Dunning's log-likelihood (G²) keyness were computed over the more than 6,800 thesis–conclusion pairs in which both sentences were present.
Results show that the opening–closing frame is robust across the cohort: over 95% of essays contain an identifiable thesis and over 90% an identifiable conclusion, with the thesis appearing as the first sentence and the conclusion as the last in roughly nine out of ten essays. However, as the same learners progress from Grade 1 to Grade 3, the lexical and structural similarity between the two sentences declines steadily: mean semantic similarity, syntactic similarity, and lemma-level Jaccard overlap all decrease over time, indicating a within-cohort developmental shift from near-verbatim repetition toward genuine paraphrase. Keyness analysis further reveals a functional asymmetry: thesis sentences over-represent topical and stance content (e.g. problem, agree, in order to), whereas conclusions are dominated by metadiscoursal connectors (e.g. for these reasons, therefore, in conclusion). Together, the findings suggest that Japanese EFL learners reliably reproduce a conventional macro-structure from the outset, while their linguistic realization of closure, initially formulaic, becomes steadily more flexible as proficiency develops.
This paper reports a corpus-assisted study of student reflection papers from an English-medium critical thinking course at a Japanese university. It examines how students linguistically represent critical thinking across six sequenced tasks that move from general understandings of critical thinking to judgment, transfer, reasoning standards, distributed cognition, and synthesis. The corpus consists of approximately 240 reflection papers written by 40 students, with each text averaging 500 to 700 words. Student reflections are treated as evidence of conceptual engagement, drawing on reflective writing and critical thinking pedagogy (Halpern, 1998; Hatton & Smith, 1995; Paul & Elder, 2007).
The reflection papers are treated as a small, specialized learner-corpus divided into six prompt-based subcorpora. Following corpus-assisted discourse analysis and small-corpus approaches, the analysis examines normalized frequency, keywords, n-grams, and collocations across the subcorpora (Baker, 2023; Koester, 2022). It then uses KWIC concordance analysis of terms related to judgment, evidence, transfer, assumptions, standards, tools, language, people, and environment. Because these terms partly reflect the assignment prompts, selected concordance lines are examined to distinguish prompt repetition from student-generated conceptual use. These lines are coded for concept naming, definition, application, author or concept distinction, self-regulation, and ecological reasoning.
Rubric scores for four concept-focused areas are used as document-level metadata but not as a primary outcome variable. Preliminary analysis indicates uneven concept-focused performance, with a decline on a complex transfer task, the distributed-cognition reflection. This pattern makes corpus analysis central to the study because it asks whether students’ language shifts from general references to thinking and understanding toward more applied and relational uses of course concepts. The paper shows how a small classroom corpus can support fine-grained analysis of conceptual uptake in EMI contexts and help instructors identify whether students reproduce course terminology or use concepts to interpret thinking in situated activity.
REFERENCES
Baker, P. (2023). Using corpora in discourse analysis (2nd ed.). Bloomsbury Academic.
Halpern, D. F. (1998). Teaching critical thinking for transfer across domains: Dispositions, skills, structure training, and metacognitive monitoring. American Psychologist, 53(4), 449–455. https://doi.org/10.1037/0003-066X.53.4.449
Hatton, N., & Smith, D. (1995). Reflection in teacher education: Towards definition and implementation. Teaching and Teacher Education, 11(1), 33–49. https://doi.org/10.1016/0742-051X(94)00012-U
Koester, A. (2010). Building small specialised corpora. In A. O’Keeffe & M. McCarthy (Eds.), The Routledge handbook of corpus linguistics (pp. 66–79). Routledge. https://doi.org/10.4324/9780203856949.ch5
Paul, R., & Elder, L. (2007). Critical thinking competency standards: Standards, principles, performance indicators, and outcomes with a critical thinking master rubric. Foundation for Critical Thinking Press.
This paper reports a corpus-assisted study of student reflection papers from an English-medium critical thinking course at a Japanese university. It examines how students linguistically represent critical thinking across six sequenced tasks that move from general understandings of critical thinking to judgment, transfer, reasoning standards, distributed cognition, and synthesis. The corpus consists of approximately 240 reflection papers written by 40 students, with each text averaging 500 to 700 words. Student reflections are treated as evidence of conceptual engagement, drawing on reflective writing and critical thinking pedagogy (Halpern, 1998; Hatton & Smith, 1995; Paul & Elder, 2007).
The reflection papers are treated as a small, specialized learner-corpus divided into six prompt-based subcorpora. Following corpus-assisted discourse analysis and small-corpus approaches, the analysis examines normalized frequency, keywords, n-grams, and collocations across the subcorpora (Baker, 2023; Koester, 2022). It then uses KWIC concordance analysis of terms related to judgment, evidence, transfer, assumptions, standards, tools, language, people, and environment. Because these terms partly reflect the assignment prompts, selected concordance lines are examined to distinguish prompt repetition from student-generated conceptual use. These lines are coded for concept naming, definition, application, author or concept distinction, self-regulation, and ecological reasoning.
Rubric scores for four concept-focused areas are used as document-level metadata but not as a primary outcome variable. Preliminary analysis indicates uneven concept-focused performance, with a decline on a complex transfer task, the distributed-cognition reflection. This pattern makes corpus analysis central to the study because it asks whether students’ language shifts from general references to thinking and understanding toward more applied and relational uses of course concepts. The paper shows how a small classroom corpus can support fine-grained analysis of conceptual uptake in EMI contexts and help instructors identify whether students reproduce course terminology or use concepts to interpret thinking in situated activity.
REFERENCES
Baker, P. (2023). Using corpora in discourse analysis (2nd ed.). Bloomsbury Academic.
Halpern, D. F. (1998). Teaching critical thinking for transfer across domains: Dispositions, skills, structure training, and metacognitive monitoring. American Psychologist, 53(4), 449–455. https://doi.org/10.1037/0003-066X.53.4.449
Hatton, N., & Smith, D. (1995). Reflection in teacher education: Towards definition and implementation. Teaching and Teacher Education, 11(1), 33–49. https://doi.org/10.1016/0742-051X(94)00012-U
Koester, A. (2010). Building small specialised corpora. In A. O’Keeffe & M. McCarthy (Eds.), The Routledge handbook of corpus linguistics (pp. 66–79). Routledge. https://doi.org/10.4324/9780203856949.ch5
Paul, R., & Elder, L. (2007). Critical thinking competency standards: Standards, principles, performance indicators, and outcomes with a critical thinking master rubric. Foundation for Critical Thinking Press.
The Science of Simplification: Disciplinary Linguistic Dimensions in Three-Minute Thesis Presentations
Three-Minute Thesis (3MT) presentations is a high-stakes academic spoken genre requiring researchers to communicate complex disciplinary knowledge to lay audiences within three minutes. While interdisciplinary differences have been reported at the discourse level (Hyland & Zou, 2021; Sun et al., 2024), Carter-Thomas and Rowley-Jolivet (2020) report genre-level homogeneity. No large-scale computational study has yet examined 3MT's overall linguistic structure to empirically resolve this tension. This study addresses this gap by asking: what dimensions of linguistic variation characterize 3MT presentations, and how do these vary across disciplinary categories?
Drawing on a purpose-built corpus of 1,045 presentations spanning 70 universities and 92 disciplines, 159 linguistic features were extracted using the Multi-Feature Tagger for English (MFTE) and classified into five Biglan-inspired disciplinary categories. Following principled feature selection (KMO = 0.718), Exploratory Factor Analysis (Principal Axis Factoring, promax rotation) identified four dimensions: Abstract Knowledge Framing (D1), Technical Quantification vs General Reference (D2), Causal Narrative Elaboration (D3), and Interactional Involvement (D4). Complementary Kruskal-Wallis test with Dunn's post-hoc analysis revealed significant disciplinary variation in 37 features, with a robust Hard vs Soft disciplinary divide and hybrid disciplines consistently occupying an intermediate position.
Findings reveal that while genre constraints suppress systematic co-occurrence patterns at the dimensional level, disciplinary identity is robustly preserved at the individual feature level, resolving an apparent tension in the existing literature. The identified disciplinary profiles and interactional patterns illuminate how researchers simplify complex knowledge for lay audiences in discipline-specific ways, offering actionable insights for pedagogy and science communication training for novice researchers.
Keywords: Three-minute thesis presentation, multi-dimensional analysis, corpus linguistics
The Science of Simplification: Disciplinary Linguistic Dimensions in Three-Minute Thesis Presentations
Three-Minute Thesis (3MT) presentations is a high-stakes academic spoken genre requiring researchers to communicate complex disciplinary knowledge to lay audiences within three minutes. While interdisciplinary differences have been reported at the discourse level (Hyland & Zou, 2021; Sun et al., 2024), Carter-Thomas and Rowley-Jolivet (2020) report genre-level homogeneity. No large-scale computational study has yet examined 3MT's overall linguistic structure to empirically resolve this tension. This study addresses this gap by asking: what dimensions of linguistic variation characterize 3MT presentations, and how do these vary across disciplinary categories?
Drawing on a purpose-built corpus of 1,045 presentations spanning 70 universities and 92 disciplines, 159 linguistic features were extracted using the Multi-Feature Tagger for English (MFTE) and classified into five Biglan-inspired disciplinary categories. Following principled feature selection (KMO = 0.718), Exploratory Factor Analysis (Principal Axis Factoring, promax rotation) identified four dimensions: Abstract Knowledge Framing (D1), Technical Quantification vs General Reference (D2), Causal Narrative Elaboration (D3), and Interactional Involvement (D4). Complementary Kruskal-Wallis test with Dunn's post-hoc analysis revealed significant disciplinary variation in 37 features, with a robust Hard vs Soft disciplinary divide and hybrid disciplines consistently occupying an intermediate position.
Findings reveal that while genre constraints suppress systematic co-occurrence patterns at the dimensional level, disciplinary identity is robustly preserved at the individual feature level, resolving an apparent tension in the existing literature. The identified disciplinary profiles and interactional patterns illuminate how researchers simplify complex knowledge for lay audiences in discipline-specific ways, offering actionable insights for pedagogy and science communication training for novice researchers.
Keywords: Three-minute thesis presentation, multi-dimensional analysis, corpus linguistics
The rapid development of generative AI (GenAI) invites corpus linguistics to reconsider its role in applied language research. In second language (L2) writing assessment, learner corpora can do more than document patterns of learner language: they can provide an empirical basis for examining how AI systems score, justify, and explain learner performance. This presentation demonstrates how learner corpora and corpus-based methods can be used to validate scores generated by GenAI.
The study is based on a corpus of 500 examination essays written by learners of English as a foreign language. The texts were produced under controlled exam conditions in response to five writing prompts and were assessed by two trained human raters each. The rating procedure used an analytic scale covering four dimensions of writing proficiency: content, organisation and cohesion, grammatical accuracy, and lexical range. The same essays were subsequently evaluated by five generative AI systems under three prompting conditions.
The analysis compares human and AI-generated scores across criteria, prompts, and models, and relates both sets of scores to corpus-based measures of learner writing. These include indices of lexical diversity and sophistication, phraseological patterning, syntactic complexity, discourse organisation, and accuracy. The results indicate that GenAI can reproduce broad human rating patterns. At the same time, the relationship between scores and linguistic features is complex. AI-generated scores tend to converge more closely with human scores than with any single group of corpus features, suggesting that GenAI assessment cannot be reduced to an aggregate of surface-level linguistic indicators. Nevertheless, corpus features help identify where AI scores align with, diverge from, or overgeneralise human judgements.
The paper argues that corpus linguistics is well placed to address this challenge. In the age of AI, corpora can serve not only as datasets, but also as instruments for validating, interrogating, and contextualising AI-mediated language assessment.
The rapid development of generative AI (GenAI) invites corpus linguistics to reconsider its role in applied language research. In second language (L2) writing assessment, learner corpora can do more than document patterns of learner language: they can provide an empirical basis for examining how AI systems score, justify, and explain learner performance. This presentation demonstrates how learner corpora and corpus-based methods can be used to validate scores generated by GenAI.
The study is based on a corpus of 500 examination essays written by learners of English as a foreign language. The texts were produced under controlled exam conditions in response to five writing prompts and were assessed by two trained human raters each. The rating procedure used an analytic scale covering four dimensions of writing proficiency: content, organisation and cohesion, grammatical accuracy, and lexical range. The same essays were subsequently evaluated by five generative AI systems under three prompting conditions.
The analysis compares human and AI-generated scores across criteria, prompts, and models, and relates both sets of scores to corpus-based measures of learner writing. These include indices of lexical diversity and sophistication, phraseological patterning, syntactic complexity, discourse organisation, and accuracy. The results indicate that GenAI can reproduce broad human rating patterns. At the same time, the relationship between scores and linguistic features is complex. AI-generated scores tend to converge more closely with human scores than with any single group of corpus features, suggesting that GenAI assessment cannot be reduced to an aggregate of surface-level linguistic indicators. Nevertheless, corpus features help identify where AI scores align with, diverge from, or overgeneralise human judgements.
The paper argues that corpus linguistics is well placed to address this challenge. In the age of AI, corpora can serve not only as datasets, but also as instruments for validating, interrogating, and contextualising AI-mediated language assessment.
Recent corpora contain scores assigned to learner language and are increasingly used to investigate relationships between linguistic features and human evaluation. However, the comparability of those findings remains unknown because different learner corpora use different tasks, rating scales, and raters. Addressing this methodological challenge, the present proof-of-concept study investigates whether GenAI can function as a common rater that links different corpora and establishes a shared measurement scale.
For this purpose, the study compared two corpora—the English Language Learner Insight, Proficiency and Skills Evaluation (ELLIPSE) corpus and the International Corpus Network of Asian Learners of English Global Rating Archives (ICNALE). The ELLIPSE corpus contains 8,890 essays rated by 27 human raters, and the ICNALE corpus contains 140 essays rated by 80 human raters. Essays from both corpora were scored by GPT-5.2 according to their respective rubrics, and these ratings were analyzed using a many-facet Rasch measurement approach to establish connectedness between datasets. The study further examined the plausibility of the relative rankings of writers by examining the relationship between writer ability measures and grammatical accuracy and lexical sophistication—two variables that are known to correlate with overall writing quality.
Results suggested that the linking procedure successfully placed the two corpora on a common scale, revealing that ELLIPSE writers were largely comparable to ICNALE’s B1-level writers but were outperformed by B2-level writers. Additionally, ICNALE raters were found to be more severe than ELLIPSE raters, likely because ELLIPSE raters adjusted their severity due to the nativelikeness criterion as a gold standard in the ELLIPSE rating scale. The linguistic analysis of essays partially supported the relative ranking of writers, suggesting that ELLIPSE writers were largely equivalent to ICNALE’s B1-level writers. These findings demonstrate the feasibility of using GenAI as an anchoring rater and highlight its potential to enhance comparability across different learner corpora.
Recent corpora contain scores assigned to learner language and are increasingly used to investigate relationships between linguistic features and human evaluation. However, the comparability of those findings remains unknown because different learner corpora use different tasks, rating scales, and raters. Addressing this methodological challenge, the present proof-of-concept study investigates whether GenAI can function as a common rater that links different corpora and establishes a shared measurement scale.
For this purpose, the study compared two corpora—the English Language Learner Insight, Proficiency and Skills Evaluation (ELLIPSE) corpus and the International Corpus Network of Asian Learners of English Global Rating Archives (ICNALE). The ELLIPSE corpus contains 8,890 essays rated by 27 human raters, and the ICNALE corpus contains 140 essays rated by 80 human raters. Essays from both corpora were scored by GPT-5.2 according to their respective rubrics, and these ratings were analyzed using a many-facet Rasch measurement approach to establish connectedness between datasets. The study further examined the plausibility of the relative rankings of writers by examining the relationship between writer ability measures and grammatical accuracy and lexical sophistication—two variables that are known to correlate with overall writing quality.
Results suggested that the linking procedure successfully placed the two corpora on a common scale, revealing that ELLIPSE writers were largely comparable to ICNALE’s B1-level writers but were outperformed by B2-level writers. Additionally, ICNALE raters were found to be more severe than ELLIPSE raters, likely because ELLIPSE raters adjusted their severity due to the nativelikeness criterion as a gold standard in the ELLIPSE rating scale. The linguistic analysis of essays partially supported the relative ranking of writers, suggesting that ELLIPSE writers were largely equivalent to ICNALE’s B1-level writers. These findings demonstrate the feasibility of using GenAI as an anchoring rater and highlight its potential to enhance comparability across different learner corpora.
With the amplification of anti-vaccine sentiment related to Covid-19, anti-vax discourse (especially on social media) has undergone extensive scrutiny while longitudinal examinations of how the problem of anti-vaccination is represented in journalistic discourse are scarce. This study addresses this gap by adopting a corpus-based discourse approach to examine the historical development of and possible shifts in representations of anti-vaccination in the mainstream New Zealand press. More specifically, it analyses how anti-vaccine positions are represented and (dis)legitimised and how “anti-vaxxer” identities are attributed or rejected in media discourse by tracing the use of salient identity markers such as “anti-vax” and linguistic devices such as reported speech.
We compiled the corpus by searching for news articles on anti-vaccination in the Newztext – a long-term archive of key New Zealand news dating from 1960 to the present. Our search returned 9,243 articles spanning 27 years from 1997 to 2023, of which 2,126 articles (totalling over 1.8 million words) were included in the corpus after irrelevant articles and duplicates were removed.
Our preliminary findings suggest a shift over time in the NZ press from “anti-immunisation” to “anti-vax.” The term “Anti-vax” first peaked in 2015, amid heated public debate on the MMR vaccine for measles; simultaneously, the discussion of vaccine expanded beyond medical experts and health professionals to include a wider range of social actors as reflected in the voices quoted in the reports. We also observe notable differences in how people labelled as “anti-immunisation” and “anti-vax” are categorised. Identity-based attributes associated with the “anti-immunisation” category appear moderate (e.g., “misguided people,” “have a point but missing the larger issue”), while language used to categorise the “anti-vaxxer” tends to be more forceful and emotive (e.g., they are “dangerous” ,“dumb”, “crazy and evil”, etc.) Our findings show that media discourse on anti-vaccination in New Zealand has shifted from an “policy -oriented” public health discussion to a more affective discourse shaped by identity-politics and wider public contestation. The “anti-vax” labelling and the expansion of this labelling in media discourse may further contribute to polarisation and social division as previous research has suggested.
With the amplification of anti-vaccine sentiment related to Covid-19, anti-vax discourse (especially on social media) has undergone extensive scrutiny while longitudinal examinations of how the problem of anti-vaccination is represented in journalistic discourse are scarce. This study addresses this gap by adopting a corpus-based discourse approach to examine the historical development of and possible shifts in representations of anti-vaccination in the mainstream New Zealand press. More specifically, it analyses how anti-vaccine positions are represented and (dis)legitimised and how “anti-vaxxer” identities are attributed or rejected in media discourse by tracing the use of salient identity markers such as “anti-vax” and linguistic devices such as reported speech.
We compiled the corpus by searching for news articles on anti-vaccination in the Newztext – a long-term archive of key New Zealand news dating from 1960 to the present. Our search returned 9,243 articles spanning 27 years from 1997 to 2023, of which 2,126 articles (totalling over 1.8 million words) were included in the corpus after irrelevant articles and duplicates were removed.
Our preliminary findings suggest a shift over time in the NZ press from “anti-immunisation” to “anti-vax.” The term “Anti-vax” first peaked in 2015, amid heated public debate on the MMR vaccine for measles; simultaneously, the discussion of vaccine expanded beyond medical experts and health professionals to include a wider range of social actors as reflected in the voices quoted in the reports. We also observe notable differences in how people labelled as “anti-immunisation” and “anti-vax” are categorised. Identity-based attributes associated with the “anti-immunisation” category appear moderate (e.g., “misguided people,” “have a point but missing the larger issue”), while language used to categorise the “anti-vaxxer” tends to be more forceful and emotive (e.g., they are “dangerous” ,“dumb”, “crazy and evil”, etc.) Our findings show that media discourse on anti-vaccination in New Zealand has shifted from an “policy -oriented” public health discussion to a more affective discourse shaped by identity-politics and wider public contestation. The “anti-vax” labelling and the expansion of this labelling in media discourse may further contribute to polarisation and social division as previous research has suggested.
Multidimensional analysis (MD analysis) is a method of linguistic analysis that leverages factor analysis to capture how linguistic features co-occur to fulfill different linguistic functions across registers. In the original MD analysis, Biber (1988) identified five major dimensions of language where features tend to co-occur. These dimensions have been frequently used as a point of comparison for understanding registers that weren’t included in the original analysis – a method known as “additive MD analysis” (Berber Sardinha, 2024). The present study is an additive MD analysis that aims to shed light on register variation in annual reports (ARs) from multinational automotive corporations (e.g., General Motors, Toyota). Specifically, two research questions are addressed: (1) How do English-language ARs from multinational automotive corporations compare to Biber’s original dimensions of English; and (2), do the linguistic profiles of ARs as measured by additive MD analysis vary by corporation, and if so, how?
To address these questions, a corpus was compiled from publicly available, self-described ARs from seven multinational automotive corporations from 2005 through 2019, resulting in a corpus of about 8.2 million words. The corpus was tagged with the Biber Tagger and post-processed with Tag Count to generate dimension scores for each report. Next, a linear mixed effects model was carried out with “corporation” as a random effect to measure and quantify variation between ARs. Findings include that ARs are most similar to the registers “official documents” and “academic prose” on Dimension 1, but featured language that is more informationally dense; on Dimension 5 (abstract vs. not abstract) ARs are most similar to research articles, but with markedly less abstract language. Furthermore, statistically significant differences with medium to large effect sizes were found between corporations on three dimensions. These findings point to sub-register variation and differing communicative goals in ARs between automotive corporations.
References
Berber Sardinha, T. (2024). AI-generated vs human-authored texts: A multidimensional
comparison. Applied Corpus Linguistics, 4, 1-9.
Biber, D. (1988). Variation Across Speech and Writing. Cambridge University Press,
Cambridge.
Multidimensional analysis (MD analysis) is a method of linguistic analysis that leverages factor analysis to capture how linguistic features co-occur to fulfill different linguistic functions across registers. In the original MD analysis, Biber (1988) identified five major dimensions of language where features tend to co-occur. These dimensions have been frequently used as a point of comparison for understanding registers that weren’t included in the original analysis – a method known as “additive MD analysis” (Berber Sardinha, 2024). The present study is an additive MD analysis that aims to shed light on register variation in annual reports (ARs) from multinational automotive corporations (e.g., General Motors, Toyota). Specifically, two research questions are addressed: (1) How do English-language ARs from multinational automotive corporations compare to Biber’s original dimensions of English; and (2), do the linguistic profiles of ARs as measured by additive MD analysis vary by corporation, and if so, how?
To address these questions, a corpus was compiled from publicly available, self-described ARs from seven multinational automotive corporations from 2005 through 2019, resulting in a corpus of about 8.2 million words. The corpus was tagged with the Biber Tagger and post-processed with Tag Count to generate dimension scores for each report. Next, a linear mixed effects model was carried out with “corporation” as a random effect to measure and quantify variation between ARs. Findings include that ARs are most similar to the registers “official documents” and “academic prose” on Dimension 1, but featured language that is more informationally dense; on Dimension 5 (abstract vs. not abstract) ARs are most similar to research articles, but with markedly less abstract language. Furthermore, statistically significant differences with medium to large effect sizes were found between corporations on three dimensions. These findings point to sub-register variation and differing communicative goals in ARs between automotive corporations.
References
Berber Sardinha, T. (2024). AI-generated vs human-authored texts: A multidimensional
comparison. Applied Corpus Linguistics, 4, 1-9.
Biber, D. (1988). Variation Across Speech and Writing. Cambridge University Press,
Cambridge.
Corpus-based multimodal analysis depends on the conversion of non-verbal language into machine-readable textual format. In the case of images, images must first be textualized through annotation procedures (Christiansen et al., 2020; Collins & Baker, 2023). Automated image tag- ging has become the dominant solution in multimodal corpus analysis of visual content (Baker et al., 2025; Christiansen et al., 2020) More recently, multimodal large language models have made it possible to generate running-text descriptions of images, raising the question of whether LLM-generated descriptions provide a different representation of visual discourse (Author, 2025). To answer this question, this study compared image tagging and text descriptions as annotation procedures for image analysis. A corpus of 500 immigration-related news images was anno- tated through Google Cloud Vision AI and gpt-image. Each annotated corpus was then analyzed through Lexical Multi-Dimensional Analysis (LMDA) (Berber Sardinha & Fitzsimmons-Doolan, 2025) in order to identify the underlying visual discourses. The analyses produced different dimen- sional models. The image-tagging corpus generated five dimensions, centered on visually concrete domains such as military personnel and border security, politicians and media appearances, and police arrests and detention. The text-description corpus produced four dimensions comprising six poles, including discourses of mediatized immigration politics, fortress-border nationalism, ideological contestation, technocratic management, and immigration control as physical domina- tion. Few similar dimensions emerged, such as an image-tagging dimension centered on policing and law enforcement with a text-description dimension focused on physical restraint and police violence. Overall, the findings showed that image tagging and LLM-generated descriptions cap- ture related but different aspects of visual discourse. Image tagging privileges visible entities, whereas LLM-generated descriptions capture higher-level characteristics, including social activity and abstract processes. The findings suggest that LLM-generated descriptions constitute a vi- able new form of multimodal corpus annotation and expand the methodological possibilities for corpus-based analysis of images.
References
Author. (2025).
Baker, P., Schmück, H., & Qian, Y. (2025). Automatic image tagging for corpus linguistics: A
multimodal study of news representations of Islam. Cambridge University Press.
Berber Sardinha, T., & Fitzsimmons-Doolan, S. (2025). Lexical multidimensional analysis: Iden- tifying discourses and ideologies. Cambridge University Press. https://doi.org/https:
//doi.org/10.1017/9781009335683
Christiansen, A., Dance, W., & Wild, A. (2020). Constructing corpora from images and text: An introduction to Visual Constituent Analysis. In S. Ruediger & D. Dayter (Eds.), Corpus approaches to social media (pp. 149–174). John Benjamins.
Collins, L., & Baker, P. (2023). Creating and analysing a multimodal corpus of news texts with Google Cloud Vision’s automatic image tagger. Applied Corpus Linguistics, 3(1), 100043.
Corpus-based multimodal analysis depends on the conversion of non-verbal language into machine-readable textual format. In the case of images, images must first be textualized through annotation procedures (Christiansen et al., 2020; Collins & Baker, 2023). Automated image tag- ging has become the dominant solution in multimodal corpus analysis of visual content (Baker et al., 2025; Christiansen et al., 2020) More recently, multimodal large language models have made it possible to generate running-text descriptions of images, raising the question of whether LLM-generated descriptions provide a different representation of visual discourse (Author, 2025). To answer this question, this study compared image tagging and text descriptions as annotation procedures for image analysis. A corpus of 500 immigration-related news images was anno- tated through Google Cloud Vision AI and gpt-image. Each annotated corpus was then analyzed through Lexical Multi-Dimensional Analysis (LMDA) (Berber Sardinha & Fitzsimmons-Doolan, 2025) in order to identify the underlying visual discourses. The analyses produced different dimen- sional models. The image-tagging corpus generated five dimensions, centered on visually concrete domains such as military personnel and border security, politicians and media appearances, and police arrests and detention. The text-description corpus produced four dimensions comprising six poles, including discourses of mediatized immigration politics, fortress-border nationalism, ideological contestation, technocratic management, and immigration control as physical domina- tion. Few similar dimensions emerged, such as an image-tagging dimension centered on policing and law enforcement with a text-description dimension focused on physical restraint and police violence. Overall, the findings showed that image tagging and LLM-generated descriptions cap- ture related but different aspects of visual discourse. Image tagging privileges visible entities, whereas LLM-generated descriptions capture higher-level characteristics, including social activity and abstract processes. The findings suggest that LLM-generated descriptions constitute a vi- able new form of multimodal corpus annotation and expand the methodological possibilities for corpus-based analysis of images.
References
Author. (2025).
Baker, P., Schmück, H., & Qian, Y. (2025). Automatic image tagging for corpus linguistics: A
multimodal study of news representations of Islam. Cambridge University Press.
Berber Sardinha, T., & Fitzsimmons-Doolan, S. (2025). Lexical multidimensional analysis: Iden- tifying discourses and ideologies. Cambridge University Press. https://doi.org/https:
//doi.org/10.1017/9781009335683
Christiansen, A., Dance, W., & Wild, A. (2020). Constructing corpora from images and text: An introduction to Visual Constituent Analysis. In S. Ruediger & D. Dayter (Eds.), Corpus approaches to social media (pp. 149–174). John Benjamins.
Collins, L., & Baker, P. (2023). Creating and analysing a multimodal corpus of news texts with Google Cloud Vision’s automatic image tagger. Applied Corpus Linguistics, 3(1), 100043.
The International Corpus of Learner English version 3 (ICLEv3) contains +9,500 argumentative essays written by university-level learners of English from 25 L1 backgrounds who are consid- ered advanced level. However, advanced level is not equivalent across international educational systems. An assessment of 500 texts (Granger et al., 2020, p. 12-13) indicated sharp differ- ences across subcorpora: e.g. 60% of the Macedonian sample was rated C2, whereas 95% of the Chinese sample was rated B2 or below. This paper presents an AI-assisted procedure for estimating ICLE v3 learner proficiency. Initially, we used prompt engineering and then fine-tuned GPT-4.1-mini, which was the most capable OpenAI model available for fine-tuning at the time, to assign CEFR labels to essays. However, CEFR labeling produced weak agreement with human assessment, with exact accuracy of .342, adjacent accuracy of .840, Cohen’s κ = .082, and quadratic weighted κ = .362, consistent with concerns about lack of transparency and coherence in CEFR-based language testing (Weir, 2005). To overcome this limitation, we adopted a com- parative judgment approach (Thwaites, Kollias, & Paquot, 2024), using GPT-5.4 to rank each essay against 20 benchmark essays from the ICLE500 dataset (Thwaites, Kollias, Kanistra, & Paquot, 2024), whose ordering was based on human many-facet Rasch rankings; the resulting AI rankings were then validated against the corresponding human rankings. Spearman’s ρ = .778 and Kendall’s τ = .597 indicated strong rank-order correspondence between the AI rankings and human rankings. The procedure was then applied to the full ICLEv3 corpus. The study argues that comparative AI-assisted ranking provides an alternative to CEFR labeling for large learner corpora, especially when the goal is not to assign official proficiency levels but to control for proficiency variation in corpus-based analyses.
References
Granger, S., Dupont, M., Meunier, F., Naets, H., & Paquot, M. (2020). The International Corpus of Learner English: Version 3. Presses universitaires de Louvain.
Thwaites, P., Kollias, C., Kanistra, V., & Paquot, M. (2024). ICLE500. Open Data UCLouvain, V1. Louvain-la-Neuve. https://doi.org/10.14428/DVN/RIOSSC
Thwaites, P., Kollias, C., & Paquot, M. (2024). Is CJ a valid, reliable form of L2 writing assessment when texts are long, homogeneous in proficiency, and feature heterogeneous prompts? Assessing Writing, 60, 100843. https://doi.org/10.1016/j.asw.2024.100843
Weir, C. J. (2005). Limitations of the Common European Framework for developing comparable examinations and tests. Language Testing, 22(3), 281–300.
The International Corpus of Learner English version 3 (ICLEv3) contains +9,500 argumentative essays written by university-level learners of English from 25 L1 backgrounds who are consid- ered advanced level. However, advanced level is not equivalent across international educational systems. An assessment of 500 texts (Granger et al., 2020, p. 12-13) indicated sharp differ- ences across subcorpora: e.g. 60% of the Macedonian sample was rated C2, whereas 95% of the Chinese sample was rated B2 or below. This paper presents an AI-assisted procedure for estimating ICLE v3 learner proficiency. Initially, we used prompt engineering and then fine-tuned GPT-4.1-mini, which was the most capable OpenAI model available for fine-tuning at the time, to assign CEFR labels to essays. However, CEFR labeling produced weak agreement with human assessment, with exact accuracy of .342, adjacent accuracy of .840, Cohen’s κ = .082, and quadratic weighted κ = .362, consistent with concerns about lack of transparency and coherence in CEFR-based language testing (Weir, 2005). To overcome this limitation, we adopted a com- parative judgment approach (Thwaites, Kollias, & Paquot, 2024), using GPT-5.4 to rank each essay against 20 benchmark essays from the ICLE500 dataset (Thwaites, Kollias, Kanistra, & Paquot, 2024), whose ordering was based on human many-facet Rasch rankings; the resulting AI rankings were then validated against the corresponding human rankings. Spearman’s ρ = .778 and Kendall’s τ = .597 indicated strong rank-order correspondence between the AI rankings and human rankings. The procedure was then applied to the full ICLEv3 corpus. The study argues that comparative AI-assisted ranking provides an alternative to CEFR labeling for large learner corpora, especially when the goal is not to assign official proficiency levels but to control for proficiency variation in corpus-based analyses.
References
Granger, S., Dupont, M., Meunier, F., Naets, H., & Paquot, M. (2020). The International Corpus of Learner English: Version 3. Presses universitaires de Louvain.
Thwaites, P., Kollias, C., Kanistra, V., & Paquot, M. (2024). ICLE500. Open Data UCLouvain, V1. Louvain-la-Neuve. https://doi.org/10.14428/DVN/RIOSSC
Thwaites, P., Kollias, C., & Paquot, M. (2024). Is CJ a valid, reliable form of L2 writing assessment when texts are long, homogeneous in proficiency, and feature heterogeneous prompts? Assessing Writing, 60, 100843. https://doi.org/10.1016/j.asw.2024.100843
Weir, C. J. (2005). Limitations of the Common European Framework for developing comparable examinations and tests. Language Testing, 22(3), 281–300.
Textbook corpus studies in English language education have typically focused on written lexical units, such as word frequency, vocabulary coverage, and lexical progression. Less attention has been paid to the phonological characteristics of textbook input, although such information is relevant to learners’ experience with spoken forms. This study examines syllable sequences and phonological complexity in Japanese secondary school English textbooks through dictionary-based phonological annotation. Textbook texts were tokenized and converted into ARPAbet phonological representations using the CMU Pronouncing Dictionary. The analysis focused on syllable n-grams, word length in syllables and phonemes, CV structures, and consonant clusters. In addition, word n-gram and syllable n-gram patterns were compared, and source word sequence diversity was calculated to examine whether frequent syllable n-grams simply reflected recurrent lexical sequences. A supplementary analysis excluding function words was also conducted. The results showed that frequent syllable n-grams were strongly influenced by high-frequency words and formulaic word sequences, indicating that syllable n-gram frequency should not be interpreted as independent of lexical repetition. However, dictionary-based phonological annotation provided additional information that word n-grams alone cannot capture. In particular, content words showed greater phonological complexity than the full word set: 52.0% of content-word tokens contained consonant clusters, compared with 29.5% when all words were included. Content words also more often ended in consonants and had longer phonological forms. These findings suggest that textbook input is shaped by both recurrent lexical sequences and phonological complexity within content words. The study demonstrates how corpus analysis of textbooks can be extended beyond written word frequency by incorporating dictionary-based phonological information.
Textbook corpus studies in English language education have typically focused on written lexical units, such as word frequency, vocabulary coverage, and lexical progression. Less attention has been paid to the phonological characteristics of textbook input, although such information is relevant to learners’ experience with spoken forms. This study examines syllable sequences and phonological complexity in Japanese secondary school English textbooks through dictionary-based phonological annotation. Textbook texts were tokenized and converted into ARPAbet phonological representations using the CMU Pronouncing Dictionary. The analysis focused on syllable n-grams, word length in syllables and phonemes, CV structures, and consonant clusters. In addition, word n-gram and syllable n-gram patterns were compared, and source word sequence diversity was calculated to examine whether frequent syllable n-grams simply reflected recurrent lexical sequences. A supplementary analysis excluding function words was also conducted. The results showed that frequent syllable n-grams were strongly influenced by high-frequency words and formulaic word sequences, indicating that syllable n-gram frequency should not be interpreted as independent of lexical repetition. However, dictionary-based phonological annotation provided additional information that word n-grams alone cannot capture. In particular, content words showed greater phonological complexity than the full word set: 52.0% of content-word tokens contained consonant clusters, compared with 29.5% when all words were included. Content words also more often ended in consonants and had longer phonological forms. These findings suggest that textbook input is shaped by both recurrent lexical sequences and phonological complexity within content words. The study demonstrates how corpus analysis of textbooks can be extended beyond written word frequency by incorporating dictionary-based phonological information.
With the growing use of generative AI in academic writing, AI-generated academic texts have become an important object of corpus-based linguistic research. This study examines the lexicogrammatical differences between AI-generated and human-written academic abstracts in terms of syntactic complexity and lexical richness. The analysis is based on a corpus of 1,000 abstracts from four disciplines: linguistics, sociology, law, and history. This corpus includes human-written abstracts and texts generated by four large language models: DeepSeek, Doubao, Gemini, and ChatGPT. Using Lu’s L2 Syntactic Complexity Analyzer (L2SCA) and L2 Lexical Complexity Analyzer (L2LCA), the study compares syntactic and lexical features across text groups, models, and disciplines. The results show that AI-generated abstracts tend to use more coordinate structures and display higher values on several lexical variety measures, whereas human-written abstracts show greater variation in verb use. Cross-disciplinary analysis further indicates that the overall pattern of human–AI differences remains relatively stable across disciplines, with only limited variation in measures related to vocabulary range. These findings provide corpus-based linguistic evidence for understanding how AI-generated academic abstracts differ from human-written ones and offer implications for the responsible use of AI in academic writing.
With the growing use of generative AI in academic writing, AI-generated academic texts have become an important object of corpus-based linguistic research. This study examines the lexicogrammatical differences between AI-generated and human-written academic abstracts in terms of syntactic complexity and lexical richness. The analysis is based on a corpus of 1,000 abstracts from four disciplines: linguistics, sociology, law, and history. This corpus includes human-written abstracts and texts generated by four large language models: DeepSeek, Doubao, Gemini, and ChatGPT. Using Lu’s L2 Syntactic Complexity Analyzer (L2SCA) and L2 Lexical Complexity Analyzer (L2LCA), the study compares syntactic and lexical features across text groups, models, and disciplines. The results show that AI-generated abstracts tend to use more coordinate structures and display higher values on several lexical variety measures, whereas human-written abstracts show greater variation in verb use. Cross-disciplinary analysis further indicates that the overall pattern of human–AI differences remains relatively stable across disciplines, with only limited variation in measures related to vocabulary range. These findings provide corpus-based linguistic evidence for understanding how AI-generated academic abstracts differ from human-written ones and offer implications for the responsible use of AI in academic writing.
Traditional corpus-based translation studies, inspired by the utopian search for translation universals, have heavily relied on standard corpus techniques, such as keyword analysis and lexical bundles to identify textual variation and to understand how translations differ from non-translations. While valuable, these isolated metrics often fail to capture the complex, co-occurring linguistic features that define institutional registers. This study demonstrates how to advance translation research beyond single-feature analyses by applying a comprehensive, full Multidimensional Analysis (MDA) framework (Biber 1988, 2016; Brezina 2018; Berber Sardinha & Veirano Pinto 2019; Goulart & Wood 2021) to capture macroscopic textual variation. While MDA has been successfully applied to study a wide range of registers (e.g. Zhao 2024), its potential for examining translated language remains largely underexplored (cf. Calzada Pérez and Sánchez Ramos 2021 for an overview).
This study fills in that gap by analyzing the Polish Eurolect, a hybrid variety shaped by multilingual law-making, translation and institutional constraints within the European Union, and comparing it to national varieties. Using a corpus of key institutional registers (legal acts, judgments, administrative reports, and citizen-facing websites), we identify four statistically-derived dimensions of variation: Argumentative vs Informational, Engaged Instruction vs Distanced Authority, Prescriptive vs Narrative, and Lexical Richness.
The findings map variation and group institutional registers, revealing significant differences between how supranational and national institutions communicate (Biel, Wasilewska, and Koźbiał 2026). EU legal acts and judgments exhibit higher degrees of prescriptiveness, legal referencing, and argumentative structuring than their domestic counterparts. Conversely, EU citizen-facing websites utilize fewer engagement and explanatory strategies, while EU reports adopt a less distanced style overall. The MDA maps variation and groups institutional registers, successfully visualizing the hybridity induced by translation. Ultimately, moving from micro-level features to holistic MDA mapping captures translation-induced hybridity across registers and provides a diagnostics framework to optimize the clarity and naturalness of translated public discourse.
Berber Sardinha, T., & Veirano Pinto, M. (2014b). Multidimensional analysis, 25 years on. John Benjamins. https://doi.org/10.1075/scl.60
Biber, D. (1988). Variation across speech and writing. Cambridge University Press. https://doi.org/10.1017/CBO9780511621024
Biber, D. (2016). Using multidimensional analysis to explore cross-linguistic universals of register variation. In L. Marie-Aude & V. Svetlana (Eds.), Genre- and register-related discourse features in contrast (pp. 7–34). John Benjamins. https://doi.org/10.1075/bct.87.02bib
Biel, Łucja, Katarzyna Wasilewska, and Dariusz Koźbiał. 2026. "Dimensions of variation across institutional legal and administrative registers." International Journal of Corpus Linguistics 31 (1):64-93. doi: https://doi.org/10.1075/ijcl.25126.bie
Brezina, V. (2018). Statistics in corpus linguistics: A practical guide. Cambridge University Press. https://doi.org/10.1017/9781316410899
Calzada Pérez, M., & Sánchez Ramos, M. d. M. (2021). MDA analysis of translated and non-translated parliamentary discourse. In M. Ji & M. P. Oakes (Eds.), Corpus exploration of lexis and discourse in translation (pp. 26–55). Routledge.
Goulart, L., & Wood, M. (2021). Methodological synthesis of research using multidimensional analysis. Journal of Research Design and Statistics in Linguistics and Communication Science, 6(2), 107–137. https://doi.org/10.1558/jrds.18454
Zhao, W. (2024). A corpus-based multidimensional analysis of the linguistic features of Aviation English. English for Specific Purposes, 76, 57–73. https://doi.org/https://doi.org/10.1016/j.esp.2024.05.004
Traditional corpus-based translation studies, inspired by the utopian search for translation universals, have heavily relied on standard corpus techniques, such as keyword analysis and lexical bundles to identify textual variation and to understand how translations differ from non-translations. While valuable, these isolated metrics often fail to capture the complex, co-occurring linguistic features that define institutional registers. This study demonstrates how to advance translation research beyond single-feature analyses by applying a comprehensive, full Multidimensional Analysis (MDA) framework (Biber 1988, 2016; Brezina 2018; Berber Sardinha & Veirano Pinto 2019; Goulart & Wood 2021) to capture macroscopic textual variation. While MDA has been successfully applied to study a wide range of registers (e.g. Zhao 2024), its potential for examining translated language remains largely underexplored (cf. Calzada Pérez and Sánchez Ramos 2021 for an overview).
This study fills in that gap by analyzing the Polish Eurolect, a hybrid variety shaped by multilingual law-making, translation and institutional constraints within the European Union, and comparing it to national varieties. Using a corpus of key institutional registers (legal acts, judgments, administrative reports, and citizen-facing websites), we identify four statistically-derived dimensions of variation: Argumentative vs Informational, Engaged Instruction vs Distanced Authority, Prescriptive vs Narrative, and Lexical Richness.
The findings map variation and group institutional registers, revealing significant differences between how supranational and national institutions communicate (Biel, Wasilewska, and Koźbiał 2026). EU legal acts and judgments exhibit higher degrees of prescriptiveness, legal referencing, and argumentative structuring than their domestic counterparts. Conversely, EU citizen-facing websites utilize fewer engagement and explanatory strategies, while EU reports adopt a less distanced style overall. The MDA maps variation and groups institutional registers, successfully visualizing the hybridity induced by translation. Ultimately, moving from micro-level features to holistic MDA mapping captures translation-induced hybridity across registers and provides a diagnostics framework to optimize the clarity and naturalness of translated public discourse.
Berber Sardinha, T., & Veirano Pinto, M. (2014b). Multidimensional analysis, 25 years on. John Benjamins. https://doi.org/10.1075/scl.60
Biber, D. (1988). Variation across speech and writing. Cambridge University Press. https://doi.org/10.1017/CBO9780511621024
Biber, D. (2016). Using multidimensional analysis to explore cross-linguistic universals of register variation. In L. Marie-Aude & V. Svetlana (Eds.), Genre- and register-related discourse features in contrast (pp. 7–34). John Benjamins. https://doi.org/10.1075/bct.87.02bib
Biel, Łucja, Katarzyna Wasilewska, and Dariusz Koźbiał. 2026. "Dimensions of variation across institutional legal and administrative registers." International Journal of Corpus Linguistics 31 (1):64-93. doi: https://doi.org/10.1075/ijcl.25126.bie
Brezina, V. (2018). Statistics in corpus linguistics: A practical guide. Cambridge University Press. https://doi.org/10.1017/9781316410899
Calzada Pérez, M., & Sánchez Ramos, M. d. M. (2021). MDA analysis of translated and non-translated parliamentary discourse. In M. Ji & M. P. Oakes (Eds.), Corpus exploration of lexis and discourse in translation (pp. 26–55). Routledge.
Goulart, L., & Wood, M. (2021). Methodological synthesis of research using multidimensional analysis. Journal of Research Design and Statistics in Linguistics and Communication Science, 6(2), 107–137. https://doi.org/10.1558/jrds.18454
Zhao, W. (2024). A corpus-based multidimensional analysis of the linguistic features of Aviation English. English for Specific Purposes, 76, 57–73. https://doi.org/https://doi.org/10.1016/j.esp.2024.05.004
The use of adverbs to amplify or attenuate the intensity of an utterance is employed in both to add emphasis or nuance to the meaning of one’s words. Such adverbs of degree can modify various parts of speech, but this study focused specifically on the use of degree adverbs to modify adjectives (Beltrama & Bochnak, 2015; Zhiber & Korotina, 2019). Although many studies have explored the use of intensifying and attenuating adverbs in L1 speech and writing (Lorenz, 1998; Indhiarti & Chaerunnisa, 2020), far fewer have examined how L2 users of English approach adjective modification, and fewer still have looked at degree adverbs in L2 spoken production (Hasselgård, 2022; Makhatadze, 2023). Thus, this study examined the use of degree adverb + adjective pairings across several CEFR proficiency levels in L2 spoken language by examining the Trinity Lancaster Corpus (Gablasova et al., 2019) and comparing these with the L1 data drawn from the Spoken BNC2014 corpus (Love et al., 2017; Brezina & Fox, 2021). The results suggest that while less proficient learners use fewer adverbs and adjectives than more proficient and L1 speakers, they tend to overuse the degree adverb + adjective pairing as well as specific adverbs such as very. Additionally, although the most frequent adjectives are similar between L1 and L2 speakers, L2 speakers draw on a narrower selection of degree adverbs to modify these adjectives and almost never use evaluative adverbs such as horribly or splendidly compared to L1 speakers. While limitations of this study should be taken into consideration, such as the differences in the nature of the spoken interactions and interpersonal relationships of the speakers, the findings of this study have pedagogical implications in the EFL context (Makhatadze, 2023) and support the claim that learners of English need more explicit instruction in the use of degree adverbs (Pérez-Paredes & Díez-Bedmar, 2019).
The use of adverbs to amplify or attenuate the intensity of an utterance is employed in both to add emphasis or nuance to the meaning of one’s words. Such adverbs of degree can modify various parts of speech, but this study focused specifically on the use of degree adverbs to modify adjectives (Beltrama & Bochnak, 2015; Zhiber & Korotina, 2019). Although many studies have explored the use of intensifying and attenuating adverbs in L1 speech and writing (Lorenz, 1998; Indhiarti & Chaerunnisa, 2020), far fewer have examined how L2 users of English approach adjective modification, and fewer still have looked at degree adverbs in L2 spoken production (Hasselgård, 2022; Makhatadze, 2023). Thus, this study examined the use of degree adverb + adjective pairings across several CEFR proficiency levels in L2 spoken language by examining the Trinity Lancaster Corpus (Gablasova et al., 2019) and comparing these with the L1 data drawn from the Spoken BNC2014 corpus (Love et al., 2017; Brezina & Fox, 2021). The results suggest that while less proficient learners use fewer adverbs and adjectives than more proficient and L1 speakers, they tend to overuse the degree adverb + adjective pairing as well as specific adverbs such as very. Additionally, although the most frequent adjectives are similar between L1 and L2 speakers, L2 speakers draw on a narrower selection of degree adverbs to modify these adjectives and almost never use evaluative adverbs such as horribly or splendidly compared to L1 speakers. While limitations of this study should be taken into consideration, such as the differences in the nature of the spoken interactions and interpersonal relationships of the speakers, the findings of this study have pedagogical implications in the EFL context (Makhatadze, 2023) and support the claim that learners of English need more explicit instruction in the use of degree adverbs (Pérez-Paredes & Díez-Bedmar, 2019).
This study examines how National Geographic Magazine (NG) has linguistically represented Korea and Japan from the 1890s to the 2020s through a diachronic Corpus-Assisted Discourse Studies (CADS) approach.
Prior research has addressed NG's construction of a visual discourse othering the non-Western world (Lutz & Collins, 1993; Steet, 2000; Kim, 2005; Song, 2006; Kim, 2015). Kim (2005) traced diachronic shifts in NG's representation of Korea through visual materials, while Song (2006) extended this to a Korea-Japan comparative framework, proposing that NG's gaze toward the two nations developed asymmetrically. However, both rely on qualitative analysis of visual materials or content, leaving unexamined how such asymmetries are realized through lexical choice. As the most basic unit through which discourse categorizes the world, lexical items allow corpus-based analysis to capture recurring or shifting discursive tendencies over time—a strength qualitative analysis alone cannot offer. This study examines through what lexical choices and keyword patterns NG's differential representations of Korea and Japan are discursively enacted.
Two corpora were compiled: a Korea-related corpus (28 articles, ca. 95,700 tokens, 1890–2020) and a Japan-related corpus (86 articles, ca. 384,000 tokens, 1894–2020), each divided into sub-corpora by historical period. Though modest in size, these constitute exhaustive collections of NG's Korea- and Japan-related articles. Given the size disparity, log-likelihood-based keyword analysis was applied, with noun and adjective keywords indicating representational and evaluative dimensions of discourse. Quantitative results were interpreted through qualitative analysis of concordance lines and visuals.
The findings broadly align with Kim (2005) and Song (2006), but the combined quantitative-qualitative analysis newly captures how shifts in NG's gaze toward Korea and Japan—and the asymmetry between them—are concretely realized linguistically. This study transforms insights grounded in qualitative interpretation and expert intuition into verifiable linguistic evidence, and extends the analysis to more recent periods.
This study examines how National Geographic Magazine (NG) has linguistically represented Korea and Japan from the 1890s to the 2020s through a diachronic Corpus-Assisted Discourse Studies (CADS) approach.
Prior research has addressed NG's construction of a visual discourse othering the non-Western world (Lutz & Collins, 1993; Steet, 2000; Kim, 2005; Song, 2006; Kim, 2015). Kim (2005) traced diachronic shifts in NG's representation of Korea through visual materials, while Song (2006) extended this to a Korea-Japan comparative framework, proposing that NG's gaze toward the two nations developed asymmetrically. However, both rely on qualitative analysis of visual materials or content, leaving unexamined how such asymmetries are realized through lexical choice. As the most basic unit through which discourse categorizes the world, lexical items allow corpus-based analysis to capture recurring or shifting discursive tendencies over time—a strength qualitative analysis alone cannot offer. This study examines through what lexical choices and keyword patterns NG's differential representations of Korea and Japan are discursively enacted.
Two corpora were compiled: a Korea-related corpus (28 articles, ca. 95,700 tokens, 1890–2020) and a Japan-related corpus (86 articles, ca. 384,000 tokens, 1894–2020), each divided into sub-corpora by historical period. Though modest in size, these constitute exhaustive collections of NG's Korea- and Japan-related articles. Given the size disparity, log-likelihood-based keyword analysis was applied, with noun and adjective keywords indicating representational and evaluative dimensions of discourse. Quantitative results were interpreted through qualitative analysis of concordance lines and visuals.
The findings broadly align with Kim (2005) and Song (2006), but the combined quantitative-qualitative analysis newly captures how shifts in NG's gaze toward Korea and Japan—and the asymmetry between them—are concretely realized linguistically. This study transforms insights grounded in qualitative interpretation and expert intuition into verifiable linguistic evidence, and extends the analysis to more recent periods.
Abstract
Autism Spectrum Disorder is a neurodevelopmental condition whose prevalence continues to rise in Indonesia. News media play a significant role in shaping public perceptions of individuals with autism through their linguistic and discursive choices (Hungerford et al., 2025). However, studies on autism representation in Indonesian news media remain scarce. This study investigates how autism is portrayed in news coverage from four major Indonesian online outlets Kompas.com, Detik.com, Tribunnews.com, and CNN Indonesia between 2016 and 2025 by employing a corpus-assisted discourse studies (CADS) framework.
News articles containing the search terms autisme (autism) and autis (autistic/autist) were collected using BootCaT (Baroni & Bernardini, 2004) and Python scripts. The primary analytical techniques focus on collocation and concordance analyses to identify lexical associations, semantic tendencies, and discursive constructions surrounding autism. Quantitative findings are subsequently interpreted qualitatively using van Leeuwen’s social actor framework (2008) and disability news representation models (Clogston, 1994; Haller, 2010).
Preliminary collocational analyses of autism and autis reveal that both search terms are predominantly child-centred, with anak (child/children) ranking first in both collocation lists. Autisme is strongly associated with medical and diagnostic discourse, reflecting the medical model. The shared high-frequency collocates penderita (sufferer) further suggests passive, patient-like constructions of individuals with autism. Meanwhile, autis shows stronger associations with caregiving and social contexts with collocates like terapi (therapy), peduli (care), orangtua (parents), sekolah (school), indicating the social pathology model, where individuals with autism are positioned as dependent on familial and state support. Representations aligned with progressive models, such as the minority/civil rights models, appear largely absent. These preliminary results indicate a need for Indonesian media professionals to adopt a more balanced portrayal of individuals with autism, specifically by prioritising narratives that highlight their agency. Shifting media narratives toward progressive disability models has the potential to reduce stigma and foster broader societal acceptance of individuals with autism.
Keywords: autism; corpus-assisted discourse studies; disability news models; Indonesian media; inclusive society
References
Baroni, M., & Bernardini, S. (2004). BootCaT: Bootstrapping corpora and terms from the web. In M. T. Lino, M. F. Xavier, F. Ferreira, R. Costa, & R. Silva (Eds.), Proceedings of LREC 2004 (pp. 1313-1316). ELRA
Clogston, J. S. (1994). Disability coverage in American newspapers. In J. A. Nelson (Ed.), The disabled, the media, and the information age (pp. 45–57). Greenwood Publishing Group.
Haller, B. A. (2010). Representing disability in an ableist world: Essays on mass media. Advocado Press.
Hungerford, C., Kornhaber, R., West, S., & Cleary, M. (2025). Autism, Stereotypes, and Stigma: The Impact of Media Representations. Issues in Mental Health Nursing, 46(3), 254–260. https://doi.org/10.1080/01612840.2025.2456698
Karaminis, T., Gabrielatos, C., Maden-Weinberger, U., & Beattie, G. (2023). Portrayals of autism in the British press: A corpus-based study. Autism, 27(4), 1092-1114. https://doi.org/10.1177/13623613221131752
van Leeuwen, T. (2008). Discourse and practice: New tools for critical discourse analysis. Oxford University Press.
Abstract
Autism Spectrum Disorder is a neurodevelopmental condition whose prevalence continues to rise in Indonesia. News media play a significant role in shaping public perceptions of individuals with autism through their linguistic and discursive choices (Hungerford et al., 2025). However, studies on autism representation in Indonesian news media remain scarce. This study investigates how autism is portrayed in news coverage from four major Indonesian online outlets Kompas.com, Detik.com, Tribunnews.com, and CNN Indonesia between 2016 and 2025 by employing a corpus-assisted discourse studies (CADS) framework.
News articles containing the search terms autisme (autism) and autis (autistic/autist) were collected using BootCaT (Baroni & Bernardini, 2004) and Python scripts. The primary analytical techniques focus on collocation and concordance analyses to identify lexical associations, semantic tendencies, and discursive constructions surrounding autism. Quantitative findings are subsequently interpreted qualitatively using van Leeuwen’s social actor framework (2008) and disability news representation models (Clogston, 1994; Haller, 2010).
Preliminary collocational analyses of autism and autis reveal that both search terms are predominantly child-centred, with anak (child/children) ranking first in both collocation lists. Autisme is strongly associated with medical and diagnostic discourse, reflecting the medical model. The shared high-frequency collocates penderita (sufferer) further suggests passive, patient-like constructions of individuals with autism. Meanwhile, autis shows stronger associations with caregiving and social contexts with collocates like terapi (therapy), peduli (care), orangtua (parents), sekolah (school), indicating the social pathology model, where individuals with autism are positioned as dependent on familial and state support. Representations aligned with progressive models, such as the minority/civil rights models, appear largely absent. These preliminary results indicate a need for Indonesian media professionals to adopt a more balanced portrayal of individuals with autism, specifically by prioritising narratives that highlight their agency. Shifting media narratives toward progressive disability models has the potential to reduce stigma and foster broader societal acceptance of individuals with autism.
Keywords: autism; corpus-assisted discourse studies; disability news models; Indonesian media; inclusive society
References
Baroni, M., & Bernardini, S. (2004). BootCaT: Bootstrapping corpora and terms from the web. In M. T. Lino, M. F. Xavier, F. Ferreira, R. Costa, & R. Silva (Eds.), Proceedings of LREC 2004 (pp. 1313-1316). ELRA
Clogston, J. S. (1994). Disability coverage in American newspapers. In J. A. Nelson (Ed.), The disabled, the media, and the information age (pp. 45–57). Greenwood Publishing Group.
Haller, B. A. (2010). Representing disability in an ableist world: Essays on mass media. Advocado Press.
Hungerford, C., Kornhaber, R., West, S., & Cleary, M. (2025). Autism, Stereotypes, and Stigma: The Impact of Media Representations. Issues in Mental Health Nursing, 46(3), 254–260. https://doi.org/10.1080/01612840.2025.2456698
Karaminis, T., Gabrielatos, C., Maden-Weinberger, U., & Beattie, G. (2023). Portrayals of autism in the British press: A corpus-based study. Autism, 27(4), 1092-1114. https://doi.org/10.1177/13623613221131752
van Leeuwen, T. (2008). Discourse and practice: New tools for critical discourse analysis. Oxford University Press.
Abstract
As a carbon-intensive industry, aviation is facing mounting pressure for sustainable development due to its adverse environmental and social problems, such as noise, emissions, pollution, and waste (ICAO, 2025). Along with sustainability policies and actions in practice, airlines are also leveraging corporate social responsibility (CSR) reporting to establish legitimacy, gain trust, and promote favorable images (Bhatia, 2012). While existing literature has examined CSR reporting in industries such as energy, finance, and retail, the aviation sector remains largely underexplored. Methodologically, prior research predominantly adopts a conventional corpus-assisted discourse studies (CADS) approach (Baker et al., 2008; Partington, 2010). Nevertheless, as demonstrated by recent scholarship (Brookes & McEnery, 2019; Jaworska & Nanda, 2018), topic modeling is a advent tool that can effectively facilitate CADS with the identification of thematic categories of large-scale datasets.
This study thus integrates topic modeling with corpus-assisted discourse study to examine the discursive features and diachronic shifts of various foci in the corporate social responsibility (CSR) reports of world-leading airlines from 2019 to 2024. The corpus comprises 29 CSR reports from five full-service airlines, totaling 1,070,606 tokens. Topic modeling identified eleven prominent topics, which were categorized into five categories: social responsibility, corporate governance, environmental protection, financial performance, and safety and security. The results show several major trajectories in topic shifts. Overall, social responsibility remains the most prominent topic, while environmental sustainability has received renewed attention in recent years. The analysis of the contexts of three selected topic words- “employee”, “SAF”, and “COVID”- further reveals how airlines utilize CSR reports for self-branding and self-legitimation.
Taken together, this study contributes to an evolving understanding of CSR reporting in the aviation industry and provides practical implications for industry practitioners. It also highlights the methodological value of topic modeling in assisting discourse analysis.
References
Baker, P., Gabrielatos, C., KhosraviNik, M., Krzyżanowski, M., McEnery, T., & Wodak, R. (2008). A useful methodological synergy? Combining critical discourse analysis and corpus linguistics to examine discourses of refugees and asylum seekers in the UK press. Discourse & Society, 19(3), 273–306. https://doi.org/10.1177/0957926508088962
Bhatia, A. (2012). The Corporate Social Responsibility Report: The Hybridization of a “Confused” Genre (2007–2011). IEEE Transactions on Professional Communication, 55(3), 221–238. https://doi.org/10.1109/TPC.2012.2205732
Brookes, G., & McEnery, T. (2019). The utility of topic modelling for discourse studies: A critical evaluation. Discourse Studies, 21(1), 3–21. https://doi.org/10.1177/1461445618814032
ICAO. (2025). Environmental Protection | International Civil Aviation Organization. https://www.icao.int/environmental-protection
Jaworska, S., & Nanda, A. (2018). Doing Well by Talking Good: A Topic Modelling-Assisted Discourse Study of Corporate Social Responsibility. Applied Linguistics, 39(3), 373–399. https://doi.org/10.1093/applin/amw014
Partington, A. (2010). Modern Diachronic Corpus-Assisted Discourse Studies (MD-CADS) on UK newspapers: An overview of the project. Corpora, 5(2), 83–108. https://doi.org/10.3366/cor.2010.0101
Abstract
As a carbon-intensive industry, aviation is facing mounting pressure for sustainable development due to its adverse environmental and social problems, such as noise, emissions, pollution, and waste (ICAO, 2025). Along with sustainability policies and actions in practice, airlines are also leveraging corporate social responsibility (CSR) reporting to establish legitimacy, gain trust, and promote favorable images (Bhatia, 2012). While existing literature has examined CSR reporting in industries such as energy, finance, and retail, the aviation sector remains largely underexplored. Methodologically, prior research predominantly adopts a conventional corpus-assisted discourse studies (CADS) approach (Baker et al., 2008; Partington, 2010). Nevertheless, as demonstrated by recent scholarship (Brookes & McEnery, 2019; Jaworska & Nanda, 2018), topic modeling is a advent tool that can effectively facilitate CADS with the identification of thematic categories of large-scale datasets.
This study thus integrates topic modeling with corpus-assisted discourse study to examine the discursive features and diachronic shifts of various foci in the corporate social responsibility (CSR) reports of world-leading airlines from 2019 to 2024. The corpus comprises 29 CSR reports from five full-service airlines, totaling 1,070,606 tokens. Topic modeling identified eleven prominent topics, which were categorized into five categories: social responsibility, corporate governance, environmental protection, financial performance, and safety and security. The results show several major trajectories in topic shifts. Overall, social responsibility remains the most prominent topic, while environmental sustainability has received renewed attention in recent years. The analysis of the contexts of three selected topic words- “employee”, “SAF”, and “COVID”- further reveals how airlines utilize CSR reports for self-branding and self-legitimation.
Taken together, this study contributes to an evolving understanding of CSR reporting in the aviation industry and provides practical implications for industry practitioners. It also highlights the methodological value of topic modeling in assisting discourse analysis.
References
Baker, P., Gabrielatos, C., KhosraviNik, M., Krzyżanowski, M., McEnery, T., & Wodak, R. (2008). A useful methodological synergy? Combining critical discourse analysis and corpus linguistics to examine discourses of refugees and asylum seekers in the UK press. Discourse & Society, 19(3), 273–306. https://doi.org/10.1177/0957926508088962
Bhatia, A. (2012). The Corporate Social Responsibility Report: The Hybridization of a “Confused” Genre (2007–2011). IEEE Transactions on Professional Communication, 55(3), 221–238. https://doi.org/10.1109/TPC.2012.2205732
Brookes, G., & McEnery, T. (2019). The utility of topic modelling for discourse studies: A critical evaluation. Discourse Studies, 21(1), 3–21. https://doi.org/10.1177/1461445618814032
ICAO. (2025). Environmental Protection | International Civil Aviation Organization. https://www.icao.int/environmental-protection
Jaworska, S., & Nanda, A. (2018). Doing Well by Talking Good: A Topic Modelling-Assisted Discourse Study of Corporate Social Responsibility. Applied Linguistics, 39(3), 373–399. https://doi.org/10.1093/applin/amw014
Partington, A. (2010). Modern Diachronic Corpus-Assisted Discourse Studies (MD-CADS) on UK newspapers: An overview of the project. Corpora, 5(2), 83–108. https://doi.org/10.3366/cor.2010.0101
English-Corpora.org has made substantial contributions to retrieving authentic usage examples and frequency data. It has become an indispensable tool for corpus linguistics. However, despite its widespread use, many potential shortcomings of searches using English-Corpora.org remain unaddressed by researchers and reviewers. For example, the Movie Corpus [MC] contained discrepancies between spoken utterances and subtitle transcripts, duplicate examples, incorrect source information, character misrecognition, and subtitles in languages other than English. Consequently, in some cases, researchers may proceed with their analyses and discussions without noticing critical errors in the data. This study aims to depict that almost all issues identified in MC are observed in the Corpus of Contemporary American English [COCA] and to present previously unreported cases. The secondary objective is to propose guidelines for the appropriate use of corpora. Both MC and COCA contain duplicate examples, character misrecognition, and data in languages other than English. Even though COCA is described as a database of “contemporary English,” it in fact includes data from Old English. Searches for sentences and frequencies using part-of-speech tags and *[asterisk] do not always yield a complete set of relevant instances; this has resulted in a failure to retrieve the most frequent cases. Drawing on previous studies and this study’s findings, a set of practical guidelines for COCA are proposed. These guidelines emphasize reporting search strategies, verifying source information, consulting the original source text, disclosing any duplicate or excluded data in frequency counts, and acknowledging the possibility of incomplete retrievals.
English-Corpora.org has made substantial contributions to retrieving authentic usage examples and frequency data. It has become an indispensable tool for corpus linguistics. However, despite its widespread use, many potential shortcomings of searches using English-Corpora.org remain unaddressed by researchers and reviewers. For example, the Movie Corpus [MC] contained discrepancies between spoken utterances and subtitle transcripts, duplicate examples, incorrect source information, character misrecognition, and subtitles in languages other than English. Consequently, in some cases, researchers may proceed with their analyses and discussions without noticing critical errors in the data. This study aims to depict that almost all issues identified in MC are observed in the Corpus of Contemporary American English [COCA] and to present previously unreported cases. The secondary objective is to propose guidelines for the appropriate use of corpora. Both MC and COCA contain duplicate examples, character misrecognition, and data in languages other than English. Even though COCA is described as a database of “contemporary English,” it in fact includes data from Old English. Searches for sentences and frequencies using part-of-speech tags and *[asterisk] do not always yield a complete set of relevant instances; this has resulted in a failure to retrieve the most frequent cases. Drawing on previous studies and this study’s findings, a set of practical guidelines for COCA are proposed. These guidelines emphasize reporting search strategies, verifying source information, consulting the original source text, disclosing any duplicate or excluded data in frequency counts, and acknowledging the possibility of incomplete retrievals.
Refusals in Emerging Romantic Relationships: A Corpus-Assisted Study of Chinese Dating Reality TV Discourse
Refusals are speech acts that “deny to engage in an action proposed by the interlocutor” (Chen et al., 1995, p. 121). As they potentially threaten the interlocutor’s face, speakers often employ various strategies to mitigate their impact (Brown & Levinson, 1987). Previous research has relied on elicited data or naturally occurring interactions (e.g., Su, 2020). While recent Chinese studies have examined refusals in reality TV (e.g., Ren & Woodfield, 2016) and scripted television dramas (e.g., Li & Wongwarapakorn, 2024), corpus-pragmatic research remains limited in exploring how refusals are interactionally negotiated across genders in semi-naturalistic dating contexts, where issues of intimacy and relational negotiation are particularly salient.
Therefore, this study examines refusal dialogues in a Chinese dating reality TV show “心动的信号 (Heart Signal)”. The programme features young adults from diverse backgrounds who cohabit for one month while developing potential romantic and interpersonal relationships. The corpus comprises approximately 68 hours of interaction from Seasons 7 and 8, involving 24 participants from the Greater Bay Area and the Yangtze River Delta region, respectively.
In total, 151 instances of refusals were identified. These refusals occurred in both same-gender and cross-gender interactions, between interlocutors with and without romantic interest, in both single-turn and multi-turn exchanges, and in contexts ranging from everyday conversational refusals to explicit romantic rejections. Refusal speech acts were analysed using a multidimensional framework, focusing on refusal triggers, strategies, syntactic patterns, intersubjectivity, adjuncts, and other interactional features. Preliminary findings suggest that adjunct-supported refusals are more common in explicit romantic rejections, and males refuse more directly in same- than cross-gender interactions.
The study contributes to corpus pragmatics and refusal research by providing empirical insights into how refusals are interactionally negotiated in semi-naturalistic romantic contexts, particularly concerning facework, relational positioning, and intimacy construction in contemporary Chinese mediated discourse.
References
Brown, P., & Levinson, S. C. (1987). Politeness: Some universals in language usage. Cambridge University Press.
Chen, X., Ye, L., & Zhang, Y. (1995). Refusing in Chinese. In G. Kasper (Ed.), Pragmatics of Chinese as native and target language (pp. 121–161). University of Hawai‘i at Mānoa.
Li, Y., & Wongwaropakorn, W. (2024). Analyzing politeness and refusal speech acts in popular Chinese television drama series. Cogent Arts & Humanities, 11(1). https://doi.org/10.1080/23311983.2024.2367327
Ren, W., & Woodfield, H. (2016). Chinese females׳ date refusals in reality TV shows: Expressing involvement or independence? Discourse, Context & Media, 13, 89–97. https://doi.org/10.1016/j.dcm.2016.05.008
Su, Y. (2020). Yes or no: Ostensible versus genuine refusals in Mandarin invitational and offering discourse. Journal of Pragmatics, 162, 1–16. https://doi.org/10.1016/j.pragma.2020.03.007
Refusals in Emerging Romantic Relationships: A Corpus-Assisted Study of Chinese Dating Reality TV Discourse
Refusals are speech acts that “deny to engage in an action proposed by the interlocutor” (Chen et al., 1995, p. 121). As they potentially threaten the interlocutor’s face, speakers often employ various strategies to mitigate their impact (Brown & Levinson, 1987). Previous research has relied on elicited data or naturally occurring interactions (e.g., Su, 2020). While recent Chinese studies have examined refusals in reality TV (e.g., Ren & Woodfield, 2016) and scripted television dramas (e.g., Li & Wongwarapakorn, 2024), corpus-pragmatic research remains limited in exploring how refusals are interactionally negotiated across genders in semi-naturalistic dating contexts, where issues of intimacy and relational negotiation are particularly salient.
Therefore, this study examines refusal dialogues in a Chinese dating reality TV show “心动的信号 (Heart Signal)”. The programme features young adults from diverse backgrounds who cohabit for one month while developing potential romantic and interpersonal relationships. The corpus comprises approximately 68 hours of interaction from Seasons 7 and 8, involving 24 participants from the Greater Bay Area and the Yangtze River Delta region, respectively.
In total, 151 instances of refusals were identified. These refusals occurred in both same-gender and cross-gender interactions, between interlocutors with and without romantic interest, in both single-turn and multi-turn exchanges, and in contexts ranging from everyday conversational refusals to explicit romantic rejections. Refusal speech acts were analysed using a multidimensional framework, focusing on refusal triggers, strategies, syntactic patterns, intersubjectivity, adjuncts, and other interactional features. Preliminary findings suggest that adjunct-supported refusals are more common in explicit romantic rejections, and males refuse more directly in same- than cross-gender interactions.
The study contributes to corpus pragmatics and refusal research by providing empirical insights into how refusals are interactionally negotiated in semi-naturalistic romantic contexts, particularly concerning facework, relational positioning, and intimacy construction in contemporary Chinese mediated discourse.
References
Brown, P., & Levinson, S. C. (1987). Politeness: Some universals in language usage. Cambridge University Press.
Chen, X., Ye, L., & Zhang, Y. (1995). Refusing in Chinese. In G. Kasper (Ed.), Pragmatics of Chinese as native and target language (pp. 121–161). University of Hawai‘i at Mānoa.
Li, Y., & Wongwaropakorn, W. (2024). Analyzing politeness and refusal speech acts in popular Chinese television drama series. Cogent Arts & Humanities, 11(1). https://doi.org/10.1080/23311983.2024.2367327
Ren, W., & Woodfield, H. (2016). Chinese females׳ date refusals in reality TV shows: Expressing involvement or independence? Discourse, Context & Media, 13, 89–97. https://doi.org/10.1016/j.dcm.2016.05.008
Su, Y. (2020). Yes or no: Ostensible versus genuine refusals in Mandarin invitational and offering discourse. Journal of Pragmatics, 162, 1–16. https://doi.org/10.1016/j.pragma.2020.03.007
Within the framework of World Englishes, this study deals with a significant gap in learner corpus research on spoken English in South Asia. While Sri Lankan English (SLE) has attracted scholarly attention within the country, previous studies have primarily focused on adult speakers or descriptive accounts of phonological features. To date, little research has systematically documented the spoken production of young SLE learners through an openly accessible learner corpus. Consequently, there is limited empirical evidence on how multilingual learners mobilize the linguistic resources of contemporary SLE. To address this gap, the present project aims to construct and disseminate an open corpus of learner monologues in SLE. The study is guided by the following research questions: (1) What phonological features appear in the English monologues of young Sri Lankan learners? and (2) What methodological considerations are necessary for developing a sustainable and publicly accessible learner spoken corpus in a multilingual context? A pilot corpus comprising 90 short monologues by 30 participants on 3 topics- self-introduction, leisure time, and opinions about social media (6,982 word tokens) was compiled and qualitatively analyzed. Particular attention was paid to recording procedures, transcription protocols, metadata design, speaker profiling, and corpus management. Preliminary analysis showed results such as recurrent realizations of the interdental fricatives /θ/ and /ð/, variable rhoticity, occurrences of centering diphthongs (/ɪə/, /eə/, /ʊə/), and vowel-quality variations. The findings have informed the redesign of data-collection protocols for a larger corpus of approximately 500 learner monologues, in collaboration with partner institutions in SL. The project contributes to learner corpus research, the documentation of Sri Lankan English, and broader discussions of linguistic diversity and World Englishes in South Asia.
Within the framework of World Englishes, this study deals with a significant gap in learner corpus research on spoken English in South Asia. While Sri Lankan English (SLE) has attracted scholarly attention within the country, previous studies have primarily focused on adult speakers or descriptive accounts of phonological features. To date, little research has systematically documented the spoken production of young SLE learners through an openly accessible learner corpus. Consequently, there is limited empirical evidence on how multilingual learners mobilize the linguistic resources of contemporary SLE. To address this gap, the present project aims to construct and disseminate an open corpus of learner monologues in SLE. The study is guided by the following research questions: (1) What phonological features appear in the English monologues of young Sri Lankan learners? and (2) What methodological considerations are necessary for developing a sustainable and publicly accessible learner spoken corpus in a multilingual context? A pilot corpus comprising 90 short monologues by 30 participants on 3 topics- self-introduction, leisure time, and opinions about social media (6,982 word tokens) was compiled and qualitatively analyzed. Particular attention was paid to recording procedures, transcription protocols, metadata design, speaker profiling, and corpus management. Preliminary analysis showed results such as recurrent realizations of the interdental fricatives /θ/ and /ð/, variable rhoticity, occurrences of centering diphthongs (/ɪə/, /eə/, /ʊə/), and vowel-quality variations. The findings have informed the redesign of data-collection protocols for a larger corpus of approximately 500 learner monologues, in collaboration with partner institutions in SL. The project contributes to learner corpus research, the documentation of Sri Lankan English, and broader discussions of linguistic diversity and World Englishes in South Asia.
Metadiscourse studies and corpus research have provided extensive accounts of how academic writers organise claims and position themselves in relation to readers, evidence, and disciplinary knowledge. Within this work, epistemic stance refers to how writers signal certainty, probability, evidential support, and commitment to a claim. Such markers are often grouped into broad categories, especially hedges and boosters. While useful, this distinction does not capture gradience among individual items: might, probably, clearly, and demonstrate, for example, all affect the force of a claim, but not to the same degree.
This paper reports a study combining experimental scaling and corpus analysis to develop an item-level scale of academic claim strength. We compile a set of 108 epistemic and stance-related expressions across six dimensions: likelihood, certainty, evidential support, claim positioning, generality, and degree. These items are evaluated using Best–Worst Scaling, a comparative judgement method in which participants select the strongest and weakest item from four-item sets. The BWS design is organised by category, with calibration anchors used to normalise scores and bridge items used to link the category-specific scales.
We apply the resulting scale to a preliminary corpus analysis of academic abstracts. Each occurrence of a target expression is assigned a strength score, allowing ordinary frequency counts to be compared with profiles showing whether texts rely mainly on weaker, moderate, or stronger stance expressions. This proof-of-concept analysis examines whether texts or subcorpora with similar stance-marker frequencies differ in the strength profile of the expressions they use. The paper contributes a scored item list, a corpus-based demonstration of strength-weighted stance profiling, and an open web-based tool for designing BWS studies.
Metadiscourse studies and corpus research have provided extensive accounts of how academic writers organise claims and position themselves in relation to readers, evidence, and disciplinary knowledge. Within this work, epistemic stance refers to how writers signal certainty, probability, evidential support, and commitment to a claim. Such markers are often grouped into broad categories, especially hedges and boosters. While useful, this distinction does not capture gradience among individual items: might, probably, clearly, and demonstrate, for example, all affect the force of a claim, but not to the same degree.
This paper reports a study combining experimental scaling and corpus analysis to develop an item-level scale of academic claim strength. We compile a set of 108 epistemic and stance-related expressions across six dimensions: likelihood, certainty, evidential support, claim positioning, generality, and degree. These items are evaluated using Best–Worst Scaling, a comparative judgement method in which participants select the strongest and weakest item from four-item sets. The BWS design is organised by category, with calibration anchors used to normalise scores and bridge items used to link the category-specific scales.
We apply the resulting scale to a preliminary corpus analysis of academic abstracts. Each occurrence of a target expression is assigned a strength score, allowing ordinary frequency counts to be compared with profiles showing whether texts rely mainly on weaker, moderate, or stronger stance expressions. This proof-of-concept analysis examines whether texts or subcorpora with similar stance-marker frequencies differ in the strength profile of the expressions they use. The paper contributes a scored item list, a corpus-based demonstration of strength-weighted stance profiling, and an open web-based tool for designing BWS studies.
In translating technical documents, such as product manuals, controlling sentence structure as well as vocabulary is crucial for ensuring clarity and consistency. While generative AI (GenAI) translation allows flexible prompt-based control, making its combination with controlled languages highly promising, existing frameworks, such as ASD-STE100 Simplified Technical English (ASD, 2025), lack specifications for fine-grained structural control, including clause order. To bridge this gap, we must establish rules that explicitly map textual functions to specific syntactic patterns. Focusing on conjunctive patterns in Japanese-to-English translation, this study addresses three research questions:
(1) What functional elements (i.e., clauses or phrases) are linked by connective expressions?
(2) Is there an optimum ordering of these elements for controlled authoring and translation?
(3) Can these identified patterns effectively improve GenAI-based translation quality?
Using a parallel corpus of 134,929 Japanese-English sentence pairs (16,282 unique pairs) from construction machinery manuals, we sample 10–20 diverse instances for each targeted English connective expression (e.g., if, before, so that, and by V-ing). Drawing on insights from Systemic Functional Grammar (Halliday & Matthiessen, 2013) and cross-referencing the Japanese source texts, we inductively label the communicative functions of elements surrounding these expressions (e.g., ACTION, GOAL, REASON, and CONDITION). This helps establish approved conjunctive patterns (e.g., an if + CONDITION clause should precede an ACTION clause). These defined patterns are integrated into GenAI prompts to achieve controlled translation. Finally, we evaluate the performance of these prompts in terms of coverage of defined patterns and consistency of translation results.
Provisionally, we have identified 20 functions and defined 2–4 approved patterns for each connective. This research offers a novel framework for a functionally oriented controlled language capable of regulating sentence structure, including information ordering, in both source and target languages. Practically, it advances prompt engineering and significantly contributes to enhancing technical translation quality.
References:
ASD (2025). ASD-STE100 Simplified Technical English (Issue 9). https://www.asd-ste100.org
Halliday, M. A. K. and Matthiessen, C. M. (2013). Halliday’s introduction to functional grammar (4th ed.). Routledge.
In translating technical documents, such as product manuals, controlling sentence structure as well as vocabulary is crucial for ensuring clarity and consistency. While generative AI (GenAI) translation allows flexible prompt-based control, making its combination with controlled languages highly promising, existing frameworks, such as ASD-STE100 Simplified Technical English (ASD, 2025), lack specifications for fine-grained structural control, including clause order. To bridge this gap, we must establish rules that explicitly map textual functions to specific syntactic patterns. Focusing on conjunctive patterns in Japanese-to-English translation, this study addresses three research questions:
(1) What functional elements (i.e., clauses or phrases) are linked by connective expressions?
(2) Is there an optimum ordering of these elements for controlled authoring and translation?
(3) Can these identified patterns effectively improve GenAI-based translation quality?
Using a parallel corpus of 134,929 Japanese-English sentence pairs (16,282 unique pairs) from construction machinery manuals, we sample 10–20 diverse instances for each targeted English connective expression (e.g., if, before, so that, and by V-ing). Drawing on insights from Systemic Functional Grammar (Halliday & Matthiessen, 2013) and cross-referencing the Japanese source texts, we inductively label the communicative functions of elements surrounding these expressions (e.g., ACTION, GOAL, REASON, and CONDITION). This helps establish approved conjunctive patterns (e.g., an if + CONDITION clause should precede an ACTION clause). These defined patterns are integrated into GenAI prompts to achieve controlled translation. Finally, we evaluate the performance of these prompts in terms of coverage of defined patterns and consistency of translation results.
Provisionally, we have identified 20 functions and defined 2–4 approved patterns for each connective. This research offers a novel framework for a functionally oriented controlled language capable of regulating sentence structure, including information ordering, in both source and target languages. Practically, it advances prompt engineering and significantly contributes to enhancing technical translation quality.
References:
ASD (2025). ASD-STE100 Simplified Technical English (Issue 9). https://www.asd-ste100.org
Halliday, M. A. K. and Matthiessen, C. M. (2013). Halliday’s introduction to functional grammar (4th ed.). Routledge.
Abstract :
Research on the translation and dissemination of Chinese classics from the perspective of Orientalist theory has yielded fruitful results. However, existing studies have largely focused on translated texts themselves, while paratextual elements, which shape the interpretation and reception of translated works, have not received sufficient attention. As an important component of a book’s external form, the cover plays a crucial role in the cross-cultural dissemination of classics through its visual signs and design elements. Against this background, drawing on Edward W. Said’s theory of Orientalism, this study examines the covers of English and Spanish translations of The Art of War . It focuses on three elements of the translated book covers—verbal elements, visual images, and image-text relations—with the aim of examining their multimodal features in different cultural contexts of dissemination from the perspective of Orientalism.
This study adopts a corpus-based research method and constructs a self-built cover corpus comprising holdings from the national libraries of five countries: the United States, the United Kingdom, Spain, Mexico, and Argentina. The corpus includes 135 English translations published in the English-speaking world and 112 Spanish translations published in the Spanish-speaking world. The findings indicate that, compared with the covers of English translations, those of Spanish ones display a more pronounced Orientalist tendency. In the verbal mode, Spanish covers are more inclined to represent The Art of War as possessing essentialized features that are singular, static, and timeless. In the visual mode, Orientalist features such as exoticization and mystification are more evident on Spanish covers, while phenomena such as cultural conflation and cultural appropriation are also more prominent. From the perspective of image-text relations, the covers of English translations demonstrate greater consistency between verbal and visual elements and correspond more closely to the book’s thematic focus on military strategy. The study further suggests that these differences result from the combined influence of factors such as the degree of canonization of the classic, the commercial logic of publishing, and divergent traditions of imagining the Orient in the English-speaking and Spanish-speaking worlds.
Keywords: Orientalism; Other-representation; The Art of War; book covers; cross-cultural dissemination
Abstract :
Research on the translation and dissemination of Chinese classics from the perspective of Orientalist theory has yielded fruitful results. However, existing studies have largely focused on translated texts themselves, while paratextual elements, which shape the interpretation and reception of translated works, have not received sufficient attention. As an important component of a book’s external form, the cover plays a crucial role in the cross-cultural dissemination of classics through its visual signs and design elements. Against this background, drawing on Edward W. Said’s theory of Orientalism, this study examines the covers of English and Spanish translations of The Art of War . It focuses on three elements of the translated book covers—verbal elements, visual images, and image-text relations—with the aim of examining their multimodal features in different cultural contexts of dissemination from the perspective of Orientalism.
This study adopts a corpus-based research method and constructs a self-built cover corpus comprising holdings from the national libraries of five countries: the United States, the United Kingdom, Spain, Mexico, and Argentina. The corpus includes 135 English translations published in the English-speaking world and 112 Spanish translations published in the Spanish-speaking world. The findings indicate that, compared with the covers of English translations, those of Spanish ones display a more pronounced Orientalist tendency. In the verbal mode, Spanish covers are more inclined to represent The Art of War as possessing essentialized features that are singular, static, and timeless. In the visual mode, Orientalist features such as exoticization and mystification are more evident on Spanish covers, while phenomena such as cultural conflation and cultural appropriation are also more prominent. From the perspective of image-text relations, the covers of English translations demonstrate greater consistency between verbal and visual elements and correspond more closely to the book’s thematic focus on military strategy. The study further suggests that these differences result from the combined influence of factors such as the degree of canonization of the classic, the commercial logic of publishing, and divergent traditions of imagining the Orient in the English-speaking and Spanish-speaking worlds.
Keywords: Orientalism; Other-representation; The Art of War; book covers; cross-cultural dissemination
The application of Large Language Models (LLMs) for the task of writing assessment has become an increasingly fertile research area in recent years (Mizumoto & Eguchi, 2023; Uchida & Negishi, 2025; Yamashita, 2024). These approaches generally use commercial LLMs, occasionally integrated with linguistic features. However, commercial language models have several limitations. First, their use limits the production of publicly available assessment tools due to API fees. Second, sending learner data to third-party APIs raises privacy concerns that the terms of many corpora, compiled before generative AI, were not written to account for. Third, data leakage of learner corpus data to commercial LLMs might inflate prediction accuracy. To address these issues, the following research questions are posed:
1. How accurately can local LLMs classify English learner writing by CEFR level? (QWK and F-Score)
2. Which architecture and strategy combination performs the strongest?
The current study used the Write & Improve Corpus (Nicholls et al., 2024) because CEFR levels were applied by annotators who were trained and certified as examiners. The corpus compilers also warn against the data leakage issue when users download the corpus. To avoid the other issues associated with commercial LLMs, CEFR level assignment to English learner writing was explored with encoder-decoder (T5Gemma) and decoder-only (Gemma 4) local LLMs. For the encoder-decoder model, an ordinal regression neural network approach (Shi et al., 2023) is compared with a text classification approach. For the decoder-only model, various prompting strategies are trialed, such as increasing the number of few-shot examples, and asking the LLM to update its prior knowledge in a conceptually Bayesian approach.
In this poster presentation, preliminary findings will be presented along with their implications. The results will indicate whether local LLMs are a viable alternative for the task of English learner writing classification by CEFR level.
References
Mizumoto, A., & Eguchi, M. (2023). Exploring the potential of using an AI language model for automated essay scoring. Research Methods in Applied Linguistics, 2(2), 100050. https://doi.org/10.1016/j.rmal.2023.100050
Nicholls, D., Caines, A., & Buttery, P. (2024). The Write & Improve Corpus 2024: Error-annotated and CEFR-labelled essays by learners of English. Cambridge University Press & Assessment. https://doi.org/10.17863/CAM.112997
Shi, X., Cao, W., & Raschka, S. (2023). Deep neural networks for rank-consistent ordinal regression based on conditional probabilities. Pattern Analysis and Applications, 26(3), 941–955. https://doi.org/10.1007/s10044-023-01181-9
Uchida, S., & Negishi, M. (2025). Assigning CEFR-J levels to English learners’ writing: An approach using lexical metrics and generative AI. Research Methods in Applied Linguistics, 4(2), 100199. https://doi.org/10.1016/j.rmal.2025.100199
Yamashita, T. (2024). An application of many-facet Rasch measurement to evaluate automated essay scoring: A case of ChatGPT-4.0. Research Methods in Applied Linguistics, 3(3), 100133. https://doi.org/10.1016/j.rmal.2024.100133
The application of Large Language Models (LLMs) for the task of writing assessment has become an increasingly fertile research area in recent years (Mizumoto & Eguchi, 2023; Uchida & Negishi, 2025; Yamashita, 2024). These approaches generally use commercial LLMs, occasionally integrated with linguistic features. However, commercial language models have several limitations. First, their use limits the production of publicly available assessment tools due to API fees. Second, sending learner data to third-party APIs raises privacy concerns that the terms of many corpora, compiled before generative AI, were not written to account for. Third, data leakage of learner corpus data to commercial LLMs might inflate prediction accuracy. To address these issues, the following research questions are posed:
1. How accurately can local LLMs classify English learner writing by CEFR level? (QWK and F-Score)
2. Which architecture and strategy combination performs the strongest?
The current study used the Write & Improve Corpus (Nicholls et al., 2024) because CEFR levels were applied by annotators who were trained and certified as examiners. The corpus compilers also warn against the data leakage issue when users download the corpus. To avoid the other issues associated with commercial LLMs, CEFR level assignment to English learner writing was explored with encoder-decoder (T5Gemma) and decoder-only (Gemma 4) local LLMs. For the encoder-decoder model, an ordinal regression neural network approach (Shi et al., 2023) is compared with a text classification approach. For the decoder-only model, various prompting strategies are trialed, such as increasing the number of few-shot examples, and asking the LLM to update its prior knowledge in a conceptually Bayesian approach.
In this poster presentation, preliminary findings will be presented along with their implications. The results will indicate whether local LLMs are a viable alternative for the task of English learner writing classification by CEFR level.
References
Mizumoto, A., & Eguchi, M. (2023). Exploring the potential of using an AI language model for automated essay scoring. Research Methods in Applied Linguistics, 2(2), 100050. https://doi.org/10.1016/j.rmal.2023.100050
Nicholls, D., Caines, A., & Buttery, P. (2024). The Write & Improve Corpus 2024: Error-annotated and CEFR-labelled essays by learners of English. Cambridge University Press & Assessment. https://doi.org/10.17863/CAM.112997
Shi, X., Cao, W., & Raschka, S. (2023). Deep neural networks for rank-consistent ordinal regression based on conditional probabilities. Pattern Analysis and Applications, 26(3), 941–955. https://doi.org/10.1007/s10044-023-01181-9
Uchida, S., & Negishi, M. (2025). Assigning CEFR-J levels to English learners’ writing: An approach using lexical metrics and generative AI. Research Methods in Applied Linguistics, 4(2), 100199. https://doi.org/10.1016/j.rmal.2025.100199
Yamashita, T. (2024). An application of many-facet Rasch measurement to evaluate automated essay scoring: A case of ChatGPT-4.0. Research Methods in Applied Linguistics, 3(3), 100133. https://doi.org/10.1016/j.rmal.2024.100133
Background
L2 learners of Japanese may find studying grammar difficult, as Japanese grammar
patterns (JGPs) are often polysemous. Data-driven learning (DDL) systems help
users grasp linguistic patterns in context (Chujo et al. 2015). Creating a DDL system
for Japanese grammar requires a corpus of sentences annotated for JGPs and their
senses. Since such a corpus does not exist and manual annotation is resource-intensive,
we will test an LLM-assisted annotation pipeline to construct one. In this study, we will
investigate LLMs’ ability to (1) determine the presence of JGPs in sentences and
(2) annotate their positions and senses.
Data and methods
We compile a set of JGPs and example sentences from a dictionary of Japanese
grammar (Group Jammassy et al. 2015). As a pilot investigation, we use 103 JGPs
and 418 example sentences, selected for diversity and feasibility of conversion from
a pedagogical resource.
To test (1), six LLMs were supplied with sentence-JGP pairs accompanied by pedagogical
descriptions, then asked to determine the target JGPs’ presence in sentences. For (2),
once the LLM confirms the presence of a JGP, it will be asked to span-annotate it and
assign it a sense label.
Provisional results
In (1), LLMs achieved accuracy, precision and recall between 83% and over 99%, showing
that they can be an effective first-pass filter for identifying sentences with specified JGPs.
Nevertheless, models systematically struggled to identify JGPs with nuanced descriptions.
Experimental setup for (2) is underway - a test set of sentences is being human-annotated
with JGPs’ positions and senses.
Potential contributions
Our results will expand the known capabilities and limitations of AI-assisted corpus
annotation. We will demonstrate an LLM-assisted annotation pipeline of polysemous
grammar patterns and contribute a corpus of span- and sense-annotated Japanese
sentences, applicable to language learning systems.
References
Chujo, Kiyomi, Kathryn Oghigian, and Shiro Akasegawa. 2015. ‘A Corpus and
Grammatical Browsing System for Remedial EFL Learners’. In Multiple Affordances
of Language Corpora for Data-Driven Learning, edited by Agnieszka Leńko-Szymańska
and Alex Boulton. Studies in Corpus Linguistics. John Benjamins Publishing Company.
https://doi.org/10.1075/scl.69.06chu.
Group Jammassy, Yuriko Sunakawa, Priscilla Ishida, et al. 2015. A Handbook
of Japanese Grammar Patterns for Teachers and Learners. Edited by Group Jammassey.
Kurosio Publishers. https://www.9640.jp/nihongo/en/detail/?678.
Background
L2 learners of Japanese may find studying grammar difficult, as Japanese grammar
patterns (JGPs) are often polysemous. Data-driven learning (DDL) systems help
users grasp linguistic patterns in context (Chujo et al. 2015). Creating a DDL system
for Japanese grammar requires a corpus of sentences annotated for JGPs and their
senses. Since such a corpus does not exist and manual annotation is resource-intensive,
we will test an LLM-assisted annotation pipeline to construct one. In this study, we will
investigate LLMs’ ability to (1) determine the presence of JGPs in sentences and
(2) annotate their positions and senses.
Data and methods
We compile a set of JGPs and example sentences from a dictionary of Japanese
grammar (Group Jammassy et al. 2015). As a pilot investigation, we use 103 JGPs
and 418 example sentences, selected for diversity and feasibility of conversion from
a pedagogical resource.
To test (1), six LLMs were supplied with sentence-JGP pairs accompanied by pedagogical
descriptions, then asked to determine the target JGPs’ presence in sentences. For (2),
once the LLM confirms the presence of a JGP, it will be asked to span-annotate it and
assign it a sense label.
Provisional results
In (1), LLMs achieved accuracy, precision and recall between 83% and over 99%, showing
that they can be an effective first-pass filter for identifying sentences with specified JGPs.
Nevertheless, models systematically struggled to identify JGPs with nuanced descriptions.
Experimental setup for (2) is underway - a test set of sentences is being human-annotated
with JGPs’ positions and senses.
Potential contributions
Our results will expand the known capabilities and limitations of AI-assisted corpus
annotation. We will demonstrate an LLM-assisted annotation pipeline of polysemous
grammar patterns and contribute a corpus of span- and sense-annotated Japanese
sentences, applicable to language learning systems.
References
Chujo, Kiyomi, Kathryn Oghigian, and Shiro Akasegawa. 2015. ‘A Corpus and
Grammatical Browsing System for Remedial EFL Learners’. In Multiple Affordances
of Language Corpora for Data-Driven Learning, edited by Agnieszka Leńko-Szymańska
and Alex Boulton. Studies in Corpus Linguistics. John Benjamins Publishing Company.
https://doi.org/10.1075/scl.69.06chu.
Group Jammassy, Yuriko Sunakawa, Priscilla Ishida, et al. 2015. A Handbook
of Japanese Grammar Patterns for Teachers and Learners. Edited by Group Jammassey.
Kurosio Publishers. https://www.9640.jp/nihongo/en/detail/?678.
Discursive Construction of Bride Price on Chinese Social Media: A Corpus-Assisted Comparative Analysis of Douban and Hupu
Bride price is a traditional Chinese custom involving monetary and material transfers from the groom to the bride. While historically embedded in a patrilocal marriage system, the practice is increasingly contested in contemporary society. Its inflation outpaces income growth, escalating financial strain on men and compounding anxieties surrounding marriage and declining fertility. Since 2019, the Chinese central government has repeatedly targeted high bride prices as an “unhealthy social trend,” with state media framing the practice negatively and promoting low- or no-bride-price marriages (Ma, 2024a). Conversely, netizen responses remain fragmented, ranging from support to rejection (Ma, 2024b). This divergence has intensified in 2026, driven by recent judicial rulings on divorce cases permitting the legal restitution of bride prices.
Therefore, focusing on netizens’ perspectives, this study examines the discursive construction of bride price on Chinese social media by contrastively analyzing online discussions on Douban and Hupu, platforms predominantly used by female and male users, respectively. Using Octoparse, a web-scraping tool, posts containing “彩礼” (cǎilǐ, bride price) published from 2019 onwards were collected. The corpus was restricted to high-engagement threads, and the data underwent corpus-assisted discourse analysis, including frequency and collocation analyses, to identify salient discursive patterns and rhetorical strategies such as irony.
Preliminary findings suggest that discourse on Hupu is predominantly agent-centric, emphasizing individual actors and personal responsibility, whereas discourse on Douban is event-centric, framing bride price in relation to broader social issues and structural conditions. The findings also indicate that support for the practice is expressed primarily by female netizens.
This study contributes to research on online discourse by offering empirical insights into netizens’ linguistic strategies in negotiating gender topics within specific online communities. Additionally, it seeks to advance existing methodological paradigms, optimizing data collection and analytical approaches for social media communication.
References
Ma, J. (2024a). Alignment, negation, and androcentricity: representation of bride price by Chinese state media. Social Semiotics, 1-20.
Ma, J. (2024b). A comparative study of bride price representation on Chinese social media. Journal of Humanities, Arts and Social Science, 8(11)
Discursive Construction of Bride Price on Chinese Social Media: A Corpus-Assisted Comparative Analysis of Douban and Hupu
Bride price is a traditional Chinese custom involving monetary and material transfers from the groom to the bride. While historically embedded in a patrilocal marriage system, the practice is increasingly contested in contemporary society. Its inflation outpaces income growth, escalating financial strain on men and compounding anxieties surrounding marriage and declining fertility. Since 2019, the Chinese central government has repeatedly targeted high bride prices as an “unhealthy social trend,” with state media framing the practice negatively and promoting low- or no-bride-price marriages (Ma, 2024a). Conversely, netizen responses remain fragmented, ranging from support to rejection (Ma, 2024b). This divergence has intensified in 2026, driven by recent judicial rulings on divorce cases permitting the legal restitution of bride prices.
Therefore, focusing on netizens’ perspectives, this study examines the discursive construction of bride price on Chinese social media by contrastively analyzing online discussions on Douban and Hupu, platforms predominantly used by female and male users, respectively. Using Octoparse, a web-scraping tool, posts containing “彩礼” (cǎilǐ, bride price) published from 2019 onwards were collected. The corpus was restricted to high-engagement threads, and the data underwent corpus-assisted discourse analysis, including frequency and collocation analyses, to identify salient discursive patterns and rhetorical strategies such as irony.
Preliminary findings suggest that discourse on Hupu is predominantly agent-centric, emphasizing individual actors and personal responsibility, whereas discourse on Douban is event-centric, framing bride price in relation to broader social issues and structural conditions. The findings also indicate that support for the practice is expressed primarily by female netizens.
This study contributes to research on online discourse by offering empirical insights into netizens’ linguistic strategies in negotiating gender topics within specific online communities. Additionally, it seeks to advance existing methodological paradigms, optimizing data collection and analytical approaches for social media communication.
References
Ma, J. (2024a). Alignment, negation, and androcentricity: representation of bride price by Chinese state media. Social Semiotics, 1-20.
Ma, J. (2024b). A comparative study of bride price representation on Chinese social media. Journal of Humanities, Arts and Social Science, 8(11)
Connective expressions play a crucial role in enhancing text clarity and cohesion. In Japanese, these expressions, such as jissai no tokoro (“as a matter of fact”) and soushitara (“accordingly” or “therefore”), pose challenges for learners and lexicographers due to their multi-word nature, orthographic variations, and multi-functional characteristics. Traditionally, compiling and classifying these expressions required labour-intensive manual analysis of corpus examples, which often lacked procedural transparency and failed to capture low-frequency linguistic functions. To address these limitations, this study proposes a multi-stage pipeline driven by large language models (LLMs) to automate the functional analysis of connectives and validates its effectiveness.
We applied this pipeline to 846 candidate expressions obtained and extended from previous studies (e.g., Ishiguro, 2016; Nishina et al., 2017). First, we instructed GPT-5.5 to filter the candidate expressions based on acceptability criteria we formulated with reference to Ishiguro (2008), yielding a refined list of 463 connectives. Next, we collected two types of usage examples: (1) synthetic data, where GPT-5.5 generated 15 diverse examples per connective, and (2) corpus data, where up to 20 examples were extracted from the Balanced Corpus of Contemporary Written Japanese (BCCWJ) (Maekawa et al., 2014), with GPT-5.5 filtering out non-connective uses. For both sources, GPT-5.5 automatically grouped the instances and assigned functional categories according to Ishiguro’s (2016) typology.
Excluding 58 expressions for which the corpus yielded no valid examples, we analysed the results for the remaining 405 connectives. Across these items, 577 and 516 functions were identified from the synthetic data and corpus data, respectively. Of these, 443 functions (68.2%) were common to both, 134 (20.6%) were unique to the synthetic data, and 73 (11.2%) were unique to the corpus data. These findings suggest that integrating both synthetic data and corpus data is essential for comprehensively identifying the diverse functions of connective expressions.
References:
Ishiguro, K. (2008). Bunsho wa setsuzokushi de kimaru [Writing is determined by connectives]. Kobunsha. (in Japanese).
Ishiguro, K. (2016). “Setsuzokushi” no gijutsu [Techniques of connectives]. JITSUMUKYOIKU-SHUPPAN. (in Japanese).
Maekawa, K., Yamazaki, M., Ogiso, T., Maruyama, T., Ogura, H., Kashino, W., Koiso, H., Yamaguchi, M., Tanaka, M., and Den, Y. (2014). Balanced corpus of contemporary written Japanese. Language Resources and Evaluation, 48, 345–371.
Nishina, K., Yagi, Y., Hodošček, B., and Abekawa, T. (2017). Construction of a connectives dictionary for academic writing assistance system. Mathematical Linguistics, 31(2), 160–176. (in Japanese).
Connective expressions play a crucial role in enhancing text clarity and cohesion. In Japanese, these expressions, such as jissai no tokoro (“as a matter of fact”) and soushitara (“accordingly” or “therefore”), pose challenges for learners and lexicographers due to their multi-word nature, orthographic variations, and multi-functional characteristics. Traditionally, compiling and classifying these expressions required labour-intensive manual analysis of corpus examples, which often lacked procedural transparency and failed to capture low-frequency linguistic functions. To address these limitations, this study proposes a multi-stage pipeline driven by large language models (LLMs) to automate the functional analysis of connectives and validates its effectiveness.
We applied this pipeline to 846 candidate expressions obtained and extended from previous studies (e.g., Ishiguro, 2016; Nishina et al., 2017). First, we instructed GPT-5.5 to filter the candidate expressions based on acceptability criteria we formulated with reference to Ishiguro (2008), yielding a refined list of 463 connectives. Next, we collected two types of usage examples: (1) synthetic data, where GPT-5.5 generated 15 diverse examples per connective, and (2) corpus data, where up to 20 examples were extracted from the Balanced Corpus of Contemporary Written Japanese (BCCWJ) (Maekawa et al., 2014), with GPT-5.5 filtering out non-connective uses. For both sources, GPT-5.5 automatically grouped the instances and assigned functional categories according to Ishiguro’s (2016) typology.
Excluding 58 expressions for which the corpus yielded no valid examples, we analysed the results for the remaining 405 connectives. Across these items, 577 and 516 functions were identified from the synthetic data and corpus data, respectively. Of these, 443 functions (68.2%) were common to both, 134 (20.6%) were unique to the synthetic data, and 73 (11.2%) were unique to the corpus data. These findings suggest that integrating both synthetic data and corpus data is essential for comprehensively identifying the diverse functions of connective expressions.
References:
Ishiguro, K. (2008). Bunsho wa setsuzokushi de kimaru [Writing is determined by connectives]. Kobunsha. (in Japanese).
Ishiguro, K. (2016). “Setsuzokushi” no gijutsu [Techniques of connectives]. JITSUMUKYOIKU-SHUPPAN. (in Japanese).
Maekawa, K., Yamazaki, M., Ogiso, T., Maruyama, T., Ogura, H., Kashino, W., Koiso, H., Yamaguchi, M., Tanaka, M., and Den, Y. (2014). Balanced corpus of contemporary written Japanese. Language Resources and Evaluation, 48, 345–371.
Nishina, K., Yagi, Y., Hodošček, B., and Abekawa, T. (2017). Construction of a connectives dictionary for academic writing assistance system. Mathematical Linguistics, 31(2), 160–176. (in Japanese).