PLENARY SPEAKERS

  • Eric Friginal (The Hong Kong Polytechnic University)


    https://www.polyu.edu.hk/engl/people/academic-staff/prof-eric-friginal/

    Plenary Session 2 [September 9, 16:50 - 17:50]

    From Corpora to Consequence: Applied Corpus Linguistics, Professional Communication, and Policy in the Age of AI
    View Abstract

    At its most powerful, corpus linguistics is a way of making language visible: revealing the patterns, assumptions, routines, silences, and inequalities that shape social and professional life. In the age of generative AI, this role is not diminished; it is intensified. As large language models increasingly mediate how people write, work, learn, decide, and communicate, applied corpus linguistics (ACL) offers essential methods for examining language at scale while remaining grounded in context, interpretation, accountability, and real-world consequence (Friginal, 2025; Thompson & Friginal, 2021). In this presentation, I reflect on the continuing relevance of ACL through the lens of my own research on professional communication across intercultural contexts and institutional settings. With domains including aviation, customer service, health communication, office talk, and other workplace and public-facing environments, my work has explored how recurrent linguistic patterns connect to professional practice, intercultural understanding, institutional norms, and decision-making. These connections have helped inform both micro-level policies—such as interactional guidelines, training practices, communication protocols, and feedback frameworks—and macro-level policies concerned with organisational communication, safety, service quality, equity, and public trust. I also situate this discussion within the broader development of the field, drawing on my role as founding Editor-in-Chief of Applied Corpus Linguistics (ACORP), as the journal celebrates its fifth anniversary. The journal was established in recognition of a major shift: corpus linguistics is no longer confined primarily to general language description or to linguistics alone. Corpus methods are now being adopted, adapted, and extended across diverse areas including forensic linguistics, social policy, food studies, anthropology, writing development, translation and interpreting, and corporate and government communication (Reppen, Goulart, & Biber, 2026). This expansion has created a need for spaces where researchers and practitioners can share case studies, develop methods, theorise practice, and communicate findings accessibly to audiences beyond corpus linguistics. I argue that the “applied” in ACL is not merely a descriptor of topic or method. It is a commitment to using corpus resources, tools, and techniques to address real-world questions: to understand communication in specific settings, to support professional reflection, to improve institutional practice, and to contribute to socially responsive policy. In the context of generative AI, ACL is uniquely positioned to scrutinise machine-generated discourse, evaluate claims about language technologies, support ethical and domain-sensitive uses of AI, and preserve the crucial link between textual patterning and situated human experience.




  • Changpeng Huan (University of Macau)


    https://fah.um.edu.mo/changpeng-huan/

    Plenary Session 4 [September 10, 16:50 - 17:50]

    Automating the resilient subject: a corpus-based discourse study of algorithmic responsibilisation in LLM advice on workplace burnout
    View Abstract

    Workplace burnout research has, from its inception, established a structural diagnosis, yet the remedy remains persistently individual. As large language models (LLMs) are increasingly positioned as first-line advisors for workplace distress, this study examines whether LLM chatbots reproduce this disjuncture when users disclose workplace burnout. Drawing on critical discourse analysis, it examines how LLM chatbots name participants, represent action, and allocate causality across burnout-related conversations from the LMSYS-Chat-1M corpus. The findings reveal a stable discursive script – acknowledge, reframe, prescribe – that systematically activates the worker as the sole agent of remedy while passivating or suppressing organisational actors. Remedies cluster overwhelmingly at the individual level: self-regulation, boundary-setting, and cognitive reframing predominate, while collective action and institutional redress are structurally absent. This ‘algorithmic responsibilisation’ discursively privatises occupational hazards, framing political problems of labour as technical problems of individual psychological management. Yet, this individualising default is not absolute; it is disrupted when users employ explicit legal or collective-action language, suggesting that access to structural advice is mediated by linguistic capital and that the individualising register functions as a contingent default shaped by alignment processes rather than an inherent property of language modelling. By automating the resilience discourse that critical scholars have long identified in neoliberal management, LLMs risk naturalising workplace harm at scale, making available a ‘resigned common sense’ in which collective resistance is discursively foreclosed. The findings are consistent with the interpretation that alignment protocols (RLHF) stabilise this depoliticised register, though the relative contributions of pre-training data and alignment remain to be disentangled.




  • Akira Murakami (University of Birmingham)


    https://www.birmingham.ac.uk/staff/profiles/elal/murakami-akira

    Plenary Session 3 [September 10, 9:20 - 10:20]

    Not all numbers are created equal: Rethinking the statistical modelling of count-based corpus-derived measures
    View Abstract

    Corpus linguistics has benefited substantially from increasingly sophisticated statistical methods. Yet choosing an appropriate model requires more than adopting a method that is widely used or technically advanced. Statistical models should also reflect what we know about the nature and generation of the variables being analysed. This talk argues for greater attention to such structural faithfulness in the statistical modelling of corpus-derived measures. I begin with syntactic complexity measures, many of which are ratios constructed from counts, such as the number of clauses per sentence. Murakami (2025) demonstrates that conventional regression models assuming normally distributed errors can be poorly suited to such measures. They disregard their theoretical bounds and treat observations based on different numbers of linguistic units as equally informative, despite differences in sampling variability. Analyses of learner corpus data and simulations show that this mismatch can produce theoretically impossible prediction intervals, heteroscedasticity, and problematic statistical inference. Alternative approaches that retain information about the underlying counts can mitigate these problems. I then broaden the discussion beyond syntactic complexity. Corpus research routinely employs derived measures whose statistical properties depend on the observations from which they are calculated. Examples include measures of lexical diversity and sophistication and association measures such as mutual information and ΔP. Murakami (2025) notes that related issues may also arise when count-based measures are used as predictors. Drawing on illustrative examples, I will explore the consequences of overlooking such properties and consider possible directions for more appropriate modelling. The broader message is simple: Corpus-derived numbers should not be treated as interchangeable continuous measurements. Understanding how a measure is constructed should be an integral part of deciding how it is statistically modelled. Reference Murakami, A. (2025). Towards more appropriate modelling of linguistic complexity measures: Beyond traditional regression models. Research Methods in Applied Linguistics, 4(1), 100182. https://doi.org/10.1016/j.rmal.2025.100182




  • Nicoletta Calzolari Zamorani (Institute of Computational Linguistics "A. Zampolli", CNR, Pisa)


    https://sites.google.com/view/nicoletta-calzolari/home

    Plenary Session 1 [September 9, 9:50 - 10:50]

    The AI of today is built on the data of yesterday: from Language Resources … to Generative AI
    View Abstract

    Modern Artificial Intelligence breakthroughs - from robust Machine Translation to Large Language Models (LLMs) - do not exist in a vacuum; they stand on the shoulders of over fifty years of foundational, infrastructural work in Computational Linguistics. This talk traces the genealogy of language AI, analyzing how the transition from explicit, hand-crafted structures to implicit, neural representations trained at scale was made possible by the theoretical and empirical foundations of Language Resources (LR) development. We revisit pivotal milestones, such as the 1980s revolution in automatic acquisition of lexical information from Machine-Readable Dictionaries (MRDs) and the landmark 1986 Grosseto Workshop. It was here that the term "Language Resources" was officially coined and a manifesto was produced, defining LRs as an essential "infrastructural" utility for technological progress - the digital equivalent of roads, aqueducts, and electricity. Through strategic standardisation initiatives like EAGLES, ISLE and ISO, our community created the "DNA of interoperability" that allows massive datasets to be merged and scaled for training today’s AI. The current "explosion of AI" is the direct payoff of these decades of data intensive international (many EC) projects and strategic initiatives, characterised by an underlying global strategy paying special attention not only to scientific issues but also to essential methodological, policy and infrastructural dimensions around the notion of language. Among these: reusability, strategic sharing, interoperability, openness, to mention a few. We argue that true advancements happen with (very BIG) data, but only because these data were painstakingly engineered to be computationally tractable. Finally, we address the next frontier: for example, reshaping our infrastructure for the Generative AI era to combat the digital extinction of under-resourced languages and address the critical constraints of AI ethics and data integrity. We conclude that the complexity of modern NLP tasks demands a "broad and open" vision rooted in global teamwork and interdisciplinary collaboration.




TOP