Open-access AN INTERVIEW WITH MARILISA SHIMAZUMI: LEARNER CORPORA AND THEIR INTERFACES WITH LANGUAGE TEACHING

UMA ENTREVISTA COM MARILISA SHIMAZUMI: CORPORA DE APRENDIZES E SUAS INTERFACES COM ENSINO DE LíNGUAS

Abstract

In this interview, Dr. Marilisa Shimazumi provides a glimpse into her academic career, which has been closely associated with the use of learner corpora and English language teaching since the 1980s. The interview begins by discussing her training as a researcher, covering her experiences in national and international research projects. The interview also provides information about her current research projects, applying Multidimensional Analysis to learner corpora and the production of language models. Her studies also investigate differences between human and LLMS-generated texts. In her final assessment of the field, she values interdisciplinary collaborations and sees a promising expansion of the field in Brazil.

Keywords:
Learner Corpora; Multidimensional Approach; Artificial Intelligence; Corpora in Language Teaching; Academic writing.

Resumo

Nesta entrevista, Dra. Marilisa Shimazumi reflete sobre sua trajetória acadêmica ligada ao ensino de inglês e ao uso de corpora de aprendizes desde os anos 1980. A entrevista se inicia discutindo sua formação como pesquisadora, passando pelas suas experiências em projetos de pesquisa nacionais e internacionais. A entrevista também traz informações sobre seus projetos de pesquisa atuais, aplicando a Análise Multidimensional a corpora de aprendizes e produção de modelos de linguagem. Seus estudos também investigam diferenças entre textos humanos e gerados por LLMS. Em sua reflexão, a pesquisadora valoriza colaborações interdisciplinares e prevê expansão promissora da área no Brasil.

Palavras-chave:
Corpora de Aprendiz; Abordagem Multidimensional; Inteligência Artificial; Corpora no Ensino de Línguas; Escrita Acadêmica.

Introduction

Prof.a Dr.a Marilisa Shimazumi has been an important reference in Brazil. She is currently a project coordinator a lecturer at Faculdade Tecnológica de São Paulo (FATEC SP, CPS). Dr. Shimazumi received her PhD in Applied Linguistics and Language Studies from LAEL/PUCSP in 2017 under the supervision of Maria Antonieta Alba Celani, an essential reference in Applied Linguistics studies, analysing the role of mentoring in teacher training, and a master’s degree in Language Teaching and Learning from the University of Liverpool, under the supervision of Geoffrey R. Thompson, an influential researcher in the field of Systemic-Functional Linguistics, examining transitivity in dialogues between doctors and their patients.3 She has actively published in the field and participated in several academic events.

In this interview, she discusses her research in the area of corpus and learner corpora since the 1980s. Her discussion opens with the beginning of her career at CEPRIL/PUCSP, through her master’s degree in Liverpool, where she came into contact with pioneers in corpus linguistics and began to explore data from learner’s corpora. Since then, she has participated in international projects, such as ICLE, in which she collaborated in the compilation of the Brazilian subcorpus (BRICLE), always focusing on academic texts by English learners. She is currently involved in some funded projects, applying Multidimensional Analysis. Her studies also investigate differences between human and AI-generated texts, highlighting the limitations of automatic generation in educational contexts. She also uses corpora in university teaching, encouraging students to build and analyse linguistic data. In her interview, Dr. Shimazumi values interdisciplinary collaborations and observes the expansion of the field in Brazil.

This interview was conducted by s. I would like to thank Prof. Dr. Marilisa Shimazumi for the opportunity and for her availability, kindly agreeing to take part in this volume.

1. Please describe some of your research history. How leaner corpora came into the picture?

I’ve always been interested in the teaching and learning of English as a foreign or additional language, especially when it comes to learner writing. Learner corpora came into the picture quite early for me. I’m talking mid-1980s here, and at that point, learner corpora didn’t really exist, actually, corpora in general weren’t available yet! But as corpus linguistics developed and learner corpora became a reality, I saw an opportunity to explore learner text data, which had always fascinated me. My interest in language teaching, and, by extension, in learner speech and writing, really started under the mentorship of Professor Antonieta Celani4 at PUC-SP in the 1980s, as an undergraduate student and then as a research assistant (at CEPRIL).5 Then, in the early 1990s, I moved to Liverpool to pursue an MA in Applied English Language Studies at the University of Liverpool. I was lucky to be taught by professors like Michael Hoey,6 Geoff Thompson,7 and Mike Scott8. Around that time, I also got to attend seminars and conferences led by key figures in corpus linguistics like John Sinclair9 and the COBUILD team,10 including Susan Thompson and Gil Francis, as well as other linguists like Mike McCarthy and Michael Coulthard.11 That was really my first deep dive into corpus linguistics, and it sparked my long-term interest in learner corpora. While working on my MA thesis, I got access to the APU archive,12: a paper-based collection of British schoolchildren’s writing. I converted it into an electronic corpus so I could analyse it using the corpus tools that were just starting to be developed. Mike Scott, who lectured at Liverpool, for example, had already created Concord with Tim Johns for Oxford University Press, and he was beginning to develop WordSmith Tools.13 I used these tools to do descriptive linguistic analyses of what became the APU corpus, and I presented my findings at seminars and research meetings both at Liverpool and elsewhere.

Later on, in the early 2000s, I moved to Flagstaff while my husband was doing a post-doc at Northern Arizona University. That gave me the chance to meet Professors Douglas Biber14 and Randi Reppen.15 Looking back, I feel really fortunate to have met, learned from, and worked with so many inspiring researchers for such a long time.

2. Could you briefly describe your research projects related to learner corpora?

I’ve been associated with large research projects involving learner corpora for several years now. One recent project is the FAPESP Temático16 called Multidimensional Analysis of Language, Discourse and Society. The idea behind this project is to develop and expand Multidimensional Analysis17 as a comprehensive approach to studying register variation, discourse, and multimodality, across different languages and domains. It brings together an international network of researchers and is organized around three main strands: methodological innovation, register description, and discourse analysis of socially relevant issues.

Now, Multidimensional Analysis hasn’t often been applied to learner corpora, and one of the goals of this project is to start addressing that gap. In the project, I’ve been working on an MD analysis of the International Corpus of Learner English, or ICLE.18 I mention this corpus in more detail in my response to question 3, but just briefly: our analysis revealed a lot of variation in how learners from different backgrounds construct argumentative writing. We identified two main dimensions. The first contrasts objective, fact-based exposition with more subjective, involved expression. The second centers on evaluative, opinion-based exposition. We used these dimensions to compare essays across different national groups.

Another project under the same Temático is one I’ve just started working on; it involves developing an AI-generated corpus of learner writing that mirrors the ICLE. The idea is to simulate learner writing using AI under the same kinds of conditions represented in the ICLE: different national backgrounds, age groups, topics, levels of English instruction, and so on. As every teacher knows, there’s growing concern that students are increasingly turning to AI to write their school essays. So what we want to do is compare the human-written essays in ICLE to AI-generated ones, using MD Analysis, and see what shifts (and doesn’t shift!) linguistically when the writer is an AI rather than a human. We’re also trying to measure how well we can tell the difference between the two based on the underlying dimensions of variation.

Still under the Temático, I’ve also been involved in a research investigation looking into how well AI can generate EFL coursebook texts. The question we wanted to explore was whether generative AI could effectively replace human authors in producing the kinds of texts typically found in English as a Foreign Language materials, and more specifically, whether it could generate writing models that would actually be used in textbooks. To do this, we adopted a corpus-based approach. We compiled a corpus of 500 EFL coursebook texts from major publishers, spanning 25 years of materials across 16 different registers, all aimed at B2 and C1 level learners. Then, we applied Multidimensional Analysis to compare these human-authored texts with AI-generated versions. The analysis revealed five major discourse dimensions: persuasive versus analytical expression; interactive and speculative discourse; structured and formulaic composition; narrative and descriptive accounts; and summarized overviews. What we found was that AI-generated texts differ quite a bit from human-written ones. In particular, AI has trouble replicating the register-specific and communicative demands of real educational materials. Often, it either misses key linguistic features or overuses others. So, while the texts might look polished on the surface, they don’t function discursively in the same way as human-authored texts. The conclusion we reached is that, at least for now, generative models can’t really serve as adequate substitutes for experienced textbook writers.

Another project I’ve been working on is called Large Language Model Reasoning Patterns: A Corpus-Based Multidimensional Approach. This one focuses on mapping the reasoning and discourse patterns produced by models like ChatGPT, Gemini, and Llama. Using both traditional and lexical MDA, we’re trying to identify the functional and ideological dimensions embedded in AI-generated texts.

Even though this isn’t a learner corpus project in the traditional sense, the methods we’re using have a lot in common with learner corpus studies. Just like MDA can help uncover communicative and ideological patterns in AI output, it can also be used to identify developmental patterns and discursive gaps in learner writing. And here I should clarify that when I say “learner writing,” I’m referring specifically to texts written by non-native English speakers in academic settings, such as university students writing in English. The idea is that the insights from this project can inform both instructional design and the development of more learner-aware AI tools.

I’m also involved in another project entitled Description and Improvement of Artificial Intelligence Text Generation Using Corpus Linguistics and Multidimensional Analysis. This one investigates the ways in which texts generated by large language models differ from those written by humans, not just in grammar, but in their functional and discourse-level features. Even when LLMs produce grammatically correct output, they often miss the deeper patterns that shape how language is actually used in context. So, in this project, we’re using MD Analysis to systematically describe how human texts behave across registers, and then we’re applying those findings to steer AI generation through Retrieval-Augmented Generation, or RAG. The goal is to help AI produce texts that are not only coherent but also communicatively adequate and closer to natural human usage. And again, this same multidimensional approach can be applied to learner corpora as well. By identifying which discursive and functional features are missing or underdeveloped in learner writing, we can support the creation of more focused instructional materials.

Lastly, I’ve also been working on a project examining variation in academic writing by university students. This one looks at the language used in research articles published between 2013 and 2023 in high-impact journals. The idea is to describe the conventions of academic English and use that information to support student writing, especially for those writing in English as an additional language. We’re using Data-Driven Learning in this project to help students become more autonomous in navigating academic writing conventions. More information is available on the project site: https://www.inpact.net.br.

3. What is the origin of the data and the compilation processes you are using? How do you select the texts included in a corpus: are there criteria based on linguistic characteristics, nationality, gender, register, or other criteria?

Probably the most ambitious project I’ve worked on was the analysis of the full ICLE corpus, which I mentioned earlier. We presented the findings from this project earlier this year at the ‘Register and Task Variation in Learner Corpus Research’ conference (VAR4LCR),19 which was held at the ICLE headquarters in Louvain, Belgium. The ICLE corpus is a large-scale international project, and the data were collected by dedicated research teams around the world. I was part of the Brazilian team back in the late 1990s and early 2000s. We followed the guidelines established by the team in Louvain to ensure consistency across all national contributions.

The full version of ICLE v.3 includes 9,529 essays written by university-level students, learners of English as a foreign language, from 27 different first-language backgrounds and before the advent of AI. The corpus totals over 5.7 million words. It’s important to note here that when we talk about ‘learner writing,’ we’re referring to academic texts written by non-native English speakers, often in higher education settings, and not beginner-level language learners. The compilation followed a strict set of design principles. The selection criteria included sociolinguistic variables like nationality, gender, age, and number of years studying English. Each national subcorpus includes about 20 essays per proficiency level: CEFR B2, C1, and C2, and also includes metadata such as whether the essays were written under exam conditions or as part of coursework.

What makes ICLE particularly valuable is its focus on argumentative writing. That decision to control for register allows us to compare how students from different backgrounds approach the same type of communicative task. It gives us a uniform context to examine variation, across countries and educational systems.

4. How do you handle issues such as metadata, representativeness, and balance in the compilation process?

As I mentioned earlier, I was directly involved in compiling the Brazilian subcorpus of the International Corpus of Learner English, or BRICLE. Like all national components of the ICLE project, it followed a very rigorous set of guidelines developed by the coordination team in Louvain. Each participant had to fill out a detailed profile questionnaire, which helped us collect key metadata, things like first language, age, gender, and the learner’s educational background in English.

Beyond those basic details, the questionnaire also gathered information about the students’ exposure to English, whether they had spent time in English-speaking environments, how many years of formal instruction they’d had, and so on. We also noted practical aspects of the writing task itself: Was the essay done in an exam setting or as homework? Was it timed or untimed? Did the student have access to any reference materials, like dictionaries or notes, while writing?

The design of the corpus was carefully controlled to ensure balance, not just in terms of regional representation, but also in terms of the register of the texts. All participants were asked to write argumentative essays, which gave us a consistent communicative task across all learners. And here again, I want to clarify that when we refer to “learner writing,” we’re talking about texts written by non-native speakers of English in academic settings, not beginner-level language learners.

By controlling for variables like proficiency level and years of English study, we were able to create a resource that supports cross-linguistic comparisons while still preserving the authenticity of real student writing. That level of standardization was needed because it was necessary to make the corpus uniform across so many different teams in different places at different times.

5. Has the use of corpora led to curricular changes in the courses you teach (e.g., new course units, new assessment methods)? What transformations have you observed? Are there indirect applications of leaner corpora (e.g., data used for the development of teaching materials) or direct ones (e.g., exploration by students or pre-service teachers)?

Yes, in our courses, students get hands-on experience with tools like COCA,20 the Corpus of Contemporary American English, developed by Professor Mark Davies, and AntConc,21 developed by Professor Lawrence Anthony. They learn to retrieve, sort, and interpret corpus data, and in some cases, they even build their own mini corpora. Through this work, they explore key concepts like frequency, co-occurrence, and variation, and they get a much better sense of how language works across different registers and contexts. That kind of experience is really valuable because it helps them move beyond intuition and start recognizing actual patterns in authentic language use.

Students are encouraged to use corpus data and notice specific language patterns based on real usage. Learner corpora can provide insights into common usage or even common misuses, and that information can shape how we design exercises or activities that better meet learners’ needs. So even when students aren’t directly analysing corpora themselves, the influence of corpus-based insights is still present in how we approach language teaching.

6. How do you view the role of LLMs in the guided exploration of corpora by students or research?

I think large language models can play a useful role in helping students and researchers explore corpora, both as a tool for learning and as a way to generate data for analysis.

In our own work, we’ve mainly used LLMs to generate text based on carefully crafted prompts. The idea is to compare these AI-generated texts with real texts written by humans, so we can see how closely the model mimics human language, especially when it comes to things like register, stance, and discursive strategies. But of course, we need to approach this with caution. LLMs can introduce all sorts of biases and distortions, and they don’t always produce language that reflects real-world use. So any analysis based on AI output has to be done with a critical eye and a strong verification process.

In the classroom, though, LLMs can be a great support for getting students more engaged with corpus work. For instance, students can use the model to test out hypotheses, make sense of concordance lines, or explore variation by asking it to produce different versions of the same sentence, for example: one formal and one informal. Then they can take those versions and compare them with real examples from a corpus using a tool like AntConc. That kind of activity really helps students notice patterns, become more aware of register differences, and think more critically about how language works in different contexts.

7. Is there integration of undergraduate and graduate research in the projects you are developing?

Yes, all these projects mentioned involve undergraduate and graduate students carrying out their own research projects at different stages and degrees of collaboration.

8. Has your research with learner corpora involved interdisciplinary collaborations (with fields such as computer science, psychology, education, or cultural studies)? If so, what has that experience been like? What are the challenges in interdisciplinary collaborations?

Our current research has really benefited from interdisciplinary collaboration. We’re working with researchers from both computer science and linguistics, in Brazil and internationally. In two of our current projects, for instance, we’ve teamed up with colleagues in computer science who specialize in natural language processing and artificial intelligence. They bring strong technical expertise in computational modelling, and we contribute with our background in corpus linguistics, learner writing, and multidimensional analysis.

It’s been a very productive exchange, we’re learning a lot from each other. These collaborations include colleagues from Northern Arizona University,22 The Hong Kong Polytechnic University,23 Georgia State,24 Arizona State,25 and Lancaster University,26 as well as partners here in Brazil from the Federal Universities of Minas Gerais27 and Rio Grande do Sul,28 the State University of Rio de Janeiro,29 São Paulo State University,30 and São Paulo Technological College,31 among others. It’s been exciting to connect across disciplines and institutions to explore language from different angles.

9. Have you obtained funding for learner corpora projects? Please comment into the process of financing such projects.

Yes, we have secured funding for our projects through Brazilian agencies, and an international colleague has obtained matching funds from their national research council. This financial support has been fundamental to the success and continuity of our research. In Brazil, we are especially grateful to FAPESP,32 CNPq,3334and CAPES, which have provided us with backing through grants and scholarships.

10. How do you evaluate the research on learner corpora in Brazil? How is the road ahead?

I think the research on learner corpora in Brazil is on very solid ground. We have a large pool of researchers across the country, many of whom come from a background in corpus linguistics, so the expertise is definitely there. We also have strong research centres dedicated to corpus work in several regions, which gives us a solid infrastructure to build on.

What’s also important is the sheer number of language teachers and learners we have in Brazil. Many of these teachers go on to graduate school and pursue MAs and PhDs, and they often bring valuable classroom insights into their research. So we have both the knowledge and the people to move learner corpus research forward.

That said, there’s still a lot more we can do. The potential is definitely there! We just need to keep building on what we already have. That’s why initiatives such as this one you are leading are so important because they bring learner corpus research to the fore, hopefully inspiring more people to take it on.

  • 3
    It was thanks to Dr. Shimazumi’s master’s that I had the opportunity to learn about her work in the mid-1990s, when I was working towards my master’s degree at PUCSP. Her approach and analysis of English was influential and inspiring for me in the development of a research methodology for Portuguese transitivity choices, a field that was still relatively unexplored at the time.
  • 4
    Maria Antonieta Alba Celani (1923-2018) was a linguist and one of the pioneers in the field of Applied Linguistics in Brazil. Celani was an emeritus professor at the Pontifical Catholic University of São Paulo (PUCSP), where she founded LAEL, the first graduate programme in Applied Linguistics in Latin America. From the 1970s through the early 1990s, she was responsible for the Projeto Nacional Inglês Instrumental nas Universidades Brasileiras (or simply projeto), whose influential impact on language teaching in Brazil can still be seen today. This project was also the pioneer at the introduction of ESP (Languages for Specific Purposes) in Brazil (Celani, 2005).
  • 5
    CEPRIL, Centro de Pesquisa, Recursos e Informação em Linguagem (previously Leitura), is a research centre at LAEL-PUCSP. It started as a centre for research, reading materials (and other ESP teaching aids), and a library specialised in Applied Linguistics in the late 1970s.
  • 6
    Michael Hoey (1948-2021) was an influential applied linguist with a seminal work in corpus linguistics, including the Lexical Priming theory (Hoey, 2012).
  • 7
    Geoff Thompson (1947-2015) was an Honorary Senior Fellow of the University of Liverpool and a very fertile researcher in Systemic-Functional Linguistics. Some of his articles and books are still very influential in the field (Hunston, 2016).
  • 8
    Mike Scott is a professor and author of WordSmith Tools. He is currently a professor at Aston University and worked at PUCSP during the 1980s in the Projeto Inglês Instrumental (along with Dr. Celani and Dr. John Holmes).
  • 9
    John Sinclair (1933-2007) was a pioneer researcher at Corpus linguistics, discourse analysis, lexicography, and language teaching from UK. He worked at the Birmingham University and was responsible for the COBUILD, one of the first electronic corpus research projects (Scott, 2007).
  • 10
    COBUILD (Collins Birmingham University International Language Database) was one of the first electronic corpus. It was implemented and coordinated by John Sinclair.
  • 11
    Michael Coulthard is a researcher in the fields of forensic and corpus linguistics.
  • 12
  • 13
    WordSmith Tools is a software for corpus analysis by Michael Scott since 1996. It is available at: https://www.lexically.net/wordsmith (Scott, 2025).Available at https://lexically.net/wordsmith/.
  • 14
    Douglas Biber is a professor at Northern Arizona University, he has developed a relevant work on Multidimensional Corpus Analysis (Biber, 1988).
  • 15
    Randi Reppen is a Professor Emerita of Applied Linguistics and TESL at Northern Arizona University. She and Dr. Biber have developed important work in corpus linguistics.
  • 16
    Projeto FAPESP Temático: Multi-Dimensional Analysis of Language, Discourse, and Society Grant # 2022/05848-7.
    Abstract :Contemporary society is marked by expanding communication formats and increasing access to technologies that process large volumes of linguistic, visual, and audio data. Yet, linguistic research often relies on limited datasets. This project addresses that disconnect by advancing Multidimensional Analysis (MDA) to cover register variation, discourse, and multimodality. The initiative involves researchers in Brazil and abroad and is organized around three main areas: (1) development of new methods and tools for MDA; (2) description of register variation across languages and domains; and (3) identification of discourses on pressing social topics. Subprojects include creating a computational tool, refining corpus design criteria, improving MDA methods, and applying them to domains such as social media, news, television, music, video games, literature, education, and politics in both English and Portuguese. Key topics include climate change, pandemics, misinformation, prejudice, and religion.
  • 17
    The Multidimensional Analysis (MD) is a methodology developed by Biber (1988) which allows to analyse different language features simultaneously and whole texts as a general context of variation (Berber Sardinha, 2000). In statistical terms, MD uses factor analysis to group multiple linguistic variables in dimensions defined by the probability of such language features to co-occur (or not). Such grouping might be interpreted in functional, grammatical and discursive ways (Brezina, 2018; Lima-Lopes, 2020).
  • 18
    “The International Corpus of Learner English (ICLE) is a corpus of essay writing by upper intermediate and advanced learners. Founded and coordinated by Sylviane Granger of the University of Louvain” (ICLE Université Catholique de Louvain, [s. d.]). Further corpus info is available at https://corpora.uclouvain.be/cecl/icle/home.
  • 19
  • 20
  • 21
    AntConc is a corpus software developed by Laurence Anthony and available at https://www.laurenceanthony.net/software/antconc/.
  • 22
    Available at https://nau.edu/.
  • 23
  • 24
    Available at https://www.gsu.edu/.
  • 25
    Available at https://www.asu.edu/.
  • 26
  • 27
    Available at https://ufmg.br/.
  • 28
    Available at https://www.ufrgs.br.
  • 29
    Available at https://www.uerj.br/.
  • 30
  • 31
  • 32
    Projeto Temático FAPESP: PROCESSO 2022/05848-7.
  • 33
    Projeto CNPq Edital 16/2024: PROCESSO 403130/2024-7. CNPq 16/2024 | Grant # 403130/2024-7. Description and Improvement of Artificial Intelligence Text Generation Using Corpus Linguistics and Multidimensional Analysis. Abstract: This project investigates the problem that texts produced by Large Language Models (LLMs) fail to exhibit human characteristics, both in how they function and in the discourses they convey. Although LLM outputs may appear human-like when judged in isolation, clear differences emerge when compared with real texts from specific text types such as conversations, research articles, essays, news reports, and song lyrics. These differences suggest that current models do not adequately represent how people actually use language. Existing research has only covered a few text types, leaving a gap in understanding how LLMs differ from human language use across a wider range of contexts. This project seeks to expand the description of variation in texts using a multidimensional approach and use those results to inform how LLMs produce language. The proposal is based on the idea that LLMs currently focus too much on sentence-level coherence, which is insufficient for capturing variation in text types and discourses. To address this, the project proposes guiding LLMs using Retrieval-Augmented Generation (RAG), with information about functional and discursive features of human texts. The effectiveness of this approach will be evaluated by comparing AI-generated texts to human ones using multidimensional parameters. The work will be carried out by a team of researchers and students from Brazil, Canada, the United States, and China with experience in Corpus Linguistics, Multidimensional Analysis, and AI.
  • 34
    Projeto CNPq Edital 22/2024: PROCESSO 444019/2024-3. CNPq 22/2024 | Grant # 444019/2024-3
    Large Language Model Reasoning Patterns: A Corpus-Based Multidimensional Approach. Abstract: This project has two goals. First, it seeks to connect Brazilian and international researchers to describe how Large Language Models (LLMs) generate language, focusing on their functional and discursive reasoning. Second, it aims to develop corpus-based multidimensional models to capture how LLMs such as ChatGPT, Gemini, and Llama produce knowledge. ‘Reasoning’ here refers to how LLMs construct responses. The project uses two types of Multidimensional Analysis (MDA): one to detect communicative functions, the other to identify ideological patterns. These methods reveal recurring combinations of linguistic features in AI-generated texts. Two obstacles complicate the analysis: the internal processes of LLMs cannot be directly observed, and their responses vary even when prompted the same way. MDA addresses this by detecting regular patterns in language use that reflect these hidden processes. Through this approach, the project aims to describe the recurring reasoning patterns behind AI-generated texts.
  • CNPq - Conselho Nacional de Desenvolvimento Científico e Tecnológico - research fellow (grant #303996/2024-2).

RESEARCH DATA AVAILABILITY STATEMENT

This is an original and never published interview. The datasets quoted by the main author, Dr. Shimazumi, had their link made available whenever possible.

References

  • BERBER SARDINHA, Tony. Análise multidimensional. DELTA: Documentação de Estudos em Lingüística Teórica e Aplicada, [s. l.], vol. 16, p. 99-127, 2000. https://doi.org/10.1590/S0102-44502000000100005 Accessed on: 12 Aug. 2025.
    » https://doi.org/10.1590/S0102-44502000000100005
  • BIBER, Douglas. Variation across speech and writing Cambridge: Cambridge University Press, 1988.
  • BREZINA, Vaclav. Statistics in corpus linguistics: A practical guide Cambridge/New York: Cambridge University Press, 2018.
  • CELANI, Maria Antonieta Alba. ESP in Brazil: 25 years of evolution and reflection Sao paulo: EDUC, 2005.
  • HOEY, Michael. Lexical Priming: A New Theory of Words and Language Hoboken: Taylor and Francis, 2012.
  • HUNSTON, Susan. An inspiring advocate for Systemic-Functional Linguistics: Geoff Thompson (1947-2015). Functions of Language, [s. l.], vol. 23, no. 1, p. 1-8, Jun. 2016. https://doi.org/10.1075/fol.23.1.001hun Accessed on: 18 Aug. 2025.
    » https://doi.org/10.1075/fol.23.1.001hun
  • ICLE UNIVERSITé CATHOLIQUE DE LOUVAIN. [S. l.]: https://www.uclouvain.be/en/research-institutes/ilc/cecl/icle, [s. d.]. Accessed on: 12 Aug. 2025.
    » https://www.uclouvain.be/en/research-institutes/ilc/cecl/icle
  • LIMA-LOPES, Rodrigo Esteves de. Immigration and the Context of Brexit: Collocate network and Multidimensional Frameworks Applied to Appraisal in SFL. Muitas Vozes, [s. l.], vol. 9, no. 1, p. 410-441, 2020. https://doi.org/10.5212/MuitasVozes.v.9i1.0024 Accessed on: 11 Mar. 2021.
    » https://doi.org/10.5212/MuitasVozes.v.9i1.0024
  • SCOTT, Mike. John Sinclair (1933-2007). Language Pioneer and Explorer. Language Awareness, [s. l.], vol. 16, no. 2, p. 79-80, May 2007. https://doi.org/10.1080/09658410708668857 Accessed on: 18 Aug. 2025.
    » https://doi.org/10.1080/09658410708668857
  • SCOTT, Mike. WordSmith Tools 9.0. [s. l.], 2025. https://doi.org/https://lexically.net/wordsmith/index.html Accessed on: 18 May 2019.
    » https://doi.org/https://lexically.net/wordsmith/index.html
  • Editora-chefe:
    Simone Tiemi Rashiguti
  • Editora-convidada:
    Silvia Araujo
  • Editora-convidada:
    Paula Tavares Pinto
  • Editora-convidada:
    Renata Palumbo
  • Editor-convidado:
    Rodrigo Esteves de Lima-Lopes

Publication Dates

  • Publication in this collection
    09 Jan 2026
  • Date of issue
    2025

History

  • Received
    01 Sept 2025
  • Accepted
    11 Sept 2025
  • Published
    26 Sept 2025
location_on
UNICAMP. Programa de Pós-Graduação em Linguística Aplicada do Instituto de Estudos da Linguagem (IEL) Rua Sérgio Buarque de Holanda, n.571 - Cidade Universitária - CEP: 13083-859, Telefone: (+55) 19 - 3521-6729 - Campinas - SP - Brazil
E-mail: tla@unicamp.br
rss_feed Acompanhe os números deste periódico no seu leitor de RSS
Ir para o topo Reportar erro