Linguistic Data Analysis II

Session 1: Introduction to LLMs as Tools for Linguistic Data Analysis

Day 1 · Introduction & First Experience

Masaki EGUCHI, Ph.D.

Tohoku University · Summer 2026

📋 Agenda

  1. Why this course
  2. Self-introductions
  3. Syllabus — the 5-day map & objectives
  4. What an LLM is
  5. Detour — History of embeddings
  6. What this buys you, and where it fails
  7. Discussion

1 · Why this course

LLMs are everywhere now

Since OpenAI released ChatGPT in late 2022, Large Language Models (LLMs) have now become common tools — for searching, translating, coding, and more.

Wang, D., & Zhang, S. (2024). Large language models in medical and healthcare fields: Applications, advances, and challenges. Artificial Intelligence Review, 57(11), 299. https://doi.org/10.1007/s10462-024-10921-0

LLMs have already changed how science is written

Kobak et al. (2025) tracked word frequencies across 15 million PubMed abstracts (2010–2024).

The authors estimate that at least 13.5% of 2024 abstracts were written with help from an LLM — over 40% in some journals and countries.

How to look at the charts:

  • Top row: words that became far more frequent right after ChatGPT.
  • Bottom row: for comparison, words that rose because of real events (COVID, Ebola) or a new method.

Kobak, D., González-Márquez, R., Horvát, E.-Á., & Lause, J. (2025). Delving into LLM-assisted writing in biomedical publications through excess vocabulary. Science Advances, 11(27), eadt3813. https://doi.org/10.1126/sciadv.adt3813

Emerging research using LLM as annotator

Already tons of research uses LLMs in Applied Linguistics:

Study What it annotated
Kim & Lu (2024), JEAP rhetorical moves in article abstracts
Mizumoto (2025), SSLA L2 errorsr = .94 with human raters, F1 = .95
Yamashita (2024), RMAL essay scores

So the question isn’t whether linguists will use them. It’s how to use them responsibly for linguistic analysis.

Linguistic annotation: promising but costly

One especially promising use is annotation — labelling language data for analysis.

  • But careful annotation is expensive: expert time, training, guidelines, double-coding.
  • If an LLM could annotate reliably, then it facilitates research.

So the question arises: can an LLM do this annotation for us?

One simple Goal of the course

Question you might have: Can we trust an LLM on annotating this?

Until today

  • Well, maybe?

OR

  • Are there any paper on this?

After this course

  • Let’s find out!

We will learn methodology to investigate LLM’s capability for lingusitic data analysis.

2 · Self-introductions

Your instructor

  • Masaki Eguchi (江口 政貴)

Position

  • Research Scientist at Equmenopolis, Inc.
  • Research Associate Processor, Perceptual Computing Lab, Waseda University

Education

  • Ph.D. in Linguistics, University of Oregon.
  • M.A. in Education, Waseda University

My research

How can computer technologies (AI, Natual Language Processing) help research and practice of language assessment?

  • Large Language Models
  • Dialog System
  • Interactional Competence

My research

Introduce yourself - Round 1

Tell us:

  • your name
  • What you enjoy doing while you are not doing research

Round 2

Tell us:

  • What is your research interest
  • Why you decided to take this course — any goals, any concerns

3 · Syllabus

Course objectives

By the end of this course, you will be able to:

Area Objective
Critical appraisal Explain what LLMs can and cannot do for linguistic analysis, and judge when an LLM-based approach is appropriate.
Annotation scheme Adapt and operationalize an existing annotation scheme and its coding guidelines.
Gold-standard datasets Build a gold-standard dataset, including assessing inter-annotator agreement.
Prompt design Design, tune, and document prompts that elicit reliable annotations from an LLM.
Evaluation Evaluate model performance with precision, recall, F1, and confusion matrices, and interpret the results critically.
Reproducibility Report methods and findings transparently and reproducibly.

What we’ll do this week

Day Focus
1 What is LLM, Python basics, call an LLM
2 Gold standard + agreement — how to define and measure quality
3 Prompt + iterate against metrics - how to use LLM
4 Method + assemble the pipeline
5 Your own data + present

Assignments & grading

Component Percent
Attendance / participation 20%
Hands-on activities (Day 1–3) + completed notebook 40%
Mini-project group presentation + Q&A (Day 5) 20%
In-class two-page report (individual) (TBA) 20%

Note

Full details on the course website: Syllabus · Readings

Session 1: Introduction to LLMs

🎯 Learning Objectives

By the end of this session, you will be able to:

  • Describe, at a high level, what a Large Language Model is and how it is trained.
  • Explain what an embedding is, and how a model comes to represent a word’s meaning in context.
  • Say what an LLM has already been used for in linguistics, and where it has been shown to struggle.

The one idea for today

An LLM produces annotations, and you must check them against a gold standard before you report them. It can be applied to the linguistic tasks you already work on.

Everything this week is a method for doing that checking. Today we explain why the checking is necessary.

4 · What is an LLM — and why it’s relevant

Question

  • Have you ever used LLMs?
    • What was it for? / Did it work?
    • What surprised you?

What do you know about LLMs?

  • What is a Large Language Model (or LLM)?

    • Share with the neighbor anything you know about LLM!

First, a 7-minute overview

Grant Sanderson (3Blue1Brown), Large Language Models explained briefly.

While you watch, write down an answer to each of these:

  1. What exactly does the model predict?
  2. Where do the numbers inside the model come from?
  3. What does attention do?

We come back to all three afterwards, with a language example.

Reflect on what the video says…

  1. What exactly does the model predict?
  2. Where do the numbers inside the model come from?
  3. What does attention do?

5 · Detour — Brief history of embedding

Before LLM — Brief history of embedding

LLM encodes language into numbers — i.e., embedding.

Embedding: “vector representations of the meaning of words that are learned directly from word distributions in texts” (Jurafsky & Martin, 2026, p. 96).

Two families, and we build up to both:

  • Static — one fixed vector per word: counting contexts, then Word2Vec.
  • Contextual — a different vector each time the word appears: BERT, and the models you use.

RGB color space

A vector is a list of numbers that fixes the location of something in a space. A colour, for example, is fixed by how much red, green and blue it has.

Move the sliders to change the colour. Drag the cube to rotate it.

The dashed path is the vector: R along red, then G along green, then B along blue. The number under each name is the distance from your colour to that colour — close in the space = similar colour. Word embeddings work the same way, with hundreds of numbers instead of three.

“You shall know a word by the company it keeps!” (Firth, 1957, p. 11)

Q: How can a machine learn meaning?

A: The answer is by studying Co-occurrence.

Left context Node Right context
is traditionally followed by cherry pie, a traditional dessert
often mixed, such as strawberry rhubarb pie. Apple pie
computer peripherals and personal digital assistants. These devices usually
a computer. This includes information available on the internet

Count, by hand: how often does each of a, computer and pie appear in each of those four windows?

Study the company to understand meaning!

The table on the right counts co-occurence frequency.

Left context Node Right context
is traditionally followed by cherry pie, a traditional dessert
often mixed, such as strawberry rhubarb pie. Apple pie
computer peripherals and personal digital assistants. These devices usually
a computer. This includes information available on the internet
a computer pie
cherry 1 0 1
strawberry 0 0 2
digital 0 1 0
information 1 1 0

Already suggests some topical clustering. Each row is that word’s vector: cherry = [1, 0, 1].

The co-occurence frequency in a real corpus

Now count every occurrence in Wikipedia.

Six of the columns becomes (Jurafsky & Martin, 2026, p. 103):

aardvark computer data result pie sugar
cherry 0 2 8 9 442 25
strawberry 0 0 0 1 60 19
digital 0 1670 1683 85 5 4
information 0 3325 3982 378 5 13

cherry and strawberry have similar rows. digital and information have similar rows.

The distributional principle

This clustering emerges without anyone writing a dictionary. Again “You shall know a word by the company it keeps!” (Firth, 1957, p. 11).

The table is too big and a lot zeros

Look at the aardvark column on the previous slide. All four words score 0. None of them ever occurs near aardvark.

aardvark computer data result pie sugar
cherry 0 2 8 9 442 25
strawberry 0 0 0 1 60 19
digital 0 1670 1683 85 5 4
information 0 3325 3982 378 5 13

Co-occurrence table is sparse, about 99 in every 100 columns are 0.

Finding the latent dimensions

We can apply factor analysis to a co-occurrence table to reduce meaning.

Factor analysis

A statistical method for uncovering the latent structure of data:

  • We give a 6-item questionnaire on enjoyment and boredom.
  • The questionnaire items are the observed variables.
  • We find latent variables that predict the values of the observed variables.
ENJ1 · look forward ENJ2 · enjoy it ENJ3 · time passes fast BOR1 · mind wanders BOR2 · lose interest BOR3 · feels slow Enjoyment Boredom
Each arrow is a loading. Solid = strong, dashed = near zero. Values are illustrative.

. . .

item Enjoy Bored
ENJ1 · I look forward to class 0.82 −0.11
ENJ2 · I enjoy the activities 0.79 −0.08
ENJ3 · Class time passes quickly 0.74 −0.15
BOR1 · My mind wanders in class −0.09 0.80
BOR2 · I lose interest quickly −0.13 0.85
BOR3 · Class feels slow −0.10 0.77

Six items, two latent variables. Each item loads high on one and near zero on the other.

Latent Semantic Analysis

Treating the words in corpus like observed variables in factor analysis, we explore underlying dimensions of meaning.

computer data result pie sugar dimension 1 dimension 2
The same diagram as the questionnaire. Context words are the observed variables; the dimensions are what factor analysis extracts.

By applying dimensional reduction, a word will associate with “latent” dimension, which explains some underlying semantic distribution in corpus.

dimension 1 dimension 2
computer 0.651 −0.009
data 0.756 0.004
result 0.066 0.019
pie 0.001 0.998
sugar 0.002 0.061

Nobody assigned these weights — they were derived from the corpus counts. This technique is called Latent Semantic Analysis (LSA) (Landauer, 1999).

word2vec: We actually do not need a table

  • LSA support distributional semantics.

BUT LSA has some limitations:

  • LSA is usually applied to document-level co-occurrence, but a document can have multiple topics.
  • The table is huge and difficult to use large corpus.

→ We want to model the dimension without huge co-occurrence table and focusing on local co-occurence.

word2vec: Learns whether a word pair co-occur or not

Take a target word and one word near it. Did this pair really occur together?

...  your  [ account   is   free   of   charge ]  a  standard  fee ...
                c1     c2    w     c3     c4

The corpus already knows the answer, so no one has to label anything.

word2vec: Vector representation

Learn the vector representation that predicts real co-occurence in corpus as YES and fake ones as NO.

a pair that really occurred free charge the target a word next to it target vector from matrix W context vector from matrix C Do they co-occur? answer YES → pull them closer
a pair drawn at random free Tolstoy the same target any word from the corpus target vector the same row of W context vector from matrix C Do they co-occur? answer NO → push them apart
  • Two matrices: W holds a vector for every word as a target, C holds one for every word as context.
  • Word2Vec learns this through thousands of trials.

  • We then often only keep W (while sometimes use both W and C).

Transformer — The breakthrough

In 2018, the paper “Attention is All You Need” published. The introduction of self-attention changed everything.

Attention Is All You Need

Self-Attention is the key

Self-attention builds the vector for bank on the fly, using the contexts (Jurafsky & Martin, 2026, p. 178).

I walked along the pond and noticed the trees along bank the word being built these two get most of the weight one vector for “bank” mostly pond and trees I deposited the cheque at the bank a different vector for “bank” mostly deposited and cheque
Arrow thickness is the weight the model gives that neighbour. Weights are illustrative.
  • Query — sent by bank, the word being built: which words here tell me what I mean?
  • Key — every other word’s answer. pond matches the query closely, the and walked do not, so pond gets a large weight.
  • Value — what each word contributes to the total.
  • Vector — the values added up, each multiplied by its weight. In the second sentence the same query meets different keys, so deposited and cheque take the weight and the sum comes out different.

Where does Attention come from?

Every word turns itself into three separate vectors. bank uses its q; every other word offers its own k for the comparison and its own v for the sum.

I walked along the pond … bank I walked pond bank kI kwalked kpond vI vwalked vpond qbank each word makes its own k and its own v score = q · k weight = softmax weight × v 0.4 1.1 6.2 0.02 0.04 0.71 … all weights add up to 1 0.02·vI 0.04·vwalked 0.71·vpond + … = the vector for “bank” here I deposited the cheque at the bank I at cheque bank kI kat kcheque vI vat vcheque qbank the same word, two different vectors score = q · k weight = softmax weight × v 0.5 0.9 5.8 0.03 0.04 0.66 … all weights add up to 1 0.03·vI 0.04·vat 0.66·vcheque + … = a different vector for “bank”
Scores and weights are illustrative. Only three context words are shown; the real sum runs over every word.

Transformer make full use of attention

  • BERT-base models 12 layers of 12 attention heads.
  • GPT-3 had 96 layer × 96 heads = 9216 attention heads

Transformer discovers Noun phrase

Transformer discovers coreference

Clark, K., Khandelwal, U., Levy, O., & Manning, C. D. (2019). What Does BERT Look At? An Analysis of BERT’s Attention (arXiv:1906.04341). arXiv. https://doi.org/10.48550/arXiv.1906.04341

6 · What this buys you, and where it fails

What an LLM lets you do

  1. It can label categories that depend on context. Most of the categories linguists work with are of that kind, and a bag-of-words model could never represent them.

  2. One model, many tasks. Generative model can do many things with a prompt. We do not train a new model for each task with a new annotated training set.

  3. Quick iteration: Annotation is painstaking, but you can iterate with LLM early.

Sounding right is not being right

An LLM essentially learns to predict next plausible word.

Plausible is not the same as correct — which is exactly what Curry et al. found when the output read well and the analysis was wrong.

That’s why we still need NLP evaluation.

What people have already done with LLMs

Historically each task had its own model — a POS tagger, an NER model, a parser.

Already tons of research uses LLMs:

Study What it annotated Track you could pick
Kim & Lu (2024), JEAP rhetorical moves in article abstracts RAAMove · CaRS-50
Mizumoto (2025), SSLA L2 errorsr = .94 with human raters, F1 = .95 AutoErrorAnalyzer
Yamashita (2024), RMAL essay scores ICNALE GRA

Three of these are on your mini-project list. You can pick one as a group to work on conceptual replication.

What this course covers

Important

This course is the evaluation half of NLP, applied to LLMs.

We do not cover the training part of NLP. The evaluation part provides exactly the rigour an LLM’s outputs need, and that is what you will learn to run this week.

  • inter-annotator agreement (percent + Cohen’s κ)
  • precision / recall / F1, confusion matrix

7 · Discussion

Your turn

  • What was new information to you about LLM?

  • Given the lecture, are you optimistic or skeptical about the use of LLM? Why do you think so?

Recap · The 5-day map

What we’ll do this week

Day Focus
1 What is LLM, Python basics, call an LLM
2 Gold standard + agreement — how to measure quality
3 Prompt + iterate against metrics
4 Method + assemble the pipeline
5 Your own data + present

Any questions?