Session 11: Plenary Pipeline Assembly
Day 4 · Methodology & Pipeline Assembly (4-2)
Warning🚧 Being prepared
Learning objectives
By the end of this session you will be able to:
- Describe the whole pipeline as a chain of inputs and outputs — what each step consumes and what it produces.
- Point to where in Days 1–4 you already ran each call, and use the same call form again.
- Write your group’s
PLAN.md: track, label set, sampling settings, the dev/test ratio, who annotates what, which agreement statistics you owe and which one number you will lead with — and the one prompt change you predict will help.
Agenda
- The pipeline, in one diagram (~15 min) — pool → sample → sheet → adjudicated gold → dev / test → prompt rounds on dev → the held-out run → report, with each arrow’s data shape named, and the 04/05 file boundary that keeps the test half closed. Plus the inventory and its ancestor column: every call has one, in Days 1–4.
- Write
PLAN.md(~45 min) — in your groups. The instructor comes round, reads it, and signs it off. - End-state check (~25 min) — each group states its I/O contract out loud: “
03_annotateconsumes the filled-in sheet and our adjudication, and produces the gold set.”
Not a re-teach. You have run every one of these helpers already; today is about being able to name the structure — which is exactly what the Q&A will ask you to do.
Reading
No new reading for this session — see the Day 4 reading (Abdurahman et al., 2025, read in full) in Session 10 and on the Readings page.
Slides & Colab
Mini-project
- The inventory: what you have to work with.
- What you write today: the
PLAN.mdgate. - Where the work happens:
lda2-proj-template.
ImportantNo group calls the model until
PLAN.md is signed
Steps 1 and 2 — sampling and annotation — need no model at all, so there is plenty to get on with. The gate exists because a mismatched label set costs an hour to unpick after you have spent quota on it.