Vocabulary: From Familiarity to Retrieval in Real Tasks
When I first maintained this guide, I also liked collecting word lists. Every added page made progress look measurable. Yet language does not appear on demand merely because it has been saved. In an email, meeting, code review, or difficult explanation, the missing ability is often not having seen a word. It is choosing the right sense, collocation, register, and responsibility under limited time.
Delay, block, and risk, for example, can all appear in a project update, but they do not describe the same state. Whether work has stopped, may be affected, or is simply later than planned changes who must act, when to escalate, and whether to wait. A grammatically correct sentence can still send collaboration toward the wrong decision.
Vocabulary size is therefore not an inventory count. Vocabulary ability means recognising an expression in real input, deciding what it means here, retrieving it when needed, and allowing another person to understand, respond, or continue work.
This chapter follows a traceable path: preserve a first encounter before lookup; decide what to skip, infer, check, learn, or escalate; verify sound, form, sense, collocation, register, and concept separately; then remove the source sentence and card and retest across topic, mode, and real task.
Chapter at a Glance
- Define the real task instead of substituting an isolated vocabulary total.
- Preserve a first-encounter baseline and separate unknown words from unparsed sentences.
- Decide whether to skip, infer, look up, learn, or professionally verify each unknown item.
- Learn form, sound, current sense, collocation, grammatical behaviour, register, and concept separately.
- Assess with unaided recall, timed retrieval, and real audience results, not familiarity after seeing an answer.
- Let errors and use adjust spacing instead of treating fixed dates as a law of memory.
- Verify AI-generated candidates with dictionaries, corpora, current documentation, and domain experts.
- Track one topic, one first encounter, and one parallel task for fourteen days.
1. Define the Task Before the Total
"Expand my vocabulary" does not say what to learn today. Begin with the conditions in which language must work:
| Task | Vocabulary ability required | Observable result |
|---|---|---|
| Listen in a meeting | Segment key chunks and responsibility in natural speech | Retell decisions, risks, owners, and dates |
| Read documentation | Identify terms, actions, and limits in the current version | Run a minimum check and point to the source |
| Explain aloud | Retrieve useful chunks under time pressure and repair | Listener retells the point without guessing from a script |
| Write an email | Select accurate sense, collocation, and relational tone | Recipient knows what to do and when to reply |
| Take an exam | Recognise or produce the range required by the task | Complete a new item rather than recognise an old card |
One learner may read specialist terms yet fail to retrieve them in a meeting, or speak everyday English while confusing near-synonyms in a contract. Audience, material, time limit, permitted tools, and acceptance standard determine whether to test reception or production, breadth or precision.
Use the Vocabulary Evidence Card to keep the task, first encounter, verification, retrieval, feedback, and delayed transfer together.
2. Preserve an Unpolished First Encounter
Choose material close to the real task: a 150-300-word text, 60-90-second recording, email, API page, or short explanation. Record permitted captions, dictionary, translation, search, and AI. Then complete the first input or output without lookup.
Material, version, and source:
Real task and acceptance standard:
First-encounter time and limit:
Permitted support:
My gist, relationships, and next step:
Chunks I know with confidence:
Sentence/timestamp where meaning stopped:
How I handled unknowns then:Do not attribute every failure to vocabulary. Sound segmentation, syntactic scope, background knowledge, layout, or attention may be the real barrier. Preserve the original cloze answer, retelling, recording, or draft so later comparison can show what lookup actually changed.
A first-encounter baseline is not a shame list. It asks which meaning arrived before tools, where guessing began, and which item truly blocked the task.
3. Make Five Decisions about Unknown Items
Not every unknown word deserves long-term review. Choose among five actions:
| Decision | Use when | Minimum record |
|---|---|---|
| Skip | It does not affect gist, action, or safety and is one-off detail | Why it can wait |
| Infer | Syntax, contrast, example, or context provides usable evidence | Candidate sense and confidence |
| Quick lookup | It blocks current understanding but has low future value | Current sense and source |
| Learn deliberately | It recurs, is task-critical, and will be used within two weeks | Sound, collocation, retrieval cue, and retest |
| Verify professionally | Safety, law, medicine, finance, versioning, or business responsibility is involved | Primary definition, scope, and accountable owner |
Inference before lookup can train contextual judgment, but do not protect a "no dictionary" rule at the expense of safety or fact. Low-risk reading can carry uncertainty; high-risk action must return to current primary documentation or a qualified person.
Select only five to eight high-value chunks for deliberate learning in one cycle. Their value comes from recurrence, task impact, connection to known language, and an imminent use, not from sounding advanced.
4. One Word Contains at Least Eight Questions
English = translation records only one candidate correspondence. Fuller vocabulary knowledge includes:
| Dimension | Question | Example/evidence |
|---|---|---|
| Form | How is it spelled, inflected, and derived? | analyse → analysis → analytical |
| Sound | How does it appear through stress, phonemes, and connected speech? | Locate it without captions |
| Current sense | What exactly does it mean in this sentence? | Explain this use instead of listing every definition |
| Collocation/chunk | Which words normally complete the meaning with it? | pose a risk to, meet a deadline |
| Grammar | What follows it, and what can serve as subject? | Original sentence plus a structural contrast |
| Register/pragmatics | Do formality, politeness, stance, and identity fit? | kids / children / minors |
| Concept/reference | Is the real object or institution actually equivalent? | Versioned definition or counterexample |
| Retrieval condition | How fast does it arrive, and what cue is still required? | Timed speaking/writing without prompts |
A word family is an estimation tool, not automatic knowledge. help, helps, helped, helpful, helpless may count as one family in some studies, while a learner may not know each derivative's sense, collocation, or tone. Abbreviations, API names, and fixed specialist expressions require current-version checks.
Words rarely work alone. Prioritise combinations that perform tasks: collocations, chunks, morphological relations, sentence slots, and register contrasts. They reduce live assembly without becoming an entire memorised script; the combination must remain adaptable in a new sentence.
5. The Relevant Sense Is Not Always Definition One
Polysemy creates the feeling of "I knew every word but misunderstood the sentence." Preserve the source before comparing senses:
Source sentence and location:
My first inferred sense:
Dictionary sense that fits this syntax and context:
Supporting collocation, example, or counterexample:
How a near-synonym changes responsibility, certainty, or tone:Do not paste the first translated gloss into every setting. Learner dictionaries are useful for common senses, pronunciation, grammar labels, collocations, and examples. Corpus evidence can show whether an expression occurs in comparable contexts. Technical terms should return to current official documentation, standards, or code.
Translation can support comparison, but one first-language word may cover several English concepts, while two institutional terms that look equivalent may differ in scope. "Close enough" is not an acceptance standard for high-risk text.
6. Receptive Vocabulary Is Not Productive Vocabulary
Receptive reading means understanding after seeing; receptive listening means segmenting and understanding when sound arrives. Productive speaking and writing require retrieval before the answer is displayed, together with form, collocation, and register.
Using matched receptive and productive translation tests with L2 learners, Webb (2008) reported larger receptive than productive vocabulary sizes. Under fuller scoring, the gap increased in lower-frequency bands. The finding warns against turning recognition into a claim of use, while the size of any gap still depends on test direction, scoring, participants, and frequency range.
Move a chunk from reception toward production through four steps:
- Explain its current sense in the source sentence.
- Close the source and retrieve it from a situational cue.
- Keep the function while changing people, time, topic, or tone.
- Use it in an email, meeting, explanation, code review, or conversation and observe the response.
Reproducing only the training sentence is not transfer. Conveying the concept with an awkward collocation means meaning arrived but form still needs repair. Locate the failed dimension instead of calling every failure "not memorised."
7. Interpret Coverage Numbers Carefully
Earlier editions said that 1,000 words yield 75% understanding and 7,000 nearly 90%. Such unconditional equations mix counting units, corpora, modes, and definitions of understanding, so they are not retained.
Vocabulary research often discusses 95% and 98% coverage of running words. As an intuition, 95% leaves about one unknown in twenty and 98% about one in fifty. Actual difficulty still changes when unknowns cluster at decisive points, background is unfamiliar, or syntax is dense.
Schmitt, Jiang, and Grabe (2011) asked 661 participants from eight countries to read two texts and complete vocabulary and comprehension measures. The relationship between known-word percentage and comprehension was relatively linear, without evidence of a sharp threshold; the authors judged 98% a more reasonable target for academic texts. That group relationship is not an individual passport.
Using 1,000-word-family lists built from the British National Corpus and assuming 98% coverage for unassisted comprehension, Nation (2006) estimated roughly 8,000-9,000 word families for written text and 6,000-7,000 for spoken text. The estimates depend on family definition, corpus, proper-noun treatment, and the comprehension assumption. They are not universal goals for every learner, domain, or task.
Coverage is useful for choosing material. When unknowns are dense and task-critical, reduce difficulty, shorten the source, or build background. When most unknowns do not affect the line, continue reading instead of turning every page into a word list.
8. A Strong Retrieval Card Tests One Decision
The front should demand recall rather than another reading of the answer:
Context: a project dependency may affect the schedule; express that it creates risk.
Gap: The dependency may ____ a risk to the schedule.The back preserves:
pose a risk toand a reliable source location;- current sense, pronunciation, part of speech, and register;
- a second situation written by the learner;
- one confusable expression and the difference;
- the prompt to remove at the next retest.
Long cards turn review into rereading. If one failure combines pronunciation, sense, and collocation, separate the cues. If a card succeeds only with its original sentence, replace it with a cue closer to the real task.
Bidirectional translation cards can be a beginning, not the final condition. Real retrieval moves from intention, situation, or object toward expression rather than forever beginning with another-language label.
9. Let Performance Control the Interval
Forgetting is real, but no fixed "memory-curve calendar" fits every item and learner. Set initial checks for the same day, day 1, days 3-7, day 14, and day 30, then adjust from evidence:
- fast accurate recall plus new use: lengthen the interval;
- effortful success: keep or slightly lengthen it;
- repeated sense/near-synonym confusion: add contrast and counterexample;
- understood in text but missed in speech: add a no-text audio cue;
- correct card but failed real task: pause new items and use timed output;
- repeated failure: split the card, reduce volume, or repair concept, sound, and background;
- no longer relevant to the task: archive it instead of reviewing forever because of sunk cost.
Webb, Yanagisawa, and Uchihara's (2020) meta-analysis included 100 effect sizes from 22 studies of flashcards, word lists, writing, and fill-in-the-blank activities. The abstract reports average immediate meaning/form recall gains of 60.1%/58.5%, falling to 39.4%/25.1% on delayed tests, with wide variation among activities. It supports deliberate learning and expected forgetting, not a guaranteed deck, interval, or individual outcome.
Anki can schedule intervals; it is not the method itself. Card count, streak, and mature-item totals deserve attention only when they change a decision to continue, split, transfer, or stop.
10. Let the Error Choose the Next Practice
Classify each weak result before adding more items:
| Error | Likely gap | Smallest next repair |
|---|---|---|
| Form not recognised | Spelling, derivation, or inflection | Mark root/part of speech and use it in a new sentence |
| Sound not recognised | Phoneme, stress, linking, or boundary | Compare a short clip, then remove text |
| Wrong sense | First-definition dependence or ignored context | Contrast senses and add a counterexample |
| Awkward collocation | Isolated-word learning | Learn a whole chunk and inspect comparable corpus use |
| Wrong register | Relationship, formality, or stance ignored | Write the same intention for two audiences |
| Retrieval too slow | Cue does not match real intention | Give a 3-5-second situational prompt and answer aloud |
| Transfer failure | Dependence on source sentence, topic, or mode | Change topic, sound/text mode, or task and produce again |
Ask a real reader or listener to retell the received meaning before discussing whether it sounds natural. Separate necessary repairs that change fact or relationship, acceptable variation, register choices, and individual style. More native-like is not automatically more suitable for the author and situation.
11. AI Generates Candidates, Not Lexical Authority
AI can produce contextual cloze items, compare near-synonyms, simulate audiences, offer counterexamples, classify errors, and generate parallel tasks. Use it in this order:
- Preserve the first encounter and your own candidate.
- Ask what layer of meaning each suggestion changes.
- Request uncertainty, regional difference, and register limits.
- Verify against a dictionary, corpus, current documentation, or domain expert.
- Close AI and retrieve under time in a changed context.
- Record acceptance, rejection, and reason.
Models can invent collocations, etymologies, quotations, and outdated technical definitions. Do not upload customer data, unpublished documents, exam answers, or unauthorised third-party material. Current primary sources and named owners decide high-risk terminology.
12. Technical Word Lists Are Indexes, Not Courses
The repository lists begin from tasks rather than daily quotas:
| List | Task entry |
|---|---|
| Common | Everyday and cross-topic work |
| Prompt | AI tasks, constraints, and acceptance language |
| Vibe Coding | Agent collaboration and code review |
| JavaScript | Browser, front end, and asynchronous flows |
| Python | Data, automation, and scripting |
| Go | Services, concurrency, and deployment |
| Java | JVM and enterprise systems |
| PHP | Web back ends and legacy maintenance |
| Rust | Ownership, performance, and systems programming |
| Swift | Apple platforms and application development |
Choose five to eight chunks from the current task and record documentation version and source. Produce once without the list and retest in a parallel task a week later. Languages, frameworks, and products change; a list cannot replace official documentation or current runtime evidence.
13. A Thirty-Five-Minute Session
- Five minutes: state the task, audience, material, and acceptance standard.
- Seven minutes: complete a no-lookup first encounter or unaided output.
- Six minutes: decide what to skip, infer, look up, learn, or verify.
- Seven minutes: check sense, sound, collocation, and register for five to eight chunks.
- Seven minutes: close the source and speak or write in a changed situation.
- Three minutes: record one main error and one cue to remove next.
Completion matters more than volume. On a low-capacity day, preserve one first encounter, one high-value chunk, and one unaided sentence. Do not spend all the time maintaining a card system instead of using language.
14. A Fourteen-Day Vocabulary Experiment
| Day | Action | Evidence |
|---|---|---|
| 1 | Choose a real topic/task and complete a first encounter | Raw understanding/output, conditions, and block |
| 2 | Make five decisions about unknown items | Skip, infer, lookup, learn, and verify queue |
| 3 | Verify eight dimensions of valuable chunks | Sense, sound, collocation, register, and sources |
| 4 | Build cues that test one decision | Cards and scaffold-removal condition |
| 5 | Complete cloze, retelling, or short writing without prompts | Retrieval time and error type |
| 6 | Change one mode across listening, reading, speaking, and writing | Receptive/productive gap |
| 7 | Test with real material or audience | Understanding, response, and repair |
| 8 | Archive low-value items and keep five to eight | Selection reason and capacity change |
| 9 | Contrast a near-synonym, counterexample, or common misuse | Meaning boundary and rejection reason |
| 10 | Keep topic, change audience and register | Relationship and tone transfer |
| 11 | Keep task, change topic | Chunk-structure transfer |
| 12 | Let AI/a peer propose candidates, then verify | Source checks and uptake record |
| 13 | Deliver under time pressure | Email, recording, explanation, or review result |
| 14 | Close old cards/prompts and complete a parallel task | Evidence to keep, split, archive, or move on |
Fourteen days is not a vocabulary-fluency deadline. It asks whether, after the source sentence, card, and immediate prompt leave, you can recognise, choose, and retrieve the chunks in a changed task.
15. Evidence That Vocabulary Is Becoming Ability
- A first encounter preserves the line instead of letting every unknown item capture attention.
- You can decide whether to skip, infer, look up, learn, or professionally verify an item.
- You can explain current sense, source, collocation, register, and counterexample instead of only a translation.
- You segment the item in natural speech and recognise its written form.
- You retrieve it under time after the answer is closed.
- A real reader/listener receives the intended meaning, relationship, and next step.
- An error enters the next practice instead of merely increasing card volume.
- Retrieval transfers across topic, audience, mode, or task.
- You can archive what is no longer useful and return attention to current life.
Vocabulary growth does not always make a list longer. Sometimes it means knowing which unknown can wait while reading, which term must be precise in a meeting, and why a more impressive phrase is wrong for this relationship. Unknowns no longer stop you all at once, and familiarity no longer hides failed retrieval.
Sources and Boundaries
- Nation (2006), How Large a Vocabulary Is Needed for Reading and Listening?: uses British National Corpus word-family lists and a 98% coverage assumption to estimate unassisted comprehension needs; counting unit, corpus, proper-noun treatment, and comprehension criterion limit transfer.
- Schmitt, Jiang & Grabe (2011), The Percentage of Words Known in a Text and Reading Comprehension: a 661-participant, eight-country study found a relatively linear coverage-comprehension relationship without a clear sudden threshold; it does not guarantee individual or cross-genre outcomes.
- Webb (2008), Receptive and Productive Vocabulary Sizes of L2 Learners: compares receptive/productive knowledge through matched translation tests; direction, scoring, frequency bands, and participant conditions limit interpretation.
- Webb, Yanagisawa & Uchihara (2020), How Effective Are Intentional Vocabulary-Learning Activities?: a meta-analysis of 100 effect sizes from 22 studies reports substantial immediate/delayed and between-activity variation; averages do not guarantee a fixed schedule or individual result.
- Pashler et al. (2008), Learning Styles: the review found no adequate evidence base for matching instruction to visual/auditory learner labels while cautioning that untested variants cannot all be declared false.
- Senses, technical terms, corpus frequencies, products, and dictionaries change. Important tasks should return to current versions, primary sources, real audiences, and named accountable owners.
Related entry points: Grammar | Listening | Reading | Speaking | Writing | Learning English with AI | Vocabulary Evidence Card | Evidence Chain Template
Closing: Let the Word Arrive When Needed
What remains in memory is often not the first line of a dictionary entry but the moment you needed the word: explaining a risk to a colleague, asking a stranger for help, apologising to someone close, or finally entering a page you once had to avoid.
A list can preserve order; it cannot live for you. A word becomes ability only after entering sound, sentence, relationship, and consequence. Nor must it remain forever. Some words finish one piece of work and can be archived; others, used repeatedly, slowly grow into your own voice.
You do not need to own an entire language at once. Begin with a few words you truly need and let them return after the answer is hidden. They do not need to prove that you are clever. At the place where meaning would otherwise break, they only need to help you take one more step toward the world.