Half of 72,454 lemmas never appear in two million sentences
Also available, written for a general readerOnly 899 of Our 3,289 New-Word Candidates Survived the FilterThis count of 38,559 is not a verdict on our vocabulary, but a reminder that our measurement tools have gaps. We found 3 failure points while dissecting our own counter-list.
Key numbers
Where these come fromMeasured from coverage_report.txt: cakupan lema terhadap korpus and 2 more
We took the lemma list we use as our standard-form reference and matched it, word by word, against two million real Indonesian sentences. The result: 38,559 of 72,454 lemmas, or 53.2 percent, never appear at all.
That number sounds like an accusation against the dictionary. It is not, and it is not a statement about any dictionary other than the list we hold. Where that list came from is named in the raw-data block below, which is where it belongs rather than in the conclusion. And the most useful part of this work only arrived when we tried reversing the question, and found that the list we built ourselves was flawed in several places.
What was actually measured
The corpus comes from the Leipzig Corpora Collection, two files combined: one million mixed web, news, and wiki sentences from 2013, plus one million news sentences from 2020. Both were retrieved on 30 July 2026 and are licensed CC BY 4.0.
Matching was done on base forms. A lemma counts as "appearing" if its base form is present in the corpus frequency list, with no full morphological analysis. That simplification matters: a word that only ever occurs in an affixed form in the corpus can be counted absent while genuinely being in use.
Of 72,454 lemmas, 33,895 received a real frequency figure. The remaining 38,559 did not. This does not mean those words are dead. A dictionary is supposed to hold rare words. A list containing only the words that turn up often in the news is not a dictionary, it is a summary of the news.
The interesting part is not the size of the number but that it could be computed at all, and then used for something practical.
The other direction: words in use, absent from the list
The more interesting question is the reverse. Which words turn up often across those two million sentences, yet appear nowhere in the lemma list?
We filtered for words occurring at least ten times with no matching lemma. That produced 3,289 words, accounting for 288,646 occurrences in total.
The top ten are almost entirely technology and internet vocabulary: test, website, link, live, smartphone, streaming, file, iphone, handphone. That is no surprise. What is surprising is how quickly the list stops being trustworthy once you look closer.
Where our own number cracks
Every word on that list carries one extra column, cap_ratio, the share of its occurrences that begin with a capital letter. It was meant as a crude signal: a word that is nearly always capitalised is probably a proper noun rather than a loanword that has genuinely entered everyday language.
That column is what demolished our own list.
Of the 3,289 candidates, 852 words, or 25.9 percent, have a cap_ratio above 0.5. More than half of their occurrences are capitalised. Link (0.558), home (0.558), and live (0.524) almost certainly come mostly from brand names, show titles, and place names rather than from use as ordinary words. Mall, with 1,205 occurrences and a cap_ratio of 0.508, is most likely part of building names rather than a common noun waiting to enter the dictionary.
What survives as a convincing candidate is far smaller. Only 899 words, or 27.3 percent, have a cap_ratio below 0.2. This is where words like website (0.092), file (0.090), smartphone (0.131), and handphone (0.115) sit, and this is the group worth taking seriously.
That group is also smaller than its count suggests. All 899 words together account for just 84,937 occurrences, or 29.4 percent of the total. Put differently, nearly three quarters of the occurrences in our "foreign words" list come from words that fail even the most basic filter.
It should be said plainly that cap_ratio is not a validated measure. It counts capital letters, and capital letters also begin sentences. A word that often starts a sentence will look more like a proper noun than it really is. We use it because it is cheap and sufficient to screen out the worst cases, not because it is correct.
The list is also far shorter than the figure 3,289 implies. Only 30 words occur a thousand times or more, and only 672 occur a hundred times or more. The remaining 2,617 occur between ten and ninety-nine times across two million sentences. The median across the whole list sits at 30 occurrences.
At that frequency, one long article repeating the same term is enough to put a word on the list. Thirty occurrences in two million sentences is roughly one occurrence every sixty-six thousand sentences. Calling that "a frequently used word" is plainly a stretch.
The imbalance is clearer from the mass side. The top thirty words alone contribute 48,493 occurrences, or 16.8 percent of all 288,646. The long tail is crowded in count but thin in weight, and it is precisely this crowded-but-thin part that most easily makes a finding sound larger than it is.
The corpus has a date on it
One more thing became visible only after the list was sorted.
The most frequent word on the entire list is test, with 4,494 occurrences. Third place goes to rapid, with 3,241. In eleventh place sits lockdown, with 1,057.
None of those is technology vocabulary. All three come from a single event. Half of this corpus is news from 2020, and rapid is not a loanword entering Indonesian, it is half of "rapid test". It ranks third because of one pandemic, not because of one shift in the language.
This is not a flaw in the corpus. A corpus is a snapshot of a period, and Leipzig states its periods clearly. The flaw is in how we read the result: sorting by frequency and reading the top ten as "words currently entering Indonesian" is a conclusion the data does not support. What the top ten actually shows is what Indonesian news wrote about in 2020.
What this counting changed
Most of this work ended not as a finding but as a table a machine uses.
Our correction engine has to decide, when a misspelled word has several possible corrections, which one the writer most likely meant. Before frequency data existed, that choice fell to a word-length rule, which is essentially a guess in a good suit. For the 33,895 lemmas that now carry a real frequency, the guess can be replaced with evidence.
For the other 38,559, nothing changed. The old rule still applies, and the output file carries a warning that it must not be removed. The temptation is real: once you have good frequency data, it becomes easy to treat words without frequency as unimportant. The opposite is true. A word absent from a news corpus is exactly the word that needs careful handling, because it is the word we know least about.
What can and cannot be concluded
After three rounds of narrowing, what remains is far smaller than the headline the first number would allow. That is how it should go.
What survives: 899 words, twenty-seven percent of the candidates, behave like ordinary words rather than proper nouns, and most of them are consumer technology vocabulary. Whether those words are a passing trend or loanwords that will settle, this data cannot say, because we have two snapshots in time rather than a time series. All that can be claimed is that the list is short enough and clear enough for a person to review one by one, and that is what it is for.
What does not survive: the claim that 53.2 percent of any dictionary is "unused". What was measured is the copy of the lemma list we held on 30 July 2026, not any published edition, and the two may well already differ. A word's absence from a news and web corpus is not evidence that people do not use it. This corpus leans toward formal written language. Speech, regional varieties, and everyday conversation are barely represented, and that is precisely where many of the "missing" lemmas most likely live.
What also does not survive: the claim that we found 3,289 words that belong in the dictionary. The honest figure is closer to 899, and even that remains a list of candidates rather than a list of decisions.
The raw data behind every number on this page is downloadable, and you can recompute it yourself. If something is wrong, that is the message we are waiting for.
Data & provenance
- coverage_report.txt: cakupan lema terhadap korpus
- foreign_candidates_ranked.tsv: 3.289 kata asing berperingkat
- PROVENANCE.md: asal korpus Leipzig & lisensinya
Limitations
The lemma list measured is the copy we held on 30 July 2026, not any published dictionary edition, and the two may already differ. The Leipzig corpus leans toward news and web text, so speech and regional varieties are barely represented; a word's absence from the corpus is not evidence it is unused. Matching was done on base forms without full morphological analysis, so a word appearing only in affixed form can be counted absent. The cap_ratio column merely counts capital letters, which also begin sentences, making it a crude signal rather than a validated measure of adoption. Half the corpus is news from 2020, so the top ranks reflect one particular period rather than a durable shift in the language.
Revision history
- Published
- Last updated
Terms in this article
- lemma
- The base form of a word that heads a dictionary entry, and the form someone looks up. A single lemma can cover several senses at once, and each sense can carry its own example sentence.
Written by the author for this article, not taken from a dictionary entry.