Twenty-five years of Turkish writing, read not word by word but
pair by pair — which ideas a society links together, and how the very meaning of its
words drifts across the decades.
7.6kwords embedded
4eras, 2006–2026
PMIin-sentence & in-article
akıllı“smart” → smartphones
Looking for the lexical history?
How individual Turkish words rose, fell, went native, and changed spelling — plus what the culture chose to write about.
The first thing co-occurrence reveals is Turkish's set phrases and the ideas it binds —
measured by how often two words appear in the same sentence and the same article, every year.
01 · the n-grams
The phrases each era spoke in
The strongest two-word collocations — words that occur together far more than chance (highest PMI) — read like a register of what Turkish was busy describing, era by era: folk tales and village life, regional food, biography, the modern world.
2006-2010
kısmi diferansiyel
bek ret
yırtıcı kuşlar
başrollerini paylaştığı
panik atak
wonder woman
politika izledi
teğmen rütbesiyle
öldüğünde yaşındaydı
antakya prensliği
ınto wild
bıçak sırtı
komplo teorileri
yazım kılavuzu
2011-2015
iyilik kötülük
yönlendirme protokolü
dağın tepesinde
kıyafet değiştirmiş
süs eşyaları
yazım kılavuzu
gözlemci çerçevesi
yorgun savaşçı
filmden fotoğraflar
sağda solda
kabartma tozu
akut kronik
joan chen
zemin hazırlamıştır
2016-2020
doğmuş büyümüştür
iyilik kötülük
hanede yaşıyordu
okunuşu şeklindedir
nina simone
sağda solda
panik atak
yorgun savaşçı
süs eşyaları
ingilizler fransızlar
gerard butler
yukarıdaki örnekte
gregory timothy
tarımla uğraşan
2021-2026
kavun karpuz
sokakta dolaşarak
feodal beyler
yöresel yiyecekler
bulgur pilavı
içli köfte
cesedi yakıldı
iyilik kötülük
dağın tepesinde
kavak söğüt
müzeye dönüştürüldü
oğuzların boyundan
büyülü sihirli
rahibe mistik
02 · rising phrases
The set-phrases Turkish increasingly writes
Now the change over time. Ranking collocations by how much their frequency grew from 2006–2012 to 2019–2026 (corrected for the corpus's overall growth) surfaces what entered the written record: kuzey makedonya (a country renamed in 2019), cumhurbaşkanı recep, covid salgını, google play, xbox one, the language of football and TV. This is the hypothesis engine — each rising phrase is a candidate story about what Turkish knowledge turned to.
03 · ideas sharing a sentence
Concepts that converged, year by year
Co-presence with time made literal: how often a pair of words lands in the same sentence, each year. New ideas assemble in real time — yapay zeka (artificial intelligence), sosyal medya, iklim değişikliği (climate change) go from never-together to everyday pairings.
Part II
Meanings on the move
Co-presence also lets us do something deeper: describe every word by the company it keeps, in
each era separately, and watch words move. A word whose neighbours change has changed its
meaning — distributional semantics, on 25 years of Turkish.
04 · changing company
When a word changes the company it keeps
For each probe word, its nearest neighbours by in-sentence association in the first era versus the last (16 of our probes kept under a quarter of their old neighbours). The highlighted neighbours are new — and they tell the story of the era: akıllı (‘smart’) leaves flying/clever behind for phones, tablets, devices; bulut (‘cloud’) leaves the sky for computing; sosyal finds media.
Every embedded word placed by meaning-neighbourhood (UMAP over per-era PPMI vectors): related words fall close together, forming regions — places, people, science, sport, the digital world. Size is how often the word is used; colour is how far its meaning drifted over 25 years. Our probe words are marked in gold.
Source: the complete Turkish Wikipedia revision history (2002–2026), prose articles
only (the bot-imported settlement stubs are excluded). For every article's year-end snapshot we
computed in-sentence and in-article co-occurrence over a 15,000-word content vocabulary, then
per-era PPMI word embeddings (SVD, dim 200) aligned across eras by orthogonal Procrustes — the
standard diachronic-embedding method (Hamilton et al. 2016). Association is PMI; drift is the
turnover in a word's nearest neighbours. Phrase growth is per-million-normalised to correct for
the corpus growing ~10× over the period. Trust the trends and neighbourhoods over exact
levels; polysemy and template residue add some blur.