កម្រងទិន្នន័យភាសាខ្មែរ · Khmer corpus linguistics

A research-grade platform for Khmer corpus linguistics.

Explore the built-in កម្ពុជសុរិយា (Kambuja Suriya) sample archive — eighty years of Cambodia's oldest scholarly journal, word-segmented and indexed for frequency, concordance, collocation, keyness, dispersion and part-of-speech analysis. On the Pro plan, upload your own Khmer text or PDF (scanned pages included) and get the same analysis on your own corpus.

Sample archive at a glance

9,490,670
Tokens
51,238
Distinct words
61
Journal issues
1926–2006
Years covered

Everything a corpus linguist needs

Frequency and lexical-richness statistics, KWIC concordancing, collocation and co-occurrence networks, keyness, dispersion, POS tagging, named-entity recognition, and more — grouped into a clean, dropdown-organised workbench.

Frequency & lexical richness

Ranked word lists, TTR/STTR/MATTR, Herdan's C, Guiraud's R.

Concordance (KWIC)

Every occurrence of a word in context, sortable by left/right/POS.

Collocations & networks

MI, t-score, log-likelihood, Dice — plus a visual "mental lexicon" graph.

Keyness

What's distinctive about one journal year versus the rest of the archive.

POS & named entities

A trained HMM tagger (~89% accuracy) and gazetteer NER for people, places, villages.

Segmentation, IPA, X-bar

Word-segment raw Khmer text, transcribe to IPA, and visualise syntax trees.

Simple pricing

The full analysis toolkit is free for everyone. Researcher adds unlimited export and creating your own corpus from uploaded text; Pro adds building a corpus straight from the web.

Free
Browse and analyze the built-in corpus
$0.00 / month
  • Full analysis toolkit — frequency, concordance, collocations, keyness, POS, NER, dispersion, segmentation, IPA, X-bar
  • View and work with the existing corpus
  • No downloads
  • Can't create a new corpus
Choose Free
Pro
Everything, plus build a corpus straight from the web
$9.99 / month
  • Everything in Researcher
  • 🕸 Web crawling — paste Khmer website links and we fetch, clean, segment and index the text as your own corpus
  • Priority support
Choose Pro