កម្រងទិន្នន័យភាសាខ្មែរ · Khmer corpus linguistics

A searchable Khmer corpus, free to use.

4 built-in corpora — 28,917,149 tokens of Khmer, spanning 1926–2026 — free with any account. Every corpus is word-segmented and indexed for frequency, concordance, collocation, keyness, dispersion and part-of-speech analysis. Compare literary Khmer against the language as it's written today, or build a corpus of your own from text, PDFs, or Khmer websites.

Built-in corpora at a glance

28,917,149
Tokens, free
4
Corpora
195
Documents
1926–2026
Years covered

Simple pricing

The full analysis toolkit is free for everyone. Researcher adds unlimited export and creating your own corpus from uploaded text; Pro adds building a corpus straight from the web.

Free
Browse and analyze every built-in corpus
$0.00 / month
  • Full analysis toolkit — frequency, concordance, collocations, keyness, POS, NER, dispersion, segmentation, IPA, X-bar
  • View and work with every built-in corpus
  • No downloads
  • Can't create a new corpus
Choose Free
Pro
Everything, plus build a corpus straight from the web
$9.99 / month
  • Everything in Researcher
  • 🕸 Web crawling — paste Khmer website links and we fetch, clean, segment and index the text as your own corpus
  • Priority support
Choose Pro