About this platform

កម្រងទិន្នន័យភាសាខ្មែរ (Khmer Corpora) is a general-purpose corpus-linguistics platform for the Khmer language: frequency, concordance, collocation, keyness, dispersion, part-of-speech tagging and named-entity recognition, all in your browser. On the Pro plan, upload your own Khmer text or PDF — including scanned pages, run through OCR automatically — and it becomes your own corpus, analyzed with the exact same toolkit and saved to your account for later.

The built-in sample: កម្ពុជសុរិយា

To explore the toolkit right away, every account gets the កម្ពុជសុរិយា (Kambuja Suriya, "Sun of Kambuja") archive — Cambodia's oldest scholarly journal, published by the Buddhist Institute from 1926 onward. This platform indexes a digitised, word-segmented run of the journal spanning 1926–2006 — roughly 9.5 million tokens across 61 yearly volumes.

The segmentation was produced automatically with a dictionary + Viterbi word segmenter anchored on the official Khmer dictionary and orthography list, then indexed for instant querying. Word counts, KWIC concordance, collocation statistics and more are computed live from the underlying text — the same pipeline used on whatever you upload yourself.

Kambuja Suriya is the only pre-loaded dataset — it's a sample, not the whole platform. Other datasets from the wider research project (Wikipedia, CC-100, OPUS, Wikisource, and others) are not included here; bring your own text instead.