About this platform

កម្រងទិន្នន័យភាសាខ្មែរ (Khmer Corpora) is a general-purpose corpus-linguistics platform for the Khmer language: frequency, concordance, collocation, keyness, dispersion, part-of-speech tagging and named-entity recognition, all in your browser. Two corpora come built in and are free with any account; on the paid plans you can add your own — pasted text, PDFs, or Khmer websites — and analyze it with the same toolkit, saved to your account for later.

Built-in corpus 1: កម្ពុជសុរិយា (1926–2006)

The កម្ពុជសុរិយា (Kambuja Suriya, "Sun of Kambuja") archive — Cambodia's oldest scholarly journal, published by the Buddhist Institute from 1926 onward. This platform indexes a digitised, word-segmented run of the journal spanning 1926–2006: roughly 9.5 million tokens across 61 yearly volumes of literary and scholarly Khmer.

The segmentation was produced automatically with a dictionary + Viterbi word segmenter anchored on the official Khmer dictionary and orthography list, then indexed for instant querying. Word counts, KWIC concordance, collocation statistics and more are computed live from the underlying text — the same pipeline used on whatever you add yourself.

Built-in corpus 2: ព័ត៌មានខ្មែរ (2019–2026)

A modern counterpart: roughly 13 million tokens of Khmer from national news outlets and government ministry websites, crawled in 2026 and segmented with the same pipeline. Crawling followed each site's own links and sitemaps, which reach well back into their archives, so coverage is not evenly spread: dates referenced in the text run from 2009 to 2026 but concentrate heavily on 2019–2023. Treat it as a corpus of recent Khmer rather than of this month's news.

Publishers were crawled only where their robots.txt permits it, at a rate limit that keeps the load light. Page furniture — navigation, headline lists, related-article blocks, footers — was stripped, and lines repeating across pages were removed: 83% of the raw crawl turned out to be site template rather than journalism, and leaving it in would have made every frequency count a measure of page design instead of language.

Why two

Running the same query against both is the most direct way to see how Khmer has moved in a century — vocabulary that has entered or fallen out of use, spellings that have settled, and the distance between literary and journalistic register. The keyness tab compares one corpus against another and is built for exactly this.

These two are the pre-loaded datasets. Other collections from the wider research project (Wikipedia, CC-100, OPUS, Wikisource, and others) aren't published here — add your own text, PDFs, or website links instead.