Sonja PhraseFlow · Sonja Projects
Author: Sonja Sahebzad · Utrecht, the Netherlands
The eight additional models use sentences by Tatoeba contributors, downloaded from the official exports dated 26 September 2026, under Creative Commons Attribution 2.0 France. Text was normalized, segmented, sampled and aggregated into statistical models. Tatoeba does not endorse this app.
The files below preserve the credited contributor, sentence ID and original sentence URL for each training sentence. A missing owner in the original export is preserved as supplied; the sentence page provides its history. Held-out evaluation answers are not included.
English, German, Finnish and Russian use the official Coursera SwiftKey corpus. The project retains its source archive and checksum. No course quiz text or answers are included in these attribution lists.
The app uses R, Shiny, data.table, jsonlite and stringi. Chinese segmentation uses ICU dictionary word boundaries. The compact language models use interpolated Kneser-Ney smoothing. Download checksums, runtime versions and transformations are recorded with the experiment scripts.