Sonja PhraseFlow suggests the next word in twelve selectable languages. English, Dutch, French, German, Finnish, Russian, Korean, Mandarin Chinese, Portuguese, Italian, Hindi and Spanish each use a separate model. The additional languages remain experimental. A suggestion is a learned continuation, not a guarantee of meaning or correctness.
The User guide / Handleiding button above the writing area opens this document. The How to use tab contains a shorter guide and a link to training-data credits. Ctrl + Enter, or Command + Enter on a Mac, triggers prediction. Switching language clears the phrase and old suggestions; it does not translate text. The interface labels remain English.
The first suggestion stays prominent; the other selected word choices are all visible beneath it. Select 5, 10 or 20 suggestions to adjust the amount shown. More choices help a writer find a suitable alternative, but do not improve the ranking of the first suggestion.
Short continuations is a separate, experimental lookup of endings from existing training sentences. It matches the last two to four normalized units and offers up to three alternatives. Three- or four-unit matches can add an observed ending of up to three units. Two-unit matches offer only the next unit because there is less context. The matched phrase appears below the buttons. No match means that no extra continuation is shown.
The lists were built from the existing training files in all twelve languages. Candidate endings are ranked by their training frequency; no quiz answers or evaluation targets were added. Only contexts with at least two retained training occurrences are kept. Selecting a button appends its words, then the predictor updates for the new text. Chinese continues without inserted spaces. The existing next-word scores are unchanged.
For example, the Dutch next-word model already ranks elkaar first after we houden van. The extra lookup can offer gelukkig after the suffix lang en, which the compact next-word ranking missed. These are inspected demonstrations, not independent accuracy measurements. A local phrase match can still be inappropriate for the meaning of the entire sentence. The app does not yet reliably understand a writer’s intention or guarantee a natural complete sentence.
Chinese text is divided into units using the ICU word dictionary. Suggested Chinese text is appended without inserting spaces. The unit is the segment found by ICU, which need not match every reader’s preferred word boundary. This version covers Mandarin corpus text, not every Chinese language or regional variety.
Korean predicts an eojeol, a unit separated by spaces that may include grammatical endings. Hindi vowel signs and other combining marks are retained; accented Latin letters and Cyrillic text are also preserved. When an input method is composing characters, automatic prediction pauses until the text is committed.
The app accepts up to 500 characters and can receive a complete sentence. Predictions use at most the four most recent normalized units. The model does not use the full sentence’s meaning, translate text or learn from the current session. For an unfamiliar phrase, shorter patterns and frequent words provide a fallback.
First means the first suggestion equals the recorded next word. Within three means that word appears among the first three. Other continuations may be reasonable, but the test scores exact matches. A 95% interval describes sampling uncertainty around a test percentage; it is not confidence in an individual suggestion. Individual confidence has not been calibrated.
The English results below are from the archived independent evaluation. The reserved English test cases were not reopened. A separate regression check confirmed identical words and scores on all 600 English development cases.
The English model matches the recorded next word first in 153/900 cases (17.0%; 95% interval 14.7–19.6%) and within three in 257/900 (28.6%; 25.7–31.6%). All 900 received a suggestion. Receiving a suggestion and receiving the correct suggestion are separate outcomes.
Dutch, French, Korean, Mandarin Chinese, Portuguese, Italian, Hindi and Spanish use Tatoeba example sentences, retrieved from the official exports dated 26 September 2026. Text was normalized and exact duplicates were removed before separating 600 held-out sentences per language. One target was chosen at random per held-out sentence. Candidate words were generated from the prefix only. Chinese prefixes were supplied without inserted spaces and segmented again without the target word.
Settings were fixed before evaluation: five-gram Kneser-Ney smoothing, discount 0.75, up to 40,000 vocabulary items, minimum count two, context support three and eight retained continuations per context. Up to 75,000 sentences were used per language; the Korean and Hindi sources supplied fewer eligible sentences. No setting was selected using these held-out outcomes.
| Language | First correct | Within three | 95% interval: within three |
|---|---|---|---|
| Dutch | 97/600 (16.2%) | 165/600 (27.5%) | 24.1% to 31.2% |
| French | 112/600 (18.7%) | 188/600 (31.3%) | 27.8% to 35.2% |
| Korean | 37/600 (6.2%) | 58/600 (9.7%) | 7.6% to 12.3% |
| Mandarin Chinese | 62/600 (10.3%) | 126/600 (21.0%) | 17.9% to 24.4% |
| Portuguese | 104/600 (17.3%) | 172/600 (28.7%) | 25.2% to 32.4% |
| Italian | 112/600 (18.7%) | 178/600 (29.7%) | 26.2% to 33.4% |
| Hindi | 134/600 (22.3%) | 197/600 (32.8%) | 29.2% to 36.7% |
| Spanish | 84/600 (14.0%) | 143/600 (23.8%) | 20.6% to 27.4% |
| Language | Training sentences | Reference top-3 | Median ms | File MiB |
|---|---|---|---|---|
| Dutch | 75,000 | 5.2% | 0.73 | 0.33 |
| French | 75,000 | 7.0% | 0.76 | 0.42 |
| Korean | 14,602 | 2.3% | 0.75 | 0.14 |
| Mandarin Chinese | 75,000 | 8.3% | 0.65 | 0.35 |
| Portuguese | 75,000 | 6.0% | 0.76 | 0.43 |
| Italian | 75,000 | 6.7% | 0.76 | 0.33 |
| Hindi | 15,466 | 13.2% | 0.76 | 0.10 |
| Spanish | 75,000 | 6.7% | 0.75 | 0.44 |
These results apply to held-out sentences from the same example-sentence collection. Similar sentence templates can remain after exact deduplication. The corpus is volunteer-written, may contain errors and is not a representative sample of everyday messages or news. Different training sizes, scripts and prediction units prevent a fair language-to-language ranking. These new measurements must not be interpreted as an improvement over the English result.
German, Finnish and Russian retain their earlier models trained on 75,000 course-corpus lines per language. Their 600-case evaluations contain 200 examples each from blogs, news and Twitter.
| Language | First correct | Within three | 95% interval: within three |
|---|---|---|---|
| German | 47/600 (7.8%) | 91/600 (15.2%) | 12.5% to 18.3% |
| Finnish | 46/600 (7.7%) | 71/600 (11.8%) | 9.5% to 14.7% |
| Russian | 52/600 (8.7%) | 92/600 (15.3%) | 12.7% to 18.4% |
The archived inference times measure ten suggestions on the local CPU after loading. The current screen computes twenty word candidates and performs the additional ending lookup; the archived times do not benchmark that whole operation. They exclude network, display, first model load and the 150 ms typing pause. Chinese word segmentation is included in its inference measurement. Cold requests can take longer. A new local timing check on 600 reused English development prefixes measured a median of 1.81 ms and a 95th percentile of 3.54 ms for twenty word candidates plus the extra lookup. This is a warm CPU measurement, not browser response time or an accuracy test.
The cache retains English and at most three additional models. The maximum combined size of these prepared R model objects is approximately 273.5 MiB. This includes the stored ending indexes but excludes the R process, Shiny, cached pages and temporary allocations; it is not total hosting memory. A language evicted from the cache loads again when selected.
Text is processed on the Shiny server for the active session. The app does not save entered phrases or send them to an external prediction API.
The next-word accuracy graphs apply to the unchanged word model. The supplementary ending lookup has not received an independent accuracy evaluation. Its training phrases are not an evaluation set.
The expansion makes twelve languages usable through the same interface, while preserving the English predictions. It does not establish high accuracy or full-sentence understanding. Korean currently has the weakest measured top-three result among the new prototypes, with limited training text and many unseen forms. That combination makes additional representative Korean data and a comparison of segmentation methods a sensible next experiment, rather than a proven solution.
Further work should compare a compact context-aware model with this baseline, test representative unseen messages in each language, obtain native-speaker review and calibrate confidence after model selection. The current language hold-outs have now been examined; after tuning, a fresh final test is needed. More language choices and 100% prediction availability do not imply 100% accuracy.
Author: Sonja Sahebzad
Utrecht, the Netherlands · Sonja Projects
This HTML was knitted from Multilingual_Guide.Rmd. The
project retains the .Rproj, app source, five-slide RStudio
Presenter .Rpres, rendered HTML, fixed protocols, download
checksums, per-case outcomes, CPU timings, checks and figure code. Model
inference requires R, Shiny, data.table, jsonlite and stringi; no GPU or
remote prediction API is needed. Open
Sonja PhraseFlow.Rproj, open app.R, then
select Run App.
Tatoeba contributors supplied the additional text under CC BY 2.0 FR. Training transformed it through normalization, segmentation and aggregation. The data credits and attribution lists retain contributor names and sentence URLs. Held-out answers remain private. The live app and five-slide presentation accompany this guide. The source repository includes the R app, editable guide and native Presenter source. Python and R research scripts are retained separately from the CPU app.
Sources: Tatoeba downloads and file definitions, official SwiftKey course archive, ICU word-boundary analysis, Chen and Goodman, 1996.