rikAI runs a translator and a reasoning model on the phone in your hand.
This page is how they were trained and what they score.
Nothing stays in place — みたい becomes two words,
the subject appears from nowhere. This is why translation takes a language model,
not a dictionary.
On the device
The pair in your pocket
Translator
RNF Translate 0.8B
Turns the sentence in front of you into English — and when the app
has already pinned down what a word means in this sentence, the translator is required
to honor it. That per-word glossary control is trained in, not bolted on.
LoRA + MO-GRPO two-stage training
6-bit bundled translator weights
Dictionary context guides translation
Reasoning
rikAI Reasoning 2B
Answers your questions about the sentence — briefed first with the
dictionary entries and grammar points the app matched, so it explains from facts
instead of vibes. Fine-tuned for exactly one job: answering from the facts it's handed.
Grounded in dictionary + grammar facts
5-bit bundled assistant · local Ask needs 4 GB memory
Earlier training experiments
Research results from earlier builds
These charts describe earlier research configurations. They are not measurements
of the current bundled models or guarantees of performance on your iPhone.
0.000COMET-QE, 660-sentence eval
0.000MetricX-24 quality
0%glossary adherence
0 msearlier Apple-silicon run
Scored blind by two neural judges and two string metrics on a
frozen 660-sentence conversational set in an earlier research run. Current release quality
and latency have not been revalidated by this website update.
One benchmark, every size
CyberAgent's CAT-Translate, untrainedthe same 0.8B, after rikAI's trainingits score, carried across the chart
COMET-QE neural judgeMetricX-24 neural judgechrF character overlapBLEU word overlap
Business Scene Dialogue — an external benchmark of everyday
conversation, run English into Japanese. Higher is better on every panel. The grey
line is the stock CAT-Translate family from CyberAgent, untrained; the blue square
is the trained model in this earlier experiment, so the blue stops at 0.8B. On all four metrics it
lands above the stock 1.4B — a model nearly twice its size. Only the desktop-class
3.3B stays ahead.
Does size buy quality? Stock CAT-Translate on everyday Japanese
0.8B0.000
3.3B0.000
COMET-QE over 1,938 everyday sentences — Tatoeba gold (438) plus
Tanaka conversational (1,500), plain prompts. Four times the parameters buys +0.003.
On everyday Japanese, size stopped being the bottleneck — so rikAI ships the 0.8B
and spends the savings on grounding.
What small buys — one sentence, translated
rikAI 0.8B · int8103 ms
7B flagship~940 ms
Earlier greedy-decoding measurements on Apple silicon using int8 weights.
The bundled translator now uses 6-bit weights; this chart is not current iPhone latency.
Earlier reasoning research
Taught to stick to the facts
These charts describe an earlier 0.8B reasoning experiment trained to answer from
supplied dictionary and grammar facts. They do not measure the current bundled 2B assistant.
AI explanations can still be wrong; verify answers against the source material.
Sticks to the facts it's given counterfactual swap test
11%
100%
Admits it when the facts are missing grounded abstention
~0%
93%
Invents synonyms the dictionary never said
33%
3%
stock 0.8B model
RNF Reasoning, after fine-tuning
On-device by default
Local models, with cloud Ask as a choice
The bundled translator uses 6-bit weights and the assistant uses 5-bit weights.
Translation and dictionary tools work offline. Local Ask requires at least 4 GB of memory;
lower-memory phones can use optional cloud Ask with permission and internet.
Settings → Model shows installed models and their sizes. If you allow Ask with ChatGPT,
questions, selected text, recent context and assistant preferences go through Cloudflare
to GPT-5 nano. Read the data-sharing details.