zmazz.dev

Project

Low-Resource Language Translation

Product · Completed · 2024

Capability-level case study

Fine-tuning and corpus work for translation on a small-corpus language, combining data design with LLM and encoder-decoder training.

Role
AI engineer / researcher
Context
Professional
Capabilities
Translation, Fine-tuning, Evaluation

Overview

Low-resource translation is mostly a data problem wearing a model problem’s clothes. This engagement combined corpus work — collection, cleaning, augmentation — with training of encoder-decoder and LLM-based translation systems, plus a small interface so linguists could see errors rather than scores only.

I led the engineering and training strategy. Linguistic validation and any official language-policy decision were not mine.

Related research on this site: Language verY Rare for All.

Problem

A language with a thin parallel corpus still needed usable translation. Off- the-shelf multilingual models were a starting point, not a product.

Approach

Corpus design first; then a mix of continued training, LLM/RAG-assisted workflows, and later RL-style fine-tuning experiments. Exact partner names, BLEU tables and “first translator for X” claims stay off this stealth page.

Tradeoffs and limitations

Small corpora overfit. Automatic metrics lie in ways a native speaker notices immediately. I do not publish a quality number here. The working paper is the right place for methods; this page is the engineering case.

Exact start and end months are omitted (stealth). Display year is 2024.

Confidentiality

Stealth / capability-only. Government and NGO partners are not named.