Clean a translation memory for AI translation
The quickest way to make an AI translate like your team is to show it your own past translations: as examples in a prompt, as files attached to a custom assistant, or retrieved from an index for every new segment. The model copies what it's shown, including your TM's contradictions, outdated terms and old machine output. Clean the TM before it becomes the model's reference.
━━ the problem
When one source sentence has three translations in your TM, a retrieval step can hand the model any of them, or all three. Keeping one agreed version per source gives it one answer to follow.
A pair from a 2017 import or an earlier MT pass looks as authoritative to a model as your reviewed work. If it's in the examples, the model imitates it.
A prompt or a retrieval result holds a limited number of examples. URLs, IDs, placeholder-only strings and untranslated pairs take slots real sentences should have, and broken characters or dropped placeholders teach the model the same mistakes.
━━ how TM Cleaner handles it
Upload the TMX and review before you clean: decide the version to keep in each duplicate group (or let most common, newest or your preferred translators decide), filter out translators and dates you don't want imitated, and switch on the quality checks. Then download the cleaned TMX, or Excel and TSV files for tools that take a spreadsheet or plain text.
- ✓one translation per source: you pick it in the review, or most common, newest or your preferred translators decide
- ✓translator and date filters drop old imports and anyone the model shouldn't imitate
- ✓machine-translated pairs are recognised by their author stamp (MT, Google, DeepL and any pattern you add) and removed
- ✓placeholder-mismatch and broken-character checks keep bad habits out of the examples; broken characters can be repaired instead
- ✓protected terms: pairs where a product or brand name got translated are removed, so names stay as they are
- ✓Excel export lists every kept pair with its translator and date; TSV gives plain source and target for scripts and retrieval pipelines
━━ frequently asked
should I clean a TM before using it with an LLM?+
yes, if you use it as examples or retrieval data. models imitate what they're shown: contradictory, outdated or machine-made examples produce contradictory, outdated or machine-like output. a cleaned TM is also smaller, so more of the examples that matter fit.
how is this different from cleaning a TM for MT training?+
training changes the model itself, so volume and balance matter: duplicates over-weight examples. as a reference, the model sees a handful of pairs at a time, so what matters most is that each pair is the version you want copied. same cleaning, different emphasis: spend your time on the duplicate choices and the translator and date filters.
which file should I use?+
the cleaned TMX if your CAT tool or AI platform reads TMX. the TSV (source and target, tab-separated) for scripts and retrieval pipelines. the Excel file if people will review the examples first, or your tool takes a spreadsheet.
does cleaning change the text of my translations?+
no, unless you turn on repair broken characters, which fixes encoding damage like “café” in place. everything else keeps or removes whole pairs; nothing is rewritten or re-aligned.
is my TM sent to an AI model?+
no. the cleaning rules are deterministic and your file is never sent to a language model. an upload is deleted when cleaning finishes (unless you ask to keep it for 24 hours) or within about a day if you only preview it, and download links expire after one hour.
━━ related pages
free · preview without an account · files up to 2 GB · try the demo