◇ for LLM reference, RAG and few-shot prompts

Clean a translation memory for AI translation

The quickest way to make an AI translate like your team is to show it your own past translations: as examples in a prompt, as files attached to a custom assistant, or retrieved from an index for every new segment. The model copies what it's shown, including your TM's contradictions, outdated terms and old machine output. Clean the TM before it becomes the model's reference.

━━ the problem

conflicting translations confuse the model

When one source sentence has three translations in your TM, a retrieval step can hand the model any of them, or all three. Keeping one agreed version per source gives it one answer to follow.

old terms and machine output get copied

A pair from a 2017 import or an earlier MT pass looks as authoritative to a model as your reviewed work. If it's in the examples, the model imitates it.

junk wastes the examples you can fit

A prompt or a retrieval result holds a limited number of examples. URLs, IDs, placeholder-only strings and untranslated pairs take slots real sentences should have, and broken characters or dropped placeholders teach the model the same mistakes.

━━ how TM Cleaner handles it

Upload the TMX and review before you clean: decide the version to keep in each duplicate group (or let most common, newest or your preferred translators decide), filter out translators and dates you don't want imitated, and switch on the quality checks. Then download the cleaned TMX, or Excel and TSV files for tools that take a spreadsheet or plain text.

━━ frequently asked

should I clean a TM before using it with an LLM?+

yes, if you use it as examples or retrieval data. models imitate what they're shown: contradictory, outdated or machine-made examples produce contradictory, outdated or machine-like output. a cleaned TM is also smaller, so more of the examples that matter fit.

how is this different from cleaning a TM for MT training?+

training changes the model itself, so volume and balance matter: duplicates over-weight examples. as a reference, the model sees a handful of pairs at a time, so what matters most is that each pair is the version you want copied. same cleaning, different emphasis: spend your time on the duplicate choices and the translator and date filters.

which file should I use?+

the cleaned TMX if your CAT tool or AI platform reads TMX. the TSV (source and target, tab-separated) for scripts and retrieval pipelines. the Excel file if people will review the examples first, or your tool takes a spreadsheet.

does cleaning change the text of my translations?+

no, unless you turn on repair broken characters, which fixes encoding damage like “café” in place. everything else keeps or removes whole pairs; nothing is rewritten or re-aligned.

is my TM sent to an AI model?+

no. the cleaning rules are deterministic and your file is never sent to a language model. an upload is deleted when cleaning finishes (unless you ask to keep it for 24 hours) or within about a day if you only preview it, and download links expire after one hour.

━━ related pages

clean your tmx now →

free · preview without an account · files up to 2 GB · try the demo