Clean Phrase TMX exports
Phrase (formerly Memsource) makes it easy to accumulate large cloud TMs shared across many linguists. That scale is exactly what breeds duplicate variants and machine-pretranslated noise. Export to TMX and clean it here.
━━ the problem
With dozens of linguists writing to one Phrase TM, the same source ends up with a dozen slightly different targets. Consensus (most-popular) or a trusted-reviewer ranking picks a single authoritative variant per segment.
Phrase's pretranslation and MTQE steps write machine-origin units into the TM. For a clean human-quality TM — or MT training data — those need to go. Filter by creationid/changeid origin signature.
URLs, numbers, tags and placeholders counted as 'words' inflate your Phrase analysis and leverage reports without producing any real match value. Strip them as structural junk.
━━ how TM Cleaner handles it
Upload the exported TMX, preview the duplicate and junk samples with projected size reduction, and download a cleaned TM ready to re-import into Phrase.
- ✓consensus or trusted-reviewer dedup for TMs written by many linguists
- ✓filter MT-pretranslated units by origin so your TM stays human-quality
- ✓author and date filters let you keep only a chosen reviewer's contributions or a date window
- ✓TSV export option for spot-checking source/target pairs in a spreadsheet
━━ frequently asked
is Phrase the same as Memsource?+
yes — Memsource rebranded to Phrase. TMX exports from either import and clean here identically.
can I keep only one reviewer's translations?+
yes. the author filter (by changeid) lets you keep only chosen contributors, or exclude specific ones. combine with a date range to isolate, say, this year's reviewed segments.
how big a TM can I clean?+
up to 2 GB per file on the paid tier. the engine streams the file so memory stays bounded even on very large cloud-TM exports.
━━ also clean
free tier · pay-as-you-go after · no subscription · see pricing