What Makes a Good Translation? Accuracy, Fluency, Purpose, and Evidence
A practical framework for judging translation quality across meaning, language, terminology, audience, presentation, and intended use.

A good translation preserves the source's intended meaning and effect, reads as valid target-language writing, and works for the audience and purpose it was commissioned to serve. It must also remain complete, consistent, and usable in its final format. Fluency alone cannot prove any of those things.
That definition changes translation review from a taste contest into an evidence-based decision. Before reviewing, specify the use case and the consequences of error. Then evaluate separate quality dimensions, record defects by type and severity, and apply a release threshold that reflects the risk.
Download the translation quality scorecard to define the criteria, sample, evidence, and acceptance decision for your own project.
Start with purpose, not a universal score
The same target text can be good enough for one purpose and unacceptable for another. A private reading aid may tolerate awkward phrasing if the meaning is recoverable. A public novel needs sustained voice and editorial polish. A safety instruction can fail because one modal verb changed, even when every sentence sounds natural.
The European Commission describes a fit-for-purpose translation as one that carries the same meaning and effect while reading naturally for the intended readership. Its current translation quality guidance also ties the depth of review to the consequences and likelihood of errors.
Write a short quality specification before evaluating:
| Decision | Question to answer | Example evidence |
|---|---|---|
| Purpose | What will readers do with the translation? | Personal reading, internal research, publication, regulated use |
| Audience | Who will read it, in which locale? | Age, domain knowledge, region, accessibility needs |
| Content risk | What would happen if a passage were wrong? | Confusion, lost narrative clue, financial loss, unsafe action |
| Required controls | Which decisions must remain stable? | Names, glossary, formality, units, citations, layout |
| Acceptance rule | What causes accept, revise, or reject? | No blockers; majors corrected and rechecked |
Without this specification, two reviewers can reach opposite conclusions because they are silently judging different products.
Evaluate seven separate dimensions
The MQM Core error typology separates translation problems into terminology, accuracy, linguistic conventions, style, locale conventions, audience appropriateness, and design and markup. These are useful review lenses because a translation can pass one while failing another.
1. Accuracy: does the target preserve the source meaning?
Compare propositions, relationships, negation, modality, sequence, quantities, and references. Look for additions, omissions, distortions, and untranslated content.
Consider this source instruction:
The valve must remain closed until the pressure falls below 2 bar.
Now imagine a fluent target sentence whose meaning, translated back for inspection, is:
The valve may remain closed until the pressure falls below 2 bar.
The grammar can be impeccable while the translation is unusable. Must became may, so an obligation became an option. Accuracy review has to compare source and target; reading the target alone cannot expose the defect.
2. Linguistic conventions: is it valid target-language writing?
Check grammar, spelling, punctuation, syntax, agreement, and mechanical correctness. Then read connected passages for cohesion. A sentence-by-sentence review can miss broken reference chains, inconsistent tense, or paragraphs that do not flow.
Naturalness is evidence for this dimension, not a substitute for accuracy. A reviewer who only asks whether the prose sounds native may approve confident mistranslations.
3. Terminology: are controlled concepts named correctly and consistently?
Search names, places, titles, product terms, domain vocabulary, acronyms, and recurring phrases across the complete target. Compare every variant with the approved glossary or project decision.
Consistency is not blind repetition. The same source word may require different target words in different senses, while several source variants may refer to one controlled concept. The book translation terminology workflow explains how to record those decisions and rejected alternatives.
4. Style and register: does the voice fit the source and situation?
Check formality, tone, rhythm, genre conventions, sentence density, characterization, humor, and editorial style. In fiction, a grammatically correct line can still fail because a child suddenly sounds bureaucratic or two distinct characters acquire the same voice.
Do not reduce style review to “I would phrase it differently.” A preference becomes a defect only when it violates the source, the agreed specification, target-language conventions, or an established style decision.
5. Locale and audience: does it work for the intended readers?
Review quotation marks, date and number formats, units, currencies, addresses, name order, directionality, cultural knowledge, reading level, and market-specific conventions. Audience review also asks whether an explanation is missing or whether an adaptation assumes knowledge the target reader does not have.
This is where translation and localization overlap, but the test remains concrete: can this audience understand and use the delivered text as intended?
6. Completeness: is every required unit present once and in order?
Check chapters, headings, paragraphs, table rows, captions, footnotes, endnotes, links, citations, appendices, and front and back matter. Search for source-language residue, duplicated boundary text, and truncated openings or endings.
Completeness often deserves its own release gate even when an evaluation framework groups omissions under accuracy. A missing chapter is not balanced out by thousands of excellent sentences elsewhere.
7. Design and markup: does the final artifact remain usable?
Open the actual PDF, EPUB, DOCX, subtitle file, or web page. Inspect reading order, navigation, tables, figures, headings, links, line breaks, embedded fonts, right-to-left behavior, and text clipping. The target words can be correct while the product fails because a table became unreadable or EPUB notes no longer link back.
The real translation samples illustrate why layout, scripts, tables, and reading order belong in translation quality rather than in a cosmetic afterthought.
Separate error type from severity
Error type tells you what failed. Severity tells you what the failure does to the intended use. Keep those judgments separate.
| Severity | Operational meaning | Example | Decision |
|---|---|---|---|
| Critical | Makes the deliverable unfit for purpose or creates serious harm | Reversed safety instruction; exposed confidential content | Reject immediately |
| Major | Materially changes meaning, reliability, or usability | Missing paragraph; recurring wrong name; broken table | Revise before acceptance and search for the pattern |
| Minor | Limited local effect; purpose remains achievable | Isolated punctuation or non-critical style defect | Correct when required; track recurring patterns |
| Neutral | Acceptable variation or reviewer note | A defensible alternative not prohibited by the specification | Do not count as a failure |
MQM's scoring guidance uses error types and severity levels to represent risk and produce configurable scores. The useful principle is not a universal magic number. It is that one severe defect should not disappear inside an average.
Review with evidence instead of impressions
A defensible evaluation leaves an audit trail. For each sampled passage or defect, record:
- the source and target location;
- the exact source and current target text;
- the error type and severity;
- the specification or evidence behind the judgment;
- the proposed correction and owner; and
- whether the final correction was independently rechecked.
Use at least two passes. The first compares source and target for completeness and meaning. The second reads the target as target-language content for coherence, style, and audience fit. Add terminology searches and format checks across the complete deliverable rather than limiting them to the language sample.
The ISO 5060:2024 overview describes analytic evaluation using error types and penalty points, and discusses evaluator competence and sampling. The downloadable scorecard here applies those general ideas as a lightweight project tool; using it is not a claim of ISO conformity.
Sample by risk, then expand when failures cluster
When complete human review is impractical, do not choose only easy or random pages. Force difficult content into the sample:
- The opening, ending, and every structural boundary.
- The densest passages for names, terms, numbers, and citations.
- Dialogue, humor, idioms, ambiguity, and high-register prose.
- Every special structure: tables, notes, figures, lists, equations, and links.
- Any instruction or claim whose failure has material consequences.
- Sections translated under different settings, models, or source revisions.
Expand the review when an error suggests a system. One wrong character name triggers a complete search for that name and its variants. One missing paragraph triggers structural comparison around every processing boundary. One flattened idiom triggers an idiom-focused review, not just a correction to the detected line.
For an operational whole-file process, use the book and long-document QA checklist. That checklist answers how to inspect the deliverable; this framework answers what quality means and whether the evidence is sufficient for the release decision.
Use accept, revise, or reject thresholds
Do not end with a percentage and no decision. Set three explicit outcomes:
- Accept: no critical errors, no unresolved majors, required complete-file checks passed, and the sample supports the intended use.
- Revise: the deliverable is recoverable, but material or systematic defects need correction and targeted reinspection.
- Reject: the wrong source or language was used, required content is missing at scale, the output is unusable, or the workflow cannot supply evidence appropriate to the risk.
Thresholds should be stricter for public, contractual, financial, clinical, legal, or safety-critical content. A high score from an automatic metric does not override a known blocker.
A practical quality gate
Before accepting a translation, confirm all of the following:
- The source version, purpose, audience, locale, and output format are recorded.
- Meaning, completeness, terminology, language, style, locale, audience, and presentation were reviewed as separate dimensions.
- Every critical or major issue has an owner, a correction, and a final recheck.
- Controlled names, terms, numbers, and do-not-translate items were searched across the complete target.
- The final artifact works in the readers or applications the audience will use.
- Sampling covered high-risk and structurally difficult content.
- The decision is accept, revise, or reject—not merely “looks good.”
If you need a complete translated draft before this review, BookTranslator can process books and documents from their original files. Its context-aware translation and glossary controls can help keep long-form decisions together, but they do not replace source comparison, target-language editing, or the risk-based release gate above.





