German PDFs fail OCR conversion in ways that generic, language-agnostic tools do not anticipate. Compound words split apart during recognition. Umlauts lose their dots. The Eszett character gets misread as a capital B. None of this is a vendor being careless. It is what happens when an OCR process built for general text meets a language with German’s specific structure. Foliage Solutions specialises in German source files precisely because these failures are well documented, predictable, and preventable, not because German OCR is inherently unsolvable.
This matters for LSPs handling German-source technical documentation, where the volume of scanned or image-based PDFs makes OCR quality a direct input into translation memory value, not just a formatting convenience, and where the specialists actually available to catch these errors manually are the same people better used on higher-value layout work.
Why compound words break during OCR specifically

German is a highly productive compounding language: words combine freely to form new, often very long single words that do not exist as separate lexicon entries. A term like Rechtsschutzversicherungsgesellschaften, a real and unremarkable German business term, is one word, not four, and a general-purpose OCR engine trained primarily on English or mixed-language text has no lexicon entry for it.
Recent benchmarking research quantifying OCR error types across languages found this specifically: deletion errors, where part of a word is dropped entirely, disproportionately shatter German compound words compared to other error types affecting English or French text. The OCR engine does not simply misread a compound. It frequently breaks it into fragments, sometimes losing a fragment outright, which changes the word’s meaning or removes it from the extracted text altogether.
This is not a font quality or scan resolution problem alone, though both make it worse. It is a structural mismatch between how German word formation works and how most OCR engines were trained to recognise word boundaries.
The same underlying compounding problem shows up in automated speech and language processing generally, not just OCR specifically. Research on German language models found that unknown compounds get broken into multiple recognised fragments, each fragment counted as a separate error, which compounds the error rate disproportionately compared to languages where words are more consistently separated by spaces. OCR faces a visual version of the same fundamental challenge: a long unbroken string of characters that does not match any single dictionary entry the engine was trained against.
Umlauts and the Eszett: a small character, a real failure point
Umlauts, ä, ö, ü, and the Eszett, ß, are frequently misrecognised even by otherwise capable OCR engines. The dots above an umlaut can be lost entirely, especially at smaller font sizes or lower scan resolutions, silently converting a word into a different, sometimes entirely unrelated word. The Eszett is commonly misread as a capital B, a substitution that is often visually plausible enough to pass an unreviewed automated check without triggering an obvious error.
This failure pattern is consistently documented across independent technical sources, not a rare edge case. It surfaces specifically in scanned documents, where the OCR engine has no digital font information to reference and is working purely from pixel recognition, which is precisely the situation most PDF-to-Word conversion work for translation actually involves.
A misread umlaut or Eszett does not just look wrong. In German, it can change a word’s actual meaning, which means an unreviewed OCR error here is a content accuracy problem, not merely a cosmetic one.
Modern OCR engines have improved substantially at general accuracy, and some now market themselves specifically around handling umlauts and the Eszett correctly. That improvement is real and worth acknowledging. It does not remove the need to actually check, since even tools that advertise strong German support can still produce inconsistent results depending on scan resolution, font, and document age, and a single missed instance in a technical term or proper noun is enough to introduce a real error into a translation memory.
What happens when this goes uncaught

An OCR error that survives into the translated document does not stay contained to that one document. If the converted text feeds into a CAT tool and becomes part of a translation memory, the corrupted segment gets stored as if it were correct language. Every future project drawing on that translation memory inherits the error, and because the mistake often looks like a plausible word rather than an obvious typo, it can go unnoticed for a long time.
This is the same underlying mechanism behind degraded fuzzy match rates on later, unrelated projects: a translation memory contaminated months earlier by an OCR error nobody caught at the time. The cost of a German-specific OCR failure is rarely paid on the project where the error occurred. It is paid later, quietly, as an LSP’s own team wonders why match rates on a long-standing client relationship have started to decline.
Consider a technical specification scanned from a German equipment manufacturer, containing several instances of long compound terms specific to that industry. A generic OCR pass splits one recurring compound into two separate words partway through the document, inconsistently, sometimes catching it correctly and sometimes not. The inconsistency itself is the problem: a translation memory now contains two different segments for what should be the same term, and future projects referencing this manufacturer’s documentation will match against whichever version happens to be closer, degrading consistency across the client relationship without an obvious single point of failure to trace it back to.
Text expansion planning is a separate, compounding problem
German text typically expands 20 to 30 percent when translated from English, and this creates a second, independent risk on top of the OCR-specific issues above. A layout planned around the source text length, without accounting for German expansion, produces overflowing text boxes and broken tables once translation is complete.
Most teams discover this after translation has already started, when it is far more expensive to fix than if expansion had been planned into the brief from the beginning. A German-aware DTP process identifies expansion risk during file preparation, before a single word has been translated, so layout adjustments happen proactively rather than as an emergency fix under deadline pressure.
This risk compounds specifically with the OCR issues covered above, rather than sitting alongside them as an unrelated concern. A document that already required manual correction for split compound words or misread umlauts is exactly the kind of document where nobody wants to discover a second, separate layout problem after translation. Catching expansion risk during the same file-preparation pass that catches OCR issues means one review cycle instead of two, and one delivery that actually accounts for everything German translation changes about the source file.
What a German-aware OCR process actually checks
Four checks separate a genuinely German-aware OCR process from a generic one applied to German content. Compound word integrity: verifying that long compound terms were not split or truncated during recognition, checked against the actual source rather than assumed correct because the output looks plausible. Umlaut and Eszett accuracy: specifically reviewing these characters rather than trusting a general accuracy score that can mask a systematic, repeated error. Layout expansion planning: identifying where German’s typical 20 to 30 percent expansion will create problems before translation begins, not after. And CAT-tool round-trip testing: confirming the converted file’s structure survives export, translation, and re-import cleanly, since a structurally sound-looking file can still corrupt a translation memory if tag structure was not preserved.
None of these four checks are exotic. They are specific, and specificity is exactly what a general-purpose OCR process, applied to German without adjustment, does not provide.
Why this affects capacity planning, not just quality

Every hour a specialist DTP team spends manually correcting German OCR errors, rebuilding a split compound word, fixing a misread umlaut, adjusting a layout for expansion discovered too late, is an hour not spent on higher-value, higher-margin work. A German-aware OCR process built to avoid these failures in the first place is not just a quality improvement. It is capacity an LSP’s own specialists get back.
This is worth quantifying in real terms, even roughly. A batch of ten German technical documents, each requiring an hour of manual correction for OCR and layout issues that a properly built process would have avoided, is ten specialist hours redirected from lower-value cleanup toward the InDesign, e-learning, and complex-layout work that actually justifies a specialist’s rate.
What Foliage does for German source files
Every German OCR conversion is checked specifically for compound word integrity, umlaut and Eszett accuracy, and expansion risk, before the file moves to translation. This is not a general OCR process with a language setting changed. It is built around the documented, specific ways German text behaves differently from English during automated recognition.
Talk to Foliage Solutions about a German PDF that has caused problems before. We can show you what a clean conversion actually looks like, compound words intact, umlauts correct, and layout already planned for what German translation does to the text.

