Container

Add New Languages for OCR

Install additional OCR language data in your self-hosted pdfRest API Toolkit Container.

Included and additional languages

The pdfRest API Toolkit Container includes English, Spanish, Japanese, and other language packs marked ✅ in the table below. Install additional languages, such as Polish (pol) and Hindi (hin), before requesting them.

This guide applies to self-hosted Container deployments that call the OCR PDF API Tool. This guide is not related to the hosted Cloud API service.

Use the languages parameter with POST /pdf-with-ocr-text to select installed languages. Supply ISO 639-2 three-letter codes, such as eng, spa, and jpn, or supported language variants, such as chi_sim. Separate multiple languages with commas: eng,spa. The default is eng. Existing supported language names remain accepted for backward compatibility; use the ISO 639-2 codes for new integrations.

If you are using an Enterprise Container image, exclude any references to LICENSE_KEY in the code samples below.

Download compatible training data

Tesseract provides downloadable language training data in the tessdata_best repository. Only tessdata_best packs are supported by this guide. Use the current compatible .traineddata files from that repository; older training-data formats are not interchangeable.

Each filename must match the language identifier: pol.traineddata is selected with languages=pol, and hin.traineddata with languages=hin. Download the raw file, not its GitHub HTML page.

Create a working directory and download Polish and Hindi packs:

mkdir -p pdfrest-ocr-languages/training-data
cd pdfrest-ocr-languages

curl --location --output training-data/pol.traineddata \
  https://raw.githubusercontent.com/tesseract-ocr/tessdata_best/main/pol.traineddata
curl --location --output training-data/hin.traineddata \
  https://raw.githubusercontent.com/tesseract-ocr/tessdata_best/main/hin.traineddata
chmod 644 training-data/*.traineddata

For repeatable deployments, pin the training-data URLs to a tested repository commit instead of main, and pin your pdfRest image to a tested version instead of latest.

Installing language packs

There are two options for adding additional language packs to a pdfRest Container deployment, you can either re-build the pdfRest image to include the language packs inside of the image, or you can mount the language packs to the deployment at runtime.

Re-build image to pre-install language packs

Adding the packs to your pdfRest Container image keeps them available whenever you recreate or scale the deployment. In the working directory, create a file named Dockerfile:

# syntax=docker/dockerfile:1
ARG PDFREST_IMAGE=pdfrest/pdf-api-toolkit
ARG PDFREST_TAG=latest
FROM ${PDFREST_IMAGE}:${PDFREST_TAG}

COPY --chmod=0644 training-data/*.traineddata /opt/tesseract/share/tessdata/

COPY --chmod makes the data readable without changing the base image's runtime user. The resulting image retains pdfRest's non-root execution configuration.

Build the image:

docker build --platform linux/amd64 \
  --build-arg PDFREST_IMAGE=pdfrest/pdf-api-toolkit \
  --build-arg PDFREST_TAG=latest \
  --tag pdfrest-with-ocr-languages:local .

For an existing Compose deployment, replace its image with pdfrest-with-ocr-languages:local, retain its environment, ports, and storage settings, and run docker compose up -d --force-recreate pdfrest. Rebuild the custom image and recreate the container after changing the packs. Restarting an old container does not switch it to a newly built image.

Mount language packs at runtime

You can keep the files on the host and mount them read-only into the container. Run this alternative from the same working directory instead of starting the derived image:

docker run --detach --name pdfrest-ocr --platform linux/amd64 \
  --publish 3000:3000 \
  --env LICENSE_KEY \
  --env PDFREST_SERVER_DOMAIN=http://localhost:3000 \
  --mount "type=bind,source=$(pwd)/training-data/pol.traineddata,target=/opt/tesseract/share/tessdata/pol.traineddata,readonly" \
  --mount "type=bind,source=$(pwd)/training-data/hin.traineddata,target=/opt/tesseract/share/tessdata/hin.traineddata,readonly" \
  pdfrest/pdf-api-toolkit:latest

Mount individual files, not the entire training-data directory, so all pre-installed language data stays available. Files must exist before startup and be readable by the container's non-root user.

For Docker Compose, save the following as compose.yaml beside training-data:

services:
  pdfrest:
    platform: linux/amd64
    image: pdfrest/pdf-api-toolkit:latest
    ports:
      - "3000:3000"
    environment:
      LICENSE_KEY: ${LICENSE_KEY:?Set LICENSE_KEY before starting Compose}
      PDFREST_SERVER_DOMAIN: http://localhost:3000
    volumes:
      - type: bind
        source: ./training-data/pol.traineddata
        target: /opt/tesseract/share/tessdata/pol.traineddata
        read_only: true
        bind:
          create_host_path: false
      - type: bind
        source: ./training-data/hin.traineddata
        target: /opt/tesseract/share/tessdata/hin.traineddata
        read_only: true
        bind:
          create_host_path: false

Start it with docker compose up -d. When you add mounts or replace training-data files, recreate the container with docker compose up -d --force-recreate pdfrest so it sees the current files. Installed-language availability is cached while pdfRest runs: even if you update an existing mounted file in place, restart the container before sending more OCR requests. For the Docker run alternative, recreate it with the updated mounts; docker restart pdfrest-ocr is sufficient only when the existing mounts still expose the updated files.

Verify with API requests

Wait until curl --include http://localhost:3000/up shows HTTP 200. Use image-based PDFs containing text in the language being tested. These requests require no client API key; licensing was configured when starting the container.

The Spanish and Japanese examples use pre-installed packs. After installing Polish, use languages=pol with a Polish scan to verify the additional pack in the same way.

Spanish

curl --request POST \
  http://localhost:3000/pdf-with-ocr-text \
  --header 'Accept: application/json' \
  --form 'file=@/path/to/spanish-scan.pdf' \
  --form 'languages=spa'

Japanese

curl --request POST \
  http://localhost:3000/pdf-with-ocr-text \
  --header 'Accept: application/json' \
  --form 'file=@/path/to/japanese-scan.pdf' \
  --form 'languages=jpn'

English and Spanish

curl --request POST \
  http://localhost:3000/pdf-with-ocr-text \
  --header 'Accept: application/json' \
  --form 'file=@/path/to/english-spanish-scan.pdf' \
  --form 'languages=eng,spa'

Select only the languages expected in the document. Requesting many languages can affect processing time, particularly Chinese, Japanese, and Korean.

A successful response includes outputId and outputUrl. Check that it has no unavailable-language warning, download outputUrl, and open the resulting PDF in a viewer. Search for known words and confirm that text can be selected and copied. A successful HTTP response alone does not establish recognition quality. In trial mode, account for any licensing-related output limitations.

Troubleshooting

SymptomWhat to check
A requested language is unavailableInstall its pack, match its filename and API code exactly, and confirm the file is at /opt/tesseract/share/tessdata/. Packs without a pre-installed check in the table require installation.
Some requested languages are missingpdfRest can process using the installed subset and return a warning identifying unavailable languages. Install the missing packs and retry; do not treat partial success as verification of every requested language.
None of the requested languages is availableThe request fails. Install at least the requested packs; pdfRest does not silently substitute English.
A new pack is still unavailableRestart after updates to existing mounted data. Recreate after image or mount changes, and ensure every deployment instance receives the same packs.
Container startup reports a mount problemConfirm each source is an existing file, not a directory. Mount individual files and keep the bundled data visible.
OCR fails after a pack is installedRedownload the raw tessdata_best file, check file permissions, and replace incomplete, incompatible, or HTML downloads. Recreate or restart the container as appropriate.
A downloaded code is rejectedCheck the compatible-language table below. A file in an upstream repository does not guarantee that its identifier is accepted by pdfRest.
Output is missing expected textTest a clear scan in the requested language, inspect response warnings, and confirm the output is searchable.

Compatible language packs

The table lists the tessdata_best language files whose identifiers are accepted unchanged by pdfRest. Each code corresponds to <code>.traineddata. A ✅ in the Pre-installed column identifies a pack included in the Container image; install unchecked packs before requesting them. Acceptance of an identifier does not verify that its pack is installed or that a particular document will be recognized accurately.

See the complete training-data repository and language catalog for the source files and language descriptions.

CodeLanguagePre-installed
afrAfrikaans
amhAmharic
araArabic
asmAssamese
azeAzerbaijani
aze_cyrlAzerbaijani - Cyrillic
belBelarusian
benBengali
bodTibetan
bosBosnian
breBreton
bulBulgarian
catCatalan; Valencian
cesCzech
chi_simChinese - Simplified✅
chi_sim_vertChinese - Simplified (vertical)✅
chi_traChinese - Traditional✅
chi_tra_vertChinese - Traditional (vertical)✅
cosCorsican
cymWelsh
danDanish
deuGerman✅
deu_latfGerman (Fraktur Latin)
divDhivehi
dzoDzongkha
ellGreek, Modern (1453-)
engEnglish✅
epoEsperanto
estEstonian
eusBasque
faoFaroese
fasPersian
finFinnish
fraFrench✅
fryWestern Frisian
glaScottish Gaelic
gleIrish
glgGalician
gujGujarati
hatHaitian; Haitian Creole
hebHebrew
hinHindi
hrvCroatian
hunHungarian
hyeArmenian
ikuInuktitut
indIndonesian
islIcelandic
itaItalian✅
ita_oldItalian - Old
javJavanese
jpnJapanese✅
jpn_vertJapanese (vertical)✅
kanKannada
katGeorgian
kat_oldGeorgian - Old
kazKazakh
khmCentral Khmer
kirKirghiz; Kyrgyz
korKorean✅
kor_vertKorean (vertical)
laoLao
latLatin
lavLatvian
litLithuanian
ltzLuxembourgish
malMalayalam
marMarathi
mkdMacedonian
mltMaltese
monMongolian
mriMaori
msaMalay
myaBurmese
nepNepali
nldDutch; Flemish✅
norNorwegian
ociOccitan (post 1500)
oriOriya
panPanjabi; Punjabi
polPolish
porPortuguese✅
pusPushto; Pashto
queQuechua
ronRomanian; Moldavian; Moldovan
rusRussian
sanSanskrit
sinSinhala; Sinhalese
slkSlovak
slvSlovenian
sndSindhi
spaSpanish; Castilian✅
spa_oldSpanish; Castilian - Old
sqiAlbanian
srpSerbian
srp_latnSerbian - Latin
sunSundanese
swaSwahili
sweSwedish
tamTamil
tatTatar
telTelugu
tgkTajik
thaThai
tirTigrinya
tonTonga
turTurkish
uigUighur; Uyghur
ukrUkrainian
urdUrdu
uzbUzbek
uzb_cyrlUzbek - Cyrillic
vieVietnamese
yidYiddish
yorYoruba

The table lists language identifiers accepted by pdfRest. Other packs in the upstream catalog may not be selectable through the languages parameter.