Add New Languages for OCR
Included and additional languages
The pdfRest API Toolkit Container includes English, Spanish, Japanese, and other language packs marked ✅ in the table below. Install additional languages, such as Polish (pol) and Hindi (hin), before requesting them.
This guide applies to self-hosted Container deployments that call the OCR PDF API Tool. This guide is not related to the hosted Cloud API service.
Use the languages parameter with POST /pdf-with-ocr-text to select installed languages. Supply ISO 639-2 three-letter codes, such as eng, spa, and jpn, or supported language variants, such as chi_sim. Separate multiple languages with commas: eng,spa. The default is eng. Existing supported language names remain accepted for backward compatibility; use the ISO 639-2 codes for new integrations.
LICENSE_KEY in the code samples below.Download compatible training data
Tesseract provides downloadable language training data in the tessdata_best repository. Only tessdata_best packs are supported by this guide. Use the current compatible .traineddata files from that repository; older training-data formats are not interchangeable.
Each filename must match the language identifier: pol.traineddata is selected with languages=pol, and hin.traineddata with languages=hin. Download the raw file, not its GitHub HTML page.
Create a working directory and download Polish and Hindi packs:
mkdir -p pdfrest-ocr-languages/training-data
cd pdfrest-ocr-languages
curl --location --output training-data/pol.traineddata \
https://raw.githubusercontent.com/tesseract-ocr/tessdata_best/main/pol.traineddata
curl --location --output training-data/hin.traineddata \
https://raw.githubusercontent.com/tesseract-ocr/tessdata_best/main/hin.traineddata
chmod 644 training-data/*.traineddata
For repeatable deployments, pin the training-data URLs to a tested repository commit instead of main, and pin your pdfRest image to a tested version instead of latest.
Installing language packs
There are two options for adding additional language packs to a pdfRest Container deployment, you can either re-build the pdfRest image to include the language packs inside of the image, or you can mount the language packs to the deployment at runtime.
Re-build image to pre-install language packs
Adding the packs to your pdfRest Container image keeps them available whenever you recreate or scale the deployment. In the working directory, create a file named Dockerfile:
# syntax=docker/dockerfile:1
ARG PDFREST_IMAGE=pdfrest/pdf-api-toolkit
ARG PDFREST_TAG=latest
FROM ${PDFREST_IMAGE}:${PDFREST_TAG}
COPY --chmod=0644 training-data/*.traineddata /opt/tesseract/share/tessdata/
COPY --chmod makes the data readable without changing the base image's runtime user. The resulting image retains pdfRest's non-root execution configuration.
Build the image:
docker build --platform linux/amd64 \
--build-arg PDFREST_IMAGE=pdfrest/pdf-api-toolkit \
--build-arg PDFREST_TAG=latest \
--tag pdfrest-with-ocr-languages:local .
For an existing Compose deployment, replace its image with pdfrest-with-ocr-languages:local, retain its environment, ports, and storage settings, and run docker compose up -d --force-recreate pdfrest. Rebuild the custom image and recreate the container after changing the packs. Restarting an old container does not switch it to a newly built image.
Mount language packs at runtime
You can keep the files on the host and mount them read-only into the container. Run this alternative from the same working directory instead of starting the derived image:
docker run --detach --name pdfrest-ocr --platform linux/amd64 \
--publish 3000:3000 \
--env LICENSE_KEY \
--env PDFREST_SERVER_DOMAIN=http://localhost:3000 \
--mount "type=bind,source=$(pwd)/training-data/pol.traineddata,target=/opt/tesseract/share/tessdata/pol.traineddata,readonly" \
--mount "type=bind,source=$(pwd)/training-data/hin.traineddata,target=/opt/tesseract/share/tessdata/hin.traineddata,readonly" \
pdfrest/pdf-api-toolkit:latest
Mount individual files, not the entire training-data directory, so all pre-installed language data stays available. Files must exist before startup and be readable by the container's non-root user.
For Docker Compose, save the following as compose.yaml beside training-data:
services:
pdfrest:
platform: linux/amd64
image: pdfrest/pdf-api-toolkit:latest
ports:
- "3000:3000"
environment:
LICENSE_KEY: ${LICENSE_KEY:?Set LICENSE_KEY before starting Compose}
PDFREST_SERVER_DOMAIN: http://localhost:3000
volumes:
- type: bind
source: ./training-data/pol.traineddata
target: /opt/tesseract/share/tessdata/pol.traineddata
read_only: true
bind:
create_host_path: false
- type: bind
source: ./training-data/hin.traineddata
target: /opt/tesseract/share/tessdata/hin.traineddata
read_only: true
bind:
create_host_path: false
Start it with docker compose up -d. When you add mounts or replace training-data files, recreate the container with docker compose up -d --force-recreate pdfrest so it sees the current files. Installed-language availability is cached while pdfRest runs: even if you update an existing mounted file in place, restart the container before sending more OCR requests. For the Docker run alternative, recreate it with the updated mounts; docker restart pdfrest-ocr is sufficient only when the existing mounts still expose the updated files.
Verify with API requests
Wait until curl --include http://localhost:3000/up shows HTTP 200. Use image-based PDFs containing text in the language being tested. These requests require no client API key; licensing was configured when starting the container.
The Spanish and Japanese examples use pre-installed packs. After installing Polish, use languages=pol with a Polish scan to verify the additional pack in the same way.
Spanish
curl --request POST \
http://localhost:3000/pdf-with-ocr-text \
--header 'Accept: application/json' \
--form 'file=@/path/to/spanish-scan.pdf' \
--form 'languages=spa'
Japanese
curl --request POST \
http://localhost:3000/pdf-with-ocr-text \
--header 'Accept: application/json' \
--form 'file=@/path/to/japanese-scan.pdf' \
--form 'languages=jpn'
English and Spanish
curl --request POST \
http://localhost:3000/pdf-with-ocr-text \
--header 'Accept: application/json' \
--form 'file=@/path/to/english-spanish-scan.pdf' \
--form 'languages=eng,spa'
Select only the languages expected in the document. Requesting many languages can affect processing time, particularly Chinese, Japanese, and Korean.
A successful response includes outputId and outputUrl. Check that it has no unavailable-language warning, download outputUrl, and open the resulting PDF in a viewer. Search for known words and confirm that text can be selected and copied. A successful HTTP response alone does not establish recognition quality. In trial mode, account for any licensing-related output limitations.
Troubleshooting
| Symptom | What to check |
|---|---|
| A requested language is unavailable | Install its pack, match its filename and API code exactly, and confirm the file is at /opt/tesseract/share/tessdata/. Packs without a pre-installed check in the table require installation. |
| Some requested languages are missing | pdfRest can process using the installed subset and return a warning identifying unavailable languages. Install the missing packs and retry; do not treat partial success as verification of every requested language. |
| None of the requested languages is available | The request fails. Install at least the requested packs; pdfRest does not silently substitute English. |
| A new pack is still unavailable | Restart after updates to existing mounted data. Recreate after image or mount changes, and ensure every deployment instance receives the same packs. |
| Container startup reports a mount problem | Confirm each source is an existing file, not a directory. Mount individual files and keep the bundled data visible. |
| OCR fails after a pack is installed | Redownload the raw tessdata_best file, check file permissions, and replace incomplete, incompatible, or HTML downloads. Recreate or restart the container as appropriate. |
| A downloaded code is rejected | Check the compatible-language table below. A file in an upstream repository does not guarantee that its identifier is accepted by pdfRest. |
| Output is missing expected text | Test a clear scan in the requested language, inspect response warnings, and confirm the output is searchable. |
Compatible language packs
The table lists the tessdata_best language files whose identifiers are accepted unchanged by pdfRest. Each code corresponds to <code>.traineddata. A ✅ in the Pre-installed column identifies a pack included in the Container image; install unchecked packs before requesting them. Acceptance of an identifier does not verify that its pack is installed or that a particular document will be recognized accurately.
See the complete training-data repository and language catalog for the source files and language descriptions.
The table lists language identifiers accepted by pdfRest. Other packs in the upstream catalog may not be selectable through the languages parameter.