Agent-friendly AI localization for projects that want translations as normal pull requests.
Localize Pipeline detects changed source strings, translates the matching target locale entries, runs deterministic and AI review checks, and opens a PR that your team can inspect before merge. It runs in your CI or on your own server with your model provider and your credentials.
Repository: https://github.com/bisq-network/localize-pipeline
Rendered docs: https://bisq-network.github.io/localize-pipeline/
- Runs in your infrastructure. Use GitHub Actions, local CLI runs, or a Docker Compose cron job.
- Bring your own model. AISuite is the default provider abstraction. Bare
OpenAI model names and explicit AISuite names such as
openai:gpt-4o-miniare supported. - Zero data egress option. Point
api_base_urlat a local OpenAI-compatible endpoint such as Ollama and keep strings inside your infrastructure. - Reviewable output. Translations are committed to a branch and opened as a pull request.
- Follow through on reviews. Optional self-hosted Guardian applies checked corrections and proposes pipeline fixes for recurring translation problems.
- Format-aware. Java
.propertiesand JSON files are built in. Mixed-format projects use a profile list. - Reusable by other projects. The stable surfaces are the
localizeCLI,localize.core,localize.formats, andlocalize.providers. - Agent discoverable.
llms.txt, examples, profiles, and docs point agents to stable commands and module boundaries.
Generate a config from an existing repository. localize init looks for common
Java .properties and JSON layouts, detects target locales, and writes a safe
dry-run config:
python3 -m venv venv
./venv/bin/pip install -e .
localize init
localize check --config config.yaml
localize doctor --config config.yaml
localize smoke --config config.yaml
localize run --dry-run --config config.yamlIf your localization files live outside the detected folder, pass the folder explicitly:
localize init --input-folder path/to/i18nFor JSON:
localize init --input-folder path/to/i18n --localization-format jsonFor JSON stored as locales/en/messages.json and
locales/de/messages.json:
localize init \
--input-folder locales \
--localization-format json \
--localization-layout locale_directoryFor a mixed project:
localize init \
--input-folder path/to/i18n \
--localization-profile java_properties:suffix \
--localization-profile json:locale_directoryThen run a real translation by setting dry_run: false in config.yaml and
providing credentials:
localize validate --config config.yaml
localize run --config config.yamlSet OPENAI_API_KEY for OpenAI-backed runs, or set api_base_url in
config.yaml for a local/OpenAI-compatible endpoint. See
docs/environment-variables.md for the full
Docker, local-runner, and Action environment reference.
Add .github/workflows/translate.yml:
name: Translate
on:
push:
branches: [main]
workflow_dispatch: {}
permissions:
contents: write
pull-requests: write
jobs:
translate:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5
with:
fetch-depth: 0
- uses: bisq-network/localize-pipeline@v0.1.21
with:
config-file: config.yaml
openai-api-key: ${{ secrets.OPENAI_API_KEY }}The action translates only files changed since the configured diff base. For a
new locale, create and review the initial translation locally first, merge that
baseline, then let the action maintain incremental source-string updates.
process-all-files: true is still available for controlled full scans, but it
should not be the default operating mode for normal CI.
Before enabling live PR creation in a target repo, add the OPENAI_API_KEY
repository secret and set the repository's Actions workflow permissions to
allow write tokens and workflow-created pull requests. The workflow-level
contents: write and pull-requests: write permissions are necessary but not
sufficient if repository or organization settings keep the default token
read-only.
Run live translation from trusted push or workflow_dispatch workflows. For
pull request checks, use dry-run mode with PR creation disabled and no model or
signing secrets. Do not run live translation from pull_request_target or from
workflow_run events triggered by pull request builds.
Full guide: docs/github-action.md.
Start with the rendered developer guide at https://bisq-network.github.io/localize-pipeline/. The source lives in docs/index.html. Agents should start with the agent-readable index in llms.txt, then use the developer guide for setup, configuration, CLI commands, and supported localization formats.
Single-format projects use:
localization_format: "json"
localization_layout:
id: "locale_directory"
source_locale: "en"Suffix layouts can set base_name when the source file also carries its locale:
localization_layout:
id: "suffix"
base_name: "Messages"
source_locale: "en"
placeholder_profile: "java-indexed" # opt-in support for %0, %1, ...Mixed-format projects use profiles:
localization_formats:
- id: "java_properties"
layout: "suffix"
- id: "json"
layout:
id: "locale_directory"
source_locale: "en"Each profile owns matching files for source lookup, parsing, validation, prompt
construction, serialization, quality gates, semantic review, and publishing.
Singular localization_format configs remain supported for existing projects.
Key settings:
| Setting | Purpose |
|---|---|
target_project_root |
Repository that contains the localization files. |
input_folder |
Localization folder, absolute or relative to target_project_root. |
localization_format |
Built-in format id for single-format projects. |
localization_layout |
suffix, locale_directory, or locale_filename; suffix layouts optionally accept base_name. |
localization_formats |
Profile list for mixed-format projects. |
translation_source |
git or transifex. New projects usually start with git. |
placeholder_profile |
standard by default; use java-indexed for %0, %1, ... runtime tokens. |
model_provider |
aisuite by default; openai_compatible is the direct SDK fallback. |
model_name, review_model_name |
Translation and review models. |
review_reasoning_effort |
Optional reasoning effort for the holistic review model. |
semantic_review.reasoning_effort |
Optional reasoning effort for the semantic review model. |
api_base_url |
OpenAI-compatible endpoint, for example Ollama. |
supported_locales |
Target locales. |
project_context |
Product/domain context injected into prompts. |
brand_technical_glossary |
Terms that must not be translated. |
translation_glossary_enforcement |
exact by default; use prompt-only for preferred lemmas that require grammatical inflection. |
ignore_key_patterns |
Python regexes matched against adapter keys; matching keys are copied from source and never translated. |
style_rules |
Locale-specific writing rules. |
Examples:
- config.example.yaml is the minimal generic starter.
- examples/generic-java-properties shows Java
.properties. - examples/generic-json shows JSON.
- profiles/bisq is the production Bisq profile with richer style and semantic QA rules.
- profiles/bisq-mobile mirrors the production mobile profile shape without secrets.
localize formats
localize init
localize check --config config.yaml
localize doctor --config config.yaml
localize smoke --config config.yaml
localize validate --config config.yaml
localize run --dry-run --config config.yaml
localize run --config config.yaml
localize quality-gate --repo-root . --input-folder i18n --config config.yaml --validation-summary logs/translation_validation_summary.json --output-json logs/quality.json --output-markdown logs/quality.md --changed-files i18n/messages_de.properties
localize bootstrap-pr --target-project-root path/to/repo --action-ref v0.1.21
localize memory stats --memory-file logs/translation_memory.jsonCustom adapter modules can be loaded before translation commands:
localize --plugin my_project.localize_adapter formatsInstalled packages can also expose the localize.format_adapters entry point
group, or users can set LOCALIZE_PLUGIN_MODULES=module_a,module_b.
Full guide: docs/localization-cli.md.
Localize Guardian follows up on translation reviews so feedback can improve both the current PR and future runs. It checks feedback against the current source and localization rules, rather than treating a reviewer's suggestion as permission to change anything.
With the relevant write modes enabled, the workflow is:
- Read trusted feedback. Each project configures which human reviewers and bots, including CodeRabbit, may authorize work for its repositories and locales.
- Apply checked corrections. Eligible value-only changes go onto an allowed translation PR with a signed commit. Bot-labelled replies explain the result, link the correction, and distinguish alternatives, deferred work, and decisions still needed from a maintainer.
- Propose prevention. Recurring problems can produce a separate pipeline improvement PR with a regression test that fails before the fix and passes afterward. These PRs open ready for review. Guardian never merges PRs or deploys prevention changes. Automatically addressing reviews on those prevention PRs is not yet supported.
A glossary disagreement stays unresolved until a maintainer chooses the policy and the operator updates the governed configuration. A bot acknowledgement, resolved thread, or merged PR never silently approves that change.
Each project runs its own Guardian and supplies its machine, credentials, reviewer allowlists, limits, and maintenance. This is not a hosted service. Codex uses the operator's ChatGPT plan by default; metered API use is opt-in.
Start in observe, the default: it records assessments locally without
commits, pushes, or GitHub comments. Review a successful observation run before
enabling writes. Follow the setup guide
for authentication, validation, scheduling, status checks, and recovery.
To onboard another repository without hand-copying files, run:
localize bootstrap-pr --target-project-root path/to/repo --action-ref v0.1.21The command refuses dirty worktrees, creates a localize/onboarding branch, and
commits:
config.yamlgenerated from detected localization filesglossary.jsoncopied from the generic example.github/workflows/translate.ymlin safedry-run: truemodedocs/localize-pipeline.mdwith the target repository's rollout checklist
Add --push --open-pr when the target repo has origin and gh configured.
For custom adapters, pass --plugin-module and --plugin-install-command; the
generated workflow will install and load the adapter before running checks.
Use these packages for reusable code:
localize.core: pipeline contracts plus reusable filesystem/reporter/processor connectors.localize.formats: format metadata, adapters, plugin registration, and conformance tests.localize.providers: AISuite/OpenAI-compatible provider factories and capabilities.
Avoid importing implementation modules directly unless you are contributing to this repository.
Java properties, JSON, Bisq, and Bisq mobile are supported profiles, not core assumptions. New validation, queue handling, prompt behavior, and publishing logic should flow through config, adapters, providers, connectors, or profiles.
Before adding another localization format, use docs/new-format-checklist.md. The minimum bar is an adapter, conformance tests, realistic placeholder/escaping coverage, an example project, dry-run integration coverage, and docs.
The runtime maintains an exact-match translation memory at
logs/translation_memory.json by default. Successful, validation-safe
translations are recorded by normalized source text, target locale, and format.
Future runs reuse matching entries before calling the model. If the same source
segment later receives competing approved targets for the same locale and
format, the entry is marked as a conflict and no longer reused.
Config knobs:
translation_memory_enabled: true
translation_memory_file_path: "logs/translation_memory.json"Use a shared path if several projects should reuse one approved memory store. Manage memory stores with:
localize memory stats --memory-file logs/translation_memory.json
localize memory export --memory-file logs/translation_memory.json --output shared-memory.json
localize memory import --memory-file logs/translation_memory.json --input shared-memory.json
localize memory promote --memory-file logs/translation_memory.json \
--source-text "Save changes" \
--target-text "Änderungen speichern" \
--locale de \
--format-id json
localize memory suggest --memory-file logs/translation_memory.json \
--source-text "Save change" \
--locale de \
--format-id jsonFuzzy memory suggestions are review aids only. The runtime still reuses exact matches only.
Pin a tagged release for production workflows once tags are available:
- uses: bisq-network/localize-pipeline@v0.1.21Use @main only when you intentionally want the latest unreleased changes.
Release notes live in CHANGELOG.md.
This section covers the translation runner, not Localize Guardian. Guardian is installed and scheduled separately on infrastructure controlled by its operator. Most projects should use the GitHub Action. The Docker path is for scheduled server jobs that pull from Transifex and push signed translation PRs.
export DOCKER_BUILDKIT=1 COMPOSE_DOCKER_CLI_BUILD=1
docker compose --env-file docker/.env -f docker/docker-compose.yml build
docker compose --env-file docker/.env -f docker/docker-compose.yml run -T --rm translatorDocker mounts profiles/${TRANSLATOR_PROFILE:-bisq}/config.yaml and
glossary.json into the container. Keep deploy keys, GPG keys, and tokens in
secrets/ and docker/.env; never commit them.
Server guide: docs/new-project-deployment.md.
- docs/maintenance/disk-space-management.md covers Docker cleanup and log retention.
- scripts/docker-cleanup.sh is the ready-to-run cleanup script for Docker deployments.
- No locale files detected: choose the layout that matches your repository:
suffix,locale_directory, orlocale_filename. Permission denied (publickey)on push: the deploy key is missing or does not have write access to the fork repository.- Skipped translation inputs: files explicitly skipped by validation or
processing are excluded before PR batching and semantic review. The publisher
preserves their imported contents and validation summary under
logs/skipped-inputs-*/. Other valid files can still be published, but the run exits unsuccessfully, retains the git-source baseline, and withholds its success heartbeat for that run. Fix the reported input problem and rerun; the pipeline does not restore old translations automatically. - Quality gate failed: inspect the PR report. The pipeline reports skipped files, placeholder errors, semantic findings, and suspicious source-identical values.
- Generated PR title or counts look surprising: inspect the "Translation run summary" section. It separates output values changed from candidate keys, model calls, translation-memory reuse, and source-identical keys skipped because no ledger baseline exists yet.
- Action pushed a branch but failed to create a PR: enable repository Actions workflow write permissions and allow workflow-created pull requests in the target repository settings.
- Strict signed-commit rules reject generated PRs: configure SSH commit signing in the action with a dedicated machine-user signing key. See docs/github-action.md.
- Unexpected archive files in a generated PR: recent action versions exclude
archive folders from staging, but you should also ignore the pipeline archive
folder under your localization input folder, for example
/src/main/resources/archive/for Java.propertiessuffix layouts. - Model parameter errors: completion-token caps are normalized at the provider
boundary. Use
model_provider: aisuiteunless you need the directopenai_compatiblefallback.
Use TDD for behavior changes:
OPENAI_API_KEY=sk-test-key venv/bin/ruff check .
OPENAI_API_KEY=sk-test-key venv/bin/pytest -qKeep public docs, examples, and tests aligned with any API or configuration changes.