Repository navigation
Conversation
Add a script to make it more easy to update the current translations with the fixes from the upstream main.
|
Thanks for the pull request, @igobranco! This repository is currently maintained by Once you've gone through the following steps feel free to tag them in a comment and let them know that your changes are ready for engineering review. 🔘 Get product approvalIf you haven't already, check this list to see if your contribution needs to go through the product review process.
🔘 Provide contextTo help your reviewers and other members of the community understand the purpose and larger context of your changes, feel free to add as much of the following information to the PR description as you can:
🔘 Get a green buildIf one or more checks are failing, continue working on your changes until this is no longer the case and your build turns green. DetailsWhere can I find more information?If you'd like to get more details on all aspects of the review process for open source pull requests (OSPRs), check out the following resources: When can I expect my changes to be merged?Our goal is to get community contributions seen and reviewed as efficiently as possible. However, the amount of time that it takes to review and merge a PR can vary significantly based on factors such as:
💡 As a result it may take up to several weeks or months to complete a review and merge your PR. |
|
Hi @openedx/committers-translations -- checking in on this PR. It's been stalled for some time, can we move it forward or perhaps close it? |
Agrendalath
left a comment
There was a problem hiding this comment.
Hi @igobranco, I just noticed this while checking the hackathon board. Thanks for putting this together - the script is nicely documented, and it's great to see test coverage. However, I'd like to step back and ask whether we need it at all, because I think plain git serves this use case better. A couple of the current behaviors may not apply to common scenarios.
My main concern is the merge model. This script does a 2-way union with no merge base, whereas git merge/git rebase does a 3-way merge that knows the fork point. That distinction matters because:
- Without a merge base, the script can't tell "I intentionally changed this" from "this is just stale", so it can never surface a real conflict. It just silently resolves one way.
- For PO files,
msgcat --use-firstlists upstream first, so upstream wins every sharedmsgid. That would overwrite any custom commits one may have added to their fork. Say they added custom fixes to override the defaults, or pulled unreviewed translations because they needed higher translation coverage for a non-English instance. This script would not preserve these changes. - For JSON,
merge_translationsonly ever adds new keys and never overwrites, so it does the opposite of the PO path. The two file types resolve conflicts in opposite directions, which is confusing.
Separately, the sorting makes future diffs harder, which is the practical cost we'll feel most. dict(sorted(...)) for JSON and --sort-output for PO reorder entire files. Most files in the repo aren't alphabetically sorted now, so a run reorders them wholesale, making diffing a custom branch against main or a release branch much noisier in the future.
A couple of smaller things:
- In all-files mode,
root_dir.rglob("*.json")will also pick uptransifex_input.jsonsource files, which I don't think we want to touch. - A missing upstream file is an expected case, but
run_commandprintsError running command/stderrto stderr beforefetch_upstream_filecatches the exception, so normal runs look confusing.
Given all this, my suggestion is to drop the script and use git for the refresh, since forks share history with this repo and a 3-way merge handles this scenario natively. A downstream that tracks main can git merge main (or git rebase onto it) and let git preserve ordering, keep its custom commits, and raise real conflicts where both sides touched the same string. If the custom changes are small, cherry-picking those commits onto a fresh upstream is even simpler and keeps clean diffs against main and release branches over time.
If there's a recurring need I'm missing - for example, reconciling two full Transifex syncs of PO files, where line-level git merge gets too noisy - then a gettext-aware helper could be a better fit. In that case, we could consider:
- Dropping the sorting.
- Flipping the precedence to preserve local changes.
- Excluding the
transifex_input.jsonsource files. - Aligning the JSON path with the PO path so they behave consistently.
Could you share a bit more about the scenario you had in mind?
brian-smith-tcril
left a comment
There was a problem hiding this comment.
While I understand maintaining translations across branches is painful, I do not believe this PR will address those problems.
The source of truth for translated strings is Transifex. This script does not change the translated strings in Transifex, meaning the changes this script makes will be overridden by the ones in Transifex when updates are made there (wiping out all changes that came from the script).
I know @OmarIthawi has done some work on a script that serves a similar purpose but ensures the strings are updated in Transifex, and we are currently investigating ways to utilize the branching functionality that was recently added to Transifex as well.
|
Thanks for the thorough reviews, @Agrendalath and @brian-smith-tcril! On using On Transifex as the source of truth (@brian-smith-tcril): The script is intentionally scoped to local git branches only - it's not meant to push anything back to Transifex. The target use case is an operator or release manager who needs to quickly backport translation fixes to a release branch before (or instead of) waiting for a full Transifex sync cycle. For context, I didn't actually use this script for our last major upgrade either - by then we always try to have all translations translated and reviewed on our own That said, I'm very interested in the Transifex branching feature you mentioned. If Transifex now supports branching natively, that could be a cleaner long-term solution. Could you share more about where that investigation stands or point me to @OmarIthawi's work? I'd be happy to close this PR if there's a better path forward through Transifex itself. |
|
Thanks @igobranco for the contribution and pardon the late review. Please note that the recommended pattern for introducing overrides involves another strategy:
Commit to github and pull from that repo. This is the preferred approach compared to few alternatives we considered earlier. I will close this PR. |
|
Thanks @OmarIthawi for your share! I think those are very good examples!
|
@igobranco I assume that you want to avoid duplicating effort and your Release better translated than Main, right? This is a bigger problem than custom translations that we're working with Transifex to support. Merging release and main have a lot of edge cases that we'd rather offloading to Transifex because it has much more metadata than us. |
|
@OmarIthawi yes, I want to avoid duplicating efforts. I have come to cases where it was missing translations and/or revisions on a Open edX release. Other cases it was required a translation fixes (not good enough translations). |
Add a script to make it more easy to update the current translations with the fixes and new translations from the upstream main.
This PR adds a Python script that updates the current translations from the upstream main translations.
So from a previous release we can make it more easy just receive the latest patched (fixed) and/or new translations.
Use cases:
Copied from the script comment: