Directional Asymmetry in Low-Resource Legal Machine Translation: A Nepali-English Case Study

Sushan Adhikari, Department of Computer Science and Engineering, Kathmandu University, Dhulikhel, Bagmati, Nepal, sushan.adhikari2060@gmail.com
Sunidhi Sharma, Department of Computer Science and Engineering, Kathmandu University, Dhulikhel, Bagmati, Nepal, sunidhisharma3002@gmail.com
Bal Krishna Bal, Department of Computer Science and Engineering, Kathmandu University, Dhulikhel, Bagmati, Nepal, bal@ku.edu.np
Darshan Lamichhane, Department of Computer Science and Engineering, Kathmandu University, Dhulikhel, Bagmati, Nepal, darshanlamichhane012@gmail.com
Rajani Chulyadyo, Department of Computer Science and Engineering, Kathmandu University, Dhulikhel, Bagmati, Nepal, rajani.chulyadyo@ku.edu.np

We present a study of Nepali–English legal machine translation (MT), using a 5,024-pair parallel corpus derived from authoritative legal texts. Fine-tuning mBART-50 and NLLB-200 on this corpus, we observe a significant directional asymmetry: while bidirectional joint training improves Nepali → English translation by approximately 5 BLEU for both models, it results in a 6.3 BLEU degradation for mBART-50 in the English → Nepali direction, while NLLB-200 remains resilient. We attribute this discrepancy to Nepali's morphological complexity, which introduces challenges in training architecture sensitive models bidirectionally. To support future research, we release the corpus, trained models, and reproduction code.

CCS Concepts: • Computing methodologies → Natural language processing; • Computing methodologies → Machine translation;

Keywords: legal machine translation, low-resource NLP, Nepali language, multilingual models, bidirectional training, directional asymmetry, morphological complexity

ACM Reference Format:
Sushan Adhikari, Sunidhi Sharma, Bal Krishna Bal, Darshan Lamichhane, and Rajani Chulyadyo. 2026. Directional Asymmetry in Low-Resource Legal Machine Translation: A Nepali-English Case Study. In 21st International Conference on Artificial Intelligence and Law (ICAIL 2026), June 08--12, 2026, Singapore, Singapore. ACM, New York, NY, USA 2 Pages. https://doi.org/10.1145/3836937.3837011

1 Introduction

Machine translation (MT) for legal texts in low-resource languages remains underexplored. Nepal's legal system exemplifies this gap: official translations of documents like the Constitution of Nepal (2015) and the Muluki Ain (Civil and Criminal Code) exist in Nepali and English, but there are no computational resources for automated translation between them.

While multilingual pre-trained models such as mBART [3] and NLLB [4] have expanded MT to Nepali, the effectiveness of bidirectional joint fine-tuning for morphologically asymmetric language pairs remains unclear. Nepali's complex word formation contrasts with English's analytic structure, challenging traditional assumptions about bidirectional training improving MT performance [2].

Bidirectional legal MT systems, such as the one in [8], have created a 125k sentence corpus, but do not evaluate directional asymmetry or compare unidirectional and bidirectional strategies. Other work on domain-adaptive learning [1] and legal QA has addressed specific aspects of legal MT, but a comprehensive evaluation remains absent.

This paper presents three contributions: (1) the first publicly available Nepali–English legal parallel corpus, (2) an empirical demonstration of directional asymmetry in bidirectional fine-tuning, and (3) a comparison of mBART-50 and NLLB-200 under both unidirectional and bidirectional training configurations.

2 Corpus Construction & Experimental Setup

Corpus. We constructed a Nepali–English legal parallel corpus using three authoritative sources: the Constitution of Nepal (2015), the Muluki Ain (Civil and Criminal Code), and selected regulatory statutes. The corpus construction followed four stages: (1) OCR with Devanagari-optimized recognition, (2) rule-based sentence segmentation preserving legal clauses, (3) semi-automatic alignment using length heuristics and cognate matching, manually verified by bilingual law students, and (4) quality control by practicing lawyers to ensure legal semantic equivalence. Approximately 8% of initial alignments required re-alignment due to mismatches in sentence-level segmentation. The final corpus consists of 5,024 sentence pairs (4,019 for training, 502 for validation, and 503 for testing), with sentence lengths typical of legal texts (34 words in English, 25 words in Nepali on average).

Models and Configurations. We fine-tune two multilingual MT model families—mBART-50 (610M parameters, 50 languages) [3] and NLLB-200 (615M distilled, 200 languages) [4] using configurations: unidirectional Nepali → English, unidirectionalEnglish → Nepali, and bidirectional joint training (8,038 pairs, including both directions). This results in 6 models (2 architectures × 3 configurations) evaluated on both translation directions.

Hyperparameters. All models use identical hyperparameters: 6 training epochs, batch size 4 (effective 32 via gradient accumulation), learning rate of 3 × 10− 5 with 200-step warmup, and a maximum sequence length of 128 tokens. mBART-50 uses AdamW with HuggingFace Seq2SeqTrainer, while NLLB-200 uses Adafactor with a custom training loop to avoid compatibility issues. Training was performed on an NVIDIA A100 (40 GB).

Evaluation. We report BLEU [5], chrF++ [6], and length ratio, all computed using SacreBLEU [7]

3 Results and Discussion

We summarize translation quality across all model configurations in Table 1. A central pattern that emerges is directional asymmetry in bidirectional training. mBART-50 shows a sharp divergence in performance between the two directions, while NLLB-200 yields more balanced outcomes.

Table 1: Translation quality across models and configurations. Best scores per direction are bolded.
Model Config. Dir. BLEU chrF++
mBART-50 Unidir. NE → EN 49.05 64.51
mBART-50 Unidir. EN → NE 36.60 64.33
mBART-50 Bidir. NE → EN 53.97 66.23
mBART-50 Bidir. EN → NE 30.26 62.36
NLLB-200 Unidir. NE → EN 46.51 61.35
NLLB-200 Unidir. EN → NE 25.71 60.72
NLLB-200 Bidir. NE → EN 51.44 60.73
NLLB-200 Bidir. EN → NE 27.19 62.20

Directional asymmetry in mBART-50. Bidirectional training yields a +4.92 BLEU gain in NE → EN (49.05 → 53.97) but a -6.34 BLEU degradation in EN → NE (36.60 → 30.26). Joint training benefits the direction targeting the morphologically simpler language (English) while harming the direction that must generate rich Nepali morphology. We attribute this to representational interference: the shared encoder may optimize toward the lower-perplexity NE → EN direction, creating conflicting gradients that undermine EN → NE generation—an effect amplified by the analytic–synthetic structural contrast. Length ratios for mBART-50 is 0.929 for NE → EN, and 0.872 for EN → NE.

NLLB-200 resilience. NLLB-200 improves in both directions under bidirectional training (+4.93 BLEU NE → EN; +1.48 BLEU EN → NE). This robustness is attributed to its broader pre-training coverage (200 languages) and use of forced decoder start tokens, which isolate directional representations and reduce cross-directional interference. Length ratios for NLLB-200 is 0.929 for NE → EN, and 0.901 for (EN → NE).

Practical implications. Our findings challenge the assumption that bidirectional training universally benefits multilingual MT. For morphologically asymmetric legal pairs: (1) unidirectional fine-tuning is the safer default; (2) NLLB-200 is preferable when a single bidirectional model is required; and (3) bidirectional mBART-50 should be avoided when EN → NE quality is critical. All models achieve length ratios between 0.87 and 0.93. Given current quality levels, expert post-editing remains essential for legally binding contexts.

4 Conclusion

We presented a low-resource Nepali–English legal MT study, introducing a 5,024-pair parallel corpus and benchmarking mBART-50 and NLLB-200 under unidirectional and bidirectional fine-tuning. Our key finding is directional asymmetry in bidirectional training: mBART-50 improves by 4.92 BLEU in NE → EN but loses 6.34 BLEU in EN → NE, challenging standard MT practices for morphologically asymmetric pairs. NLLB-200 shows more robustness, improving in both directions.

We release the corpus, models, and code to support further research in legal NLP. Data and code are available at https://github.com/Sushan-Adhikari/LegalNLP.

Acknowledgments

We thank the bilingual law students and legal professional at Kathmandu University for their contributions to corpus annotation and quality verification, and the Law Commission of Nepal for providing corresponding legal documents in Nepali and English languages.

References

  • Sharad Duwal, Suraj Prasai, and Suresh Manandhar. 2024. Domain-Adaptative Continual Learning for Low-Resource Tasks: Evaluation on Nepali. arXiv preprint arXiv:2412.13860 (2024).
  • Melvin Johnson, Mike Schuster, Quoc V. Le, Maxim Krikun, Yonghui Wu, Zhifeng Chen, Nikhil Thorat, Fernanda Viégas, Martin Wattenberg, Greg Corrado, et al. 2017. Google's Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation. Transactions of the Association for Computational Linguistics 5 (2017), 339–351. https://doi.org/10.1162/tacl_a_00034
  • Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. Multilingual Denoising Pre-Training for Neural Machine Translation. Transactions of the Association for Computational Linguistics 8 (2020), 726–742. https://doi.org/10.1162/tacl_a_00343
  • NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, et al. 2022. No Language Left Behind: Scaling Human-Centered Machine Translation. arXiv preprint arXiv:2207.04672 (2022).
  • Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: A Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics. 311–318.
  • Maja Popović. 2015. chrF: Character N-gram F-score for Automatic MT Evaluation. In Proceedings of the Tenth Workshop on Statistical Machine Translation. 392–395.
  • Matt Post. 2018. A Call for Clarity in Reporting BLEU Scores. arXiv preprint arXiv:1804.08771 (2018).
  • Shabdapurush Poudel, Bal Krishna Bal, and Praveen Acharya. 2024. Bidirectional English-Nepali Machine Translation (MT) System for Legal Domain. (2024), 53–58.

CC-BY license image
This work is licensed under a Creative Commons Attribution 4.0 International License.

ICAIL 2026, Singapore, Singapore

© 2026 Copyright held by the owner/author(s).
ACM ISBN 979-8-4007-2756-6/26/06.
DOI: https://doi.org/10.1145/3836937.3837011