ParaMed: a parallel corpus for English–Chinese translation in the biomedical domain-Reference-Cited by-同舟云学术

ParaMed: a parallel corpus for English–Chinese translation in the biomedical domain

Published:2021-09-06 Issue:1 Volume:21 Page:
ISSN:1472-6947
Container-title:BMC Medical Informatics and Decision Making
language:en
Short-container-title:BMC Med Inform Decis Mak

Author:

Liu Boxiang^ORCID,Huang Liang

Abstract

Abstract Background Biomedical language translation requires multi-lingual fluency as well as relevant domain knowledge. Such requirements make it challenging to train qualified translators and costly to generate high-quality translations. Machine translation represents an effective alternative, but accurate machine translation requires large amounts of in-domain data. While such datasets are abundant in general domains, they are less accessible in the biomedical domain. Chinese and English are two of the most widely spoken languages, yet to our knowledge, a parallel corpus does not exist for this language pair in the biomedical domain. Description We developed an effective pipeline to acquire and process an English-Chinese parallel corpus from the New England Journal of Medicine (NEJM). This corpus consists of about 100,000 sentence pairs and 3,000,000 tokens on each side. We showed that training on out-of-domain data and fine-tuning with as few as 4000 NEJM sentence pairs improve translation quality by 25.3 (13.4) BLEU for en

$$\rightarrow$$

→ zh (zh

$$\rightarrow$$

→ en) directions. Translation quality continues to improve at a slower pace on larger in-domain data subsets, with a total increase of 33.0 (24.3) BLEU for en

$$\rightarrow$$

→ zh (zh

$$\rightarrow$$

→ en) directions on the full dataset. Conclusions The code and data are available at https://github.com/boxiangliu/ParaMed.

Publisher

Springer Science and Business Media LLC

Subject

Health Informatics,Health Policy,Computer Science Applications

Link

https://link.springer.com/content/pdf/10.1186/s12911-021-01621-8.pdf

Reference42 articles.

1. Bamforth I. Biomedical translation. BMJ. 1998;316(7124):2–7124.

2. Das A. Medical interpreters. BMJ. 2009;338:2354.

3. Hassan H, Aue A, Chen C, Chowdhary V, Clark J, Federmann C, Huang X, Junczys-Dowmunt M, Lewis W, Li M, et al. Achieving human parity on automatic chinese to english news translation; 2018. arXiv preprint arXiv:1803.05567

4. Bodenreider O. The unified medical language system (UMLs): integrating biomedical terminology. Nucleic Acids Res. 2004;32(suppl–1):267–70.

5. Sennrich R, Haddow B, Birch A. Improving neural machine translation models with monolingual data; 2015. arXiv preprint arXiv:1511.06709

Cited by 10 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Artificial Intelligence in Multilingual Interpretation and Radiology Assessment for Clinical Language Evaluation (AI-MIRACLE);Journal of Personalized Medicine;2024-08-30

2. Domain Adaptation for Arabic Machine Translation: Financial Texts as a Case Study;Applied Sciences;2024-08-13

3. Time Convolutional Network-Transformer based Chinese English Translation Model;2024 Second International Conference on Data Science and Information System (ICDSIS);2024-05-17

4. WCC-EC 2.0: Enhancing Neural Machine Translation with a 1.6M+ Web-Crawled English-Chinese Parallel Corpus;Electronics;2024-04-05

5. Computer Information Extraction Algorithm Based on English Corpus;2024 International Conference on Distributed Computing and Optimization Techniques (ICDCOT);2024-03-15