Building multilingual term mapping models: a case study in industry terminology
Main Article Content
Abstract
Cross‑lingual natural language processing has moved from academic curiosity to industrial necessity. Modern software development is global, with teams distributed across continents and documentation written in a handful of lingua franca. Terms like autoscaling group, immutable infrastructure or zero‑trust network segmentation emerge in English‑language cloud platforms and quickly travel to DevOps manuals, system administration runbooks and cyber‑security guidelines. Yet those terms often lack precise equivalents in other languages, and literal translations can lead to misunderstandings. This article examines how to build multilingual term mapping models that align industry terminology across languages, with a focus on the information‑technology (IT) domain. We describe a layered methodology that combines statistical term extraction, semantic alignment of cross‑lingual embeddings, iterative fine‑tuning of neural translators and careful curation of expert glossaries. A large‑scale case study covers cloud computing, DevOps practices and cyber‑security, analyzing the effectiveness of different approaches on English–Russian and English–Kazakh pairs. Experiments reveal that the proposed hybrid model reduces errors by over 20 percentage points compared with dictionary‑based baselines, increases term coverage to over 80 % and substantially cuts manual post‑editing time.
Article Details
References
Bayekeyeva A. et al. Controlled multilingual thesauri for Kazakh industry-specific terms //Social Inclusion. – 2021. – Т. 9. – № . 1. – С. 35-44. https://www.cogitatiopress.com/socialinclusion/article/view/3527 DOI: https://doi.org/10.17645/si.v9i1.3527
Köksal Ö., Tekinerdogan B. Automated classification of unstructured bilingual software bug reports: An industrial case study research //Applied Sciences. – 2021. – Т. 12. – № . 1. – С. 338. https://www.mdpi.com/2076-3417/12/1/338 DOI: https://doi.org/10.3390/app12010338
Sandhu A. K. Big data with cloud computing: Discussions and challenges //Big Data Mining and Analytics. – 2021. – Т. 5. – № . 1. – С. 32-40. https://ieeexplore.ieee.org/abstract/document/9663258/ DOI: https://doi.org/10.26599/BDMA.2021.9020016
Shetty J. P., Panda R. An overview of cloud computing in SMEs //Journal of Global Entrepreneurship Research. – 2021. – Т. 11. – № 1. – С. 175-188. https://link.springer.com/article/10.1007/s40497-021-00273-2 DOI: https://doi.org/10.1007/s40497-021-00273-2
Homayouni A. NLP-based Failure log Clustering to Enable Batch Log Processing in Industrial DevOps Setting. – 2022. https://www.diva-portal.org/smash/record.jsf?pid=diva2:1673361
Asudani D. S., Nagwani N. K., Singh P. Impact of word embedding models on text analytics in deep learning environment: a review //Artificial intelligence review. – 2023. – Т. 56. – № . 9. – С. 10345-10425. https://link.springer.com/article/10.1007/S10462-023-10419-1 DOI: https://doi.org/10.1007/s10462-023-10419-1
Ali H., Khan E., Sajad M.A. Phytoremediation of heavy metals—concepts and applications // Chemosphere. – 2013. – Vol. 91, № 7. – P. 869–881. – URL: https://doi.org/10.1016/j.chemosphere.2013.01.075 DOI: https://doi.org/10.1016/j.chemosphere.2013.01.075
Mittal A. K., Chisti Y., Banerjee U. C. Synthesis of metallic nanoparticles using plant extracts // Biotechnology Advances. – 2013. – Vol. 31, No. 2. – P. 346–356. – DOI: https://doi.org/10.1016/j.biotechadv.2013.01.003 DOI: https://doi.org/10.1016/j.biotechadv.2013.01.003
