Logo PTI Logo FedCSIS

Position Papers of the 21st Conference on Computer Science and Intelligence Systems

Annals of Computer Science and Information Systems, Volume 48

Multi-Candidate Synthesis in Multimodal Machine Translation

, , , ,

DOI: http://dx.doi.org/10.15439/2026F4324

Citation: Jan Kolwicz, , , ,

Full text

Abstract. Multimodal Machine Translation (MMT) typically relies on passing text and corresponding images to a model to resolve ambiguities. However, the direct presence of an image can occasionally mislead the model, causing regressions compared to text-only translation. To address this, we propose a 5-step pipeline for evaluating and refining MMT. Instead of single-pass translations, our approach generates four diverse candidate translations leveraging different subsets of multimodal context (text-only, text+image, text+caption, text+image+caption). A final synthesis step acts as a judge and editor to produce the optimal translation. Evaluated on the ConECT and Multi30k datasets using large language models (Qwen3.6-35B and Gemma 4 26B), our synthesis pipeline consistently outperforms text-only and standard multimodal approaches. We observe up to a +1.34 improvement in COMET scores.

References

  1. O. Caglayan, P. Madhyastha, L. Specia, and L. Barrault, “Probing the need for visual context in multimodal machine translation,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), J. Burstein, C. Doran, and T. Solorio, Eds. Minneapolis, Minnesota: Association for Computational Linguistics, Jun. 2019, pp. 4159–4170. [Online]. Available: https://aclanthology.org/N19-1422/
  2. M. Pokrywka, W. Kusa, M. Rutkowski, and M. Koszowski, “Conect dataset: Overcoming data scarcity in context-aware e-commerce mt,” 2025. [Online]. Available: https://arxiv.org/abs/2506.04929
  3. L. Specia, S. Frank, K. Sima’an, and D. Elliott, “A shared task on multimodal machine translation and crosslingual image description,” in Proceedings of the First Conference on Machine Translation: Volume 2, Shared Task Papers, O. Bojar, C. Buck, R. Chatterjee, C. Federmann, L. Guillou, B. Haddow, M. Huck, A. J. Yepes, A. Névéol, M. Neves, P. Pecina, M. Popel, P. Koehn, C. Monz, M. Negri, M. Post, L. Specia, K. Verspoor, J. Tiedemann, and M. Turchi, Eds. Berlin, Germany: Association for Computational Linguistics, Aug. 2016, pp. 543–553. [Online]. Available: https://aclanthology.org/W16-2346/
  4. Y. Feng, C. Li, J. He, Z. Hou, and V. Ng, “Multimodal neural machine translation: A survey of the state of the art,” in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng, Eds. Suzhou, China: Association for Computational Linguistics, Nov. 2025, pp. 22 130–22 147. [Online]. Available: https://aclanthology.org/2025.emnlp-main.1125/
  5. M. Zheng, Z. Li, B. Qu, M. Song, Y. Du, M. Sun, and D. Wang, “Hunyuan-mt technical report,” 2025. [Online]. Available: https://arxiv.org/abs/2509.05209
  6. D. Elliott, S. Frank, K. Sima’an, and L. Specia, “Multi30k: Multilingual english-german image descriptions,” arXiv preprint https://arxiv.org/abs/1605.00459, 2016.
  7. Q. Team, “Qwen3 technical report,” 2025. [Online]. Available: https://arxiv.org/abs/2505.09388
  8. G. Team, “Gemma 3 technical report,” 2025. [Online]. Available: https://arxiv.org/abs/2503.19786
  9. M. Post, “A call for clarity in reporting BLEU scores,” in Proceedings of the Third Conference on Machine Translation: Research Papers. Belgium, Brussels: Association for Computational Linguistics, Oct. 2018, pp. 186–191. [Online]. Available: https: //www.aclweb.org/anthology/W18-6319
  10. R. Rei, J. G. C. de Souza, D. Alves, C. Zerva, A. C. Farinhas, T. Glushkova, A. Lavie, L. Coheur, and A. F. T. Martins, “COMET22: Unbabel-IST 2022 submission for the metrics shared task,” in Proceedings of the Seventh Conference on Machine Translation (WMT). Abu Dhabi, United Arab Emirates (Hybrid): Association for Computational Linguistics, Dec. 2022, pp. 578–585. [Online]. Available: https://aclanthology.org/2022.wmt-1.52