Artificial Intelligence in Simultaneous Interpreting: Capabilities, Limitations, and Human–Machine Collaboration

Authors

DOI:

https://doi.org/10.57125/FS.2026.03.20.10

Keywords:

artificial intelligence; simultaneous interpreting, automatic speech recognition, neural machine translation, human–machine collaboration, conference interpreting, cognitive load, computer-assisted interpreting, remote simultaneous interpreting, natural language processing.

Abstract

The integration of artificial intelligence (AI) into simultaneous interpreting (SI) constitutes one of the most consequential developments in language mediation studies of the past decade. Recent advances in neural machine translation (NMT), automatic speech recognition (ASR), and large language models (LLMs) have produced systems capable of rendering speech across language boundaries in near-real time, yet the fundamental cognitive, pragmatic, and cultural demands of professional conference interpreting continue to expose critical architectural limitations in current AI designs. This article provides a comprehensive interdisciplinary analysis of the state of AI in SI, structured around three principal aims: (1) to map the cognitive architecture of human SI against the computational strategies of current AI systems; (2) to synthesise empirical evidence on AI interpreting performance across language pairs, registers, and delivery conditions; and (3) to evaluate human–AI collaboration frameworks — including computer-assisted interpreting (CAI) and remote simultaneous interpreting (RSI) platforms — for their practical and ethical implications. We argue that the prevalent framing of AI as a replacement for human interpreters is both technically premature and epistemically inadequate. A more productive paradigm — augmentation — positions AI tools as extending interpreter competencies while preserving the pragmatic, affective, and cultural mediation capacities that remain beyond the reach of existing architectures.

References

AIIC. (2022). AIIC position paper on artificial intelligence and machine interpreting. International Association of Conference Interpreters. https://aiic.org/page/9107

Automatic Language Processing Advisory Committee. (1966). Languages and machines: Computers in translation and linguistics. National Academy of Sciences.

Arivazhagan, N., Cherry, C., Macherey, W., Chiu, C.-C., Yavuz, S., Pang, R., Li, W., & Wu, Y. (2019). Monotonic infinite lookback attention for simultaneous machine translation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 1313–1323). Association for Computational Linguistics. https://doi.org/10.18653/v1/P19-1126

Baddeley, A. D., & Hitch, G. J. (1974). Working memory. In G. Bower (Ed.), The psychology of learning and motivation (Vol. 8, pp. 47–89). Academic Press.

Bahdanau, D., Cho, K., & Bengio, Y. (2015). Neural machine translation by jointly learning to align and translate. In Proceedings of the 3rd International Conference on Learning Representations (ICLR 2015). https://doi.org/10.48550/arXiv.1409.0473

Bar-Hillel, Y. (1960). The present status of automatic translation of languages. Advances in Computers, 1, 91–163.

Barrault, L., Bhandare, Y., Bhatt, I., Bhosale, S., Duquenne, P.-A., Heffernan, N., Hwang, J., & Le, H. (2023). SeamlessM4T: Massively multilingual and multimodal machine translation. arXiv. https://doi.org/10.48550/arXiv.2308.11596

Becker, M., Schubert, T., Strobach, T., Gallinat, J., & Kühn, S. (2016). Simultaneous interpreters vs. professional multilingual controls: Group differences in cognitive control as well as brain structure and function. NeuroImage, 134, 250–260. https://doi.org/10.1016/j.neuroimage.2016.03.079

Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT '21) (pp. 610–623). Association for Computing Machinery. https://doi.org/10.1145/3442188.3445922

Bentivogli, L., Bertero, M., Di Gangi, M., Gaido, M., Negri, M., & Turchi, M. (2021). Is it worth it? Comparing six evaluation sets for English to German simultaneous speech translation. In Proceedings of the 18th International Conference on Spoken Language Translation (IWSLT 2021). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.iwslt-1.26

Braun, S. (2020). Towards a pedagogical framework for the integration of remote interpreting in interpreter education. Journal of Specialised Translation, 33, 324–343.

Brown, P. F., Cocke, J., Della Pietra, S. A., Della Pietra, V. J., Jelinek, F., Lafferty, J., Mercer, R., & Roossin, P. (1990). A statistical approach to machine translation. Computational Linguistics, 16(2), 79–85.

Cho, K., van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., & Bengio, Y. (2014). Learning phrase representations using RNN encoder-decoder for statistical machine translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 1724–1734). Association for Computational Linguistics. https://doi.org/10.3115/v1/D14-1179

Christoffels, I. K., & De Groot, A. M. B. (2005). Simultaneous interpreting: A cognitive perspective. In J. F. Kroll & A. M. B. De Groot (Eds.), Handbook of bilingualism: Psycholinguistic approaches (pp. 454–479). Oxford University Press.

Corpas Pastor, G., & Gaber, M. (2020). Ad hoc interpreting and the language of COVID-19: Terminological challenges for non-professional interpreters. Revista de Lenguas para Fines Específicos, 26(2), 10–33.

DePalma, D. A., & Kelly, N. (2008). Project management for translators. In J. Drugan (Ed.), Quality in professional translation (pp. 123–147). Continuum.

Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT 2019) (pp. 4171–4186). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423

Di Gangi, M. A., Gaido, M., Bentivogli, L., Negri, M., & Turchi, M. (2019). Enhancing transformer for end-to-end speech-to-text translation. In Proceedings of Machine Translation Summit XVII (Vol. 1, pp. 21–31).

Dong, L., & Xu, C. (2020). CIF: Continuous integrate-and-fire for end-to-end speech recognition. In Proceedings of the 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 6079–6083). IEEE. https://doi.org/10.1109/ICASSP40776.2020.9054250

Elbayad, M., Besacier, L., & Verbeek, J. (2020). Efficient wait-k models for simultaneous machine translation. In Proceedings of Interspeech 2020 (pp. 1461–1465). ISCA. https://doi.org/10.21437/Interspeech.2020-1480

Elmer, S., Hänggi, J., & Jäncke, L. (2014). Processing demands upon cognitive, linguistic, and articulatory functions promote grey matter plasticity in the adult multilingual brain. Cortex, 54, 179–191. https://doi.org/10.1016/j.cortex.2013.11.001

Engelbart, D. C. (1962). Augmenting human intellect: A conceptual framework (Summary Report AFOSR-3223). Stanford Research Institute.

Fantinuoli, C. (2017). Computer-assisted interpreting: Challenges and future perspectives. In G. Corpas Pastor (Ed.), Researching and using terminological resources in interpreting (pp. 224–250). Cambridge Scholars Publishing.

Fantinuoli, C. (2018). Interpreting and technology: The rising of the machine? In C. Fantinuoli (Ed.), Interpreting and technology (pp. 1–12). Language Science Press.

Fantinuoli, C. (2021). Machine interpreting: On the boundary of science and fiction. In Proceedings of the 1st Workshop on AI for Simultaneous Interpreting (AI4SI), LREC 2020.

FIT Europe. (2021). Technology and the future of translation: A position paper. International Federation of Translators, European Regional Centre.

Gaiba, F. (1998). The origins of simultaneous interpretation: The Nuremberg Trial. University of Ottawa Press.

Gile, D. (1995). Basic concepts and models for interpreter and translator training. John Benjamins.

Gile, D. (2009). Basic concepts and models for interpreter and translator training (2nd ed.). John Benjamins.

González-Davies, M., & Scott-Tennent, C. (2005). A problem-solving and student-centred approach to the translation of cultural references. Meta: Journal des traducteurs, 50(1), 160–179. https://doi.org/10.7202/010663ar

Hervais-Adelman, A., Moser-Mercer, B., & Golestani, N. (2015). Brain functional plasticity associated with the emergence of expertise in extreme language control. NeuroImage, 114, 264–274. https://doi.org/10.1016/j.neuroimage.2015.03.048

Hutchins, W. J. (2007). Machine translation: A concise history. Computer Aided Translation: Theory and Practice, 13, 29–70.

Jia, Y., Ramanovich, M., Remez, R., Pomeranz, L., Moreno, I., & Weiss, R. J. (2019). Direct speech-to-speech translation with a sequence-to-sequence model. In Proceedings of Interspeech 2019 (pp. 1123–1127). ISCA. https://doi.org/10.21437/Interspeech.2019-2820

Jia, Y., Ramanovich, M., Wang, Q., Zen, H., Wu, Y., Zhang, Y., & Bapna, A. (2022). Translatotron 2: High-quality direct speech-to-speech translation with voice preservation. In Proceedings of the 39th International Conference on Machine Learning (ICML 2022). https://doi.org/10.48550/arXiv.2107.08661

Kelly, N. (2022). The economics of the AI interpreting market (Nimdzi Research Report 2022). Nimdzi Insights.

Koehn, P., Hoang, H., Birch, A., Callison-Burch, C., Federico, M., Bertoldi, N., Cowan, B., & Ney, H. (2007). Moses: Open source toolkit for statistical machine translation. In Proceedings of the 45th Annual Meeting of the Association for Computational Linguistics: Demo and Poster Sessions (pp. 177–180). Association for Computational Linguistics. https://doi.org/10.3115/1557769.1557821

Koehn, P., & Knowles, R. (2017). Six challenges for neural machine translation. In Proceedings of the 1st Workshop on Neural Machine Translation (pp. 28–39). Association for Computational Linguistics. https://doi.org/10.18653/v1/W17-3204

Köpke, B., & Nespoulous, J.-L. (2006). Working memory performance in expert and novice interpreters. Interpreting, 8(1), 1–23. https://doi.org/10.1075/intp.8.1.03kop

Koshkin, R., Plaksin, A., & Ossadtchi, A. (2018). Measuring cognitive load in simultaneous interpreters using pupillometric data. PLOS ONE, 13(2), Article e0190736. https://doi.org/10.1371/journal.pone.0190736

Lederer, M. (1994). La traduction aujourd'hui: Le modèle interprétatif. Hachette.

Ma, M., Pino, J., Cross, J., Dong, L., & Diab, M. (2020). Monotonic multihead attention. In Proceedings of the 8th International Conference on Learning Representations (ICLR 2020). https://doi.org/10.48550/arXiv.1909.02480

Macháček, D., Dabre, R., & Bojar, O. (2023). Turning Whisper into real-time transcription system. arXiv. https://doi.org/10.48550/arXiv.2307.14743

Marcus, G., & Davis, E. (2019). Rebooting AI: Building artificial intelligence we can trust. Pantheon Books.

Mikkelson, H. (2017). Introduction to court interpreting (2nd ed.). Routledge.

Moorkens, J., Lewis, D., Kenny, D., Way, A., & Doğru Ünal, A. (2018). Correlations of perceived post-editing effort with measurements of machine translation quality. Machine Translation, 32(1), 93–114. https://doi.org/10.1007/s10590-018-9224-8

Moser-Mercer, B. (2005). Remote interpreting: Issues of multi-sensory integration in a multilingual task. Meta: Journal des traducteurs, 50(2), 727–738. https://doi.org/10.7202/011016ar

Moser-Mercer, B., Künzli, A., & Korac, M. (1998). Prolonged turns in interpreting: Effects on quality, physiological and psychological stress. Interpreting, 3(1), 47–64. https://doi.org/10.1075/intp.3.1.03mos

Nekoto, W., Marivate, V., Matsila, T., Fasubaa, T., Kolawole, T., Fagbohungbe, T., Akinola, S. O., & Muhammad, S. H. (2020). Participatory research for low-resourced machine translation: A case study in African languages. In Findings of the Association for Computational Linguistics: EMNLP 2020 (pp. 2144–2160). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.findings-emnlp.195

OpenAI. (2023). GPT-4 technical report. arXiv. https://doi.org/10.48550/arXiv.2303.08774

Paas, F., Renkl, A., & Sweller, J. (2003). Cognitive load theory and instructional design: Recent developments. Educational Psychologist, 38(1), 1–4. https://doi.org/10.1207/S15326985EP3801_1

Pagura, J., Arjona, E., & Stenzl, C. (1994). Methodology for quality assessment of simultaneous interpretation. In S. Lambert & B. Moser-Mercer (Eds.), Bridging the gap: Empirical research in simultaneous interpretation (pp. 317–325). John Benjamins.

Papineni, K., Roukos, S., Ward, T., & Zhu, W.-J. (2002). BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (ACL 2002) (pp. 311–318). Association for Computational Linguistics. https://doi.org/10.3115/1073083.1073135

Papi, S., Bentivogli, L., Gaido, M., Karakanta, A., Negri, M., & Turchi, M. (2023). Over-generation cannot be rewarded: Length-adaptive average lagging for simultaneous speech translation. arXiv. https://doi.org/10.48550/arXiv.2309.16715

Pöchhacker, F. (2001). Quality assessment in conference and community interpreting. Meta: Journal des traducteurs, 46(2), 410–425. https://doi.org/10.7202/004519ar

Pöchhacker, F. (2016). Introducing interpreting studies (2nd ed.). Routledge.

Rabinovich, E., Tsvetkov, Y., & Wintner, S. (2017). Personalized machine translation: Preserving original author traits. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2017) (pp. 1074–1084). Association for Computational Linguistics. https://doi.org/10.18653/v1/E17-1101

Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., & Sutskever, I. (2022). Robust speech recognition via large-scale weak supervision. arXiv. https://doi.org/10.48550/arXiv.2212.04356

Radford, A., Narasimhan, K., Salimans, T., & Sutskever, I. (2018). Improving language understanding by generative pre-training. OpenAI. https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf

Rei, R., Stewart, C., Farinha, A. C., & Lavie, A. (2020). COMET: A neural framework for MT evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 2685–2702). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.emnlp-main.213

Rinne, J. O., Tommola, J., Laine, M., Krause, B. J., Schmidt, D., Kaasinen, V., Teräs, M., & Solin, O. (2000). The translating brain: Cerebral activation patterns during simultaneous interpreting. Neuroscience Letters, 294(2), 85–88. https://doi.org/10.1016/S0304-3940(00)01540-8

Rinsche, A., & Portera-Zanotti, N. (2009). The size of the language industry in the EU (DGT Studies). European Commission.

Sandrelli, A. (2021). Technology in conference interpreting: From AIIC's early concerns to the post-COVID digital revolution. Translation, Cognition & Behavior, 4(2), 169–193.

Seeber, K. G. (2011). Cognitive load in simultaneous interpreting: Existing theories — new models. Interpreting, 13(2), 176–204. https://doi.org/10.1075/intp.13.2.03see

Seeber, K. G., & Kerzel, D. (2012). Cognitive load in simultaneous interpreting: Model meets data. International Journal of Bilingualism, 16(2), 228–242. https://doi.org/10.1177/1367006911403202

Seleskovitch, D. (1975). Langage, langues et mémoire: Étude de la prise de notes en interprétation consécutive. Minard Lettres Modernes.

Sellam, T., Das, D., & Parikh, A. (2020). BLEURT: Learning robust metrics for text generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL 2020) (pp. 7881–7892). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.704

Signorelli, T. M., Haarmann, H. J., & Obler, L. K. (2012). Working memory in simultaneous interpreters: Effects of task and age. International Journal of Bilingualism, 16(2), 198–212. https://doi.org/10.1177/1367006911403200

Sun, T., Gaut, A., Tang, S., Huang, Y., ElSherief, M., Zhao, J., Mirber, D., & Wang, W. Y. (2021). Mitigating gender bias in natural language processing: Literature review. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics (ACL 2021). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.acl-long.416

Sutskever, I., Vinyals, O., & Le, Q. V. (2014). Sequence to sequence learning with neural networks. In Advances in Neural Information Processing Systems 27 (NeurIPS 2014). Curran Associates.

Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2), 257–285. https://doi.org/10.1207/s15516709cog1202_4

Tatman, R., & Kasten, C. (2017). Effects of talker dialect, gender and race on accuracy of Bing Speech and YouTube automatic captions. In Proceedings of Interspeech 2017 (pp. 934–938). ISCA. https://doi.org/10.21437/Interspeech.2017-1746

Van Besien, F. (1999). Anticipation in simultaneous interpretation. Meta: Journal des traducteurs, 44(2), 250–259. https://doi.org/10.7202/003686ar

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems 30 (NeurIPS 2017). Curran Associates. https://doi.org/10.48550/arXiv.1706.03762

Viezzi, M. (1996). Aspetti della qualità in interpretazione. SERT, Scuola Superiore di Lingue Moderne per Interpreti e Traduttori.

Wordly User Research. (2023). User satisfaction with AI-powered simultaneous interpretation in enterprise settings: Preliminary findings. Wordly Inc.

Zhang, Y., Han, W., Qin, J., Wang, Y., Bapna, A., Chen, Z., Chen, N., & Wu, Y. (2023). Google USM: Scaling automatic speech recognition beyond 100 languages. arXiv. https://doi.org/10.48550/arXiv.2303.01037

Downloads

Published

2026-03-15

How to Cite

Milcu, M. (2026). Artificial Intelligence in Simultaneous Interpreting: Capabilities, Limitations, and Human–Machine Collaboration. Futurity of Social Sciences, 4(1), 169–187. https://doi.org/10.57125/FS.2026.03.20.10