Skip to main navigation menu Skip to main content Skip to site footer

Research-based papers

ANO XXVIII - NÚMERO 42 - 2026

BackTranslationLLM: An Agentic AI Architecture for Translating and Validating the Content of Psychometric Instruments

DOI
https://doi.org/10.51914/brjmt.42.2026.486
Submitted
January 17, 2026
Published
29-07-2026

Abstract

Cross-cultural adaptation of psychometric instruments is a methodologically complex, time-consuming, and resource-intensive process.  This study introduces and evaluates BackTranslationLLM, a multi-agent Artificial Intelligence (AI) architecture designed to automate the translation and content validation of psychological tests. Guided by principles of AI agent engineering, the system employs a sequential pipeline with specialized agents for translation, back-translation, and refinement, followed by a diversified committee of AI judges for evaluation. The framework was applied to adapt the Cuestionario de Impacto de la Sesión de Musicoterapia from Spanish to Brazilian Portuguese. Results were benchmarked against a parallel validation conducted by five human expert judges. Exploratory Graph Analysis confirmed that the automated translation preserved the instrument's semantic structure. Both the AI and human committees assigned adequate Content Validity Coefficients (CVCs) to the adapted instrument. However, a lack of correlation between the item-level CVCs (e.g., Clarity, ρ = -0.37) revealed divergent judgment heuristics: the AI acted as an auditor of methodological compliance, whereas human experts excelled at evaluating pragmatic and cultural nuances. We conclude that BackTranslationLLM is an effective framework that complements, rather than replaces, human expertise by optimizing the validation process.

References

  1. AERA - American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association.
  2. Bhagwat, S. (2024). Principles of Building AI Agents (2nd ed.). Mastra.ai.
  3. Boateng, G. O., Neilands, T. B., Frongillo, E. A., Melgar-Quiñonez, H. R., & Young, S. L. (2018). Best practices for developing and validating scales for health, social, and behavioral research: A primer. Frontiers in Public Health, 6, 149. DOI: https://doi.org/10.3389/fpubh.2018.00149
  4. Borsa, L. C.; Seize, M.. (2017). Construção e adaptação de instrumentos psicológicos: dois caminhos possíveis. In: Damásio, B. F.; Borsa, J. C.. Manual de desenvolvimento de instrumentos psicológicos. São Paulo: Vetor.
  5. Cassepp-Borges, V., Balbinotti, M. A. A., & Teodoro, M. L. M. (2010). Tradução e validação de conteúdo: uma proposta para a adaptação de instrumentos. In L. Pasquali (Org.), Instrumentação psicológica: Fundamentos e práticas (pp. 506-520). Artmed.
  6. Cripps, C.; Tsiris, G.; Spiro, N. (2016). Outcome measures in music therapy: A resource developed by the Nordoff Robbins research team. 1. ed. London: Nordoff Robbins. https://eresearch.qmu.ac.uk/handle/20.500.12289/4429
  7. Danon, L., Díaz-Guilera, A., Duch, J., & Arenas, A. (2005). Comparing community structure identification. Journal of Statistical Mechanics: Theory and Experiment, 2005(09), P09008–P09008. https://doi.org/10.1088/1742-5468/2005/09/P09008. DOI: https://doi.org/10.1088/1742-5468/2005/09/P09008
  8. DeepL SE. (2024). DeepL Translator. https://www.deepl.com/
  9. Google. (2024a). Gemini: A family of highly capable multimodal models. https://deepmind.google/technologies/gemini/
  10. Google. (2024b). Gemma: Open models based on Gemini research and technology. https://ai.google/discover/gemma/
  11. Hernandez-Nieto, R. A. (2002). Contributions to statistical analysis. Universidad de Los Andes.
  12. Kyriazos, T. A., & Stalikas, A. (2018). Applied psychometrics: The steps of scale development and standardization process. Psychology, 9(11), 2531-2560. DOI: https://doi.org/10.4236/psych.2018.911145
  13. Laverghetta Jr., A., & Licato, J. (2023). Generating better items for cognitive assessments using large language models. In E. Kochmar, J. Burstein, A. Horbach, R. Laarmann-Quante, N. Madnani, A. Tack, V. Yaneva, Z. Yuan, & T. Zesch (Eds.), Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023) (pp. 414–428). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.bea-1.34 DOI: https://doi.org/10.18653/v1/2023.bea-1.34
  14. Laverghetta Jr., A., Luchini, S., Linell, A., Reiter-Palmon, R., & Beaty, R. (2024). The creative psychometric item generator: A framework for item generation and validation using large language models. arXiv. https://arxiv.org/abs/2409.00202
  15. Nakano, T. C., & Siqueira, L. G. G. (2012). Validade de conteúdo da Gifted Rating Scale (versão escolar) para a população brasileira. Avaliação Psicológica, 11(1), 123-140.
  16. Pasquali, L. (2009). Psicometria: Teoria dos testes na psicologia e na educação. Vozes.
  17. Pasquali, L. (2010). Instrumentação psicológica: Fundamentos e práticas. Artmed.
  18. Pedrosa, F. G. (2023). Escala de Avaliação dos Efeitos da Musicoterapia em Grupo na Dependência Química (MTDQ) [Tese, Universidade Federal de Minas Gerais]. https://doi.org/10.13140/RG.2.2.21769.04962 DOI: https://doi.org/10.35699/2317-6377.2023.45027
  19. R Core Team. (2025). R: A language and environment for statistical computing. R Foundation for Statistical Computing. https://www.R-project.org/
  20. Russell-Lasalandra, L.L., Christensen, A.P. & Golino, H. (2026) Generative psychometrics via AI-GENIE: Automatic item generation and validation with network-integrated evaluation. Behavior Research Methods 58, 217. https://doi.org/10.3758/s13428-026-03082- DOI: https://doi.org/10.3758/s13428-026-03082-1
  21. Silveira, M. B., et al. (2018). Construção e validade de conteúdo de um instrumento para avaliação de quedas em idosos. Einstein (São Paulo), 16(2).
  22. Vinh, N. X., Epps, J., & Bailey, J. (2010). Information theoretic measures for clusterings comparison. Journal of Machine Learning Research. 11(95):2837−2854.
  23. Vercher, I. B., Soler, A. A., & Ferrari, K. D. (2023). Cuestionario CISMA - Cuestionario del Impacto de las Sesiones de Musicoterapia en Pacientes Adultos. Brazilian Journal of Music Therapy, (33). https://doi.org/10.51914/brjmt.33.2022.385 DOI: https://doi.org/10.51914/brjmt.33.2022.385
  24. Zmitrowicz, M., & Moura, C. de F. (2018). Musicoterapia e reabilitação: uma revisão de escopo. Cadernos Brasileiros de Terapia Ocupacional, 26(4), 891-903.

Similar Articles

<< < 2 3 4 5 6 7 8 9 10 11 > >> 

You may also start an advanced similarity search for this article.

Most read articles by the same author(s)