Short Courses of WebMedia 2025
Keywords:
Multimodal Large Language Models, 3GPP, Human Subjects Research, User Experience, Digital Accessibility, Inclusive WebSynopsis
Traditionally, the Brazilian Symposium on Multimedia and the Web provides participants with the opportunity to explore and learn about relevant topics in the field through its Short Courses and Tutorials track. In its 31st edition, held in the city of Rio de Janeiro from November 11 to 14, 2025, the event featured 3 tutorials and 5 short courses, selected through a process based on proposals submitted in response to a nationwide public call.
As in previous editions of WebMedia, the instructors of the short courses were invited to produce a text corresponding to the content presented, for publication in the form of a book chapter. Thus, this volume brings together, retrospectively, the material corresponding to three short courses.
Chapter 1, entitled "Multimodal Large Language Models (MLLMs): From Theory to Practice", addresses the topic of current MLLMs from the perspective of multimodality, combining the ability to understand and generate natural language with perceptions derived from modalities such as images and audio. The chapter explores practical techniques for preprocessing, prompt engineering, and the construction of multimodal pipelines, while also highlighting promising trends.
Chapter 2, "Human Subjects Research in Computing: Ethical Aspects, Regulation, and Educational Scenarios", presents the historical and normative foundations of research ethics in Computing research involving human subjects. From an educational perspective, hypothetical scenarios illustrating recurring situations in Computing research are analyzed, highlighting how ethical principles apply and the need for institutional ethical review.
Chapter 3, "UX Research Involving People with Disabilities: A Path Toward Inclusive Web Media", addresses the issue of digital accessibility in Brazil, discussing the importance of conducting research focused on User Experience (UX Research) that involves the direct participation of people with disabilities. In addition, it presents technical guidelines as well as the practical and ethical challenges involved in designing democratic and accessible Web and multimedia systems.
Chapters
-
1. Multimodal Large Language Models (MLLMs): From Theory to Practice
-
2. Human Subjects Research in Computing: Ethical Aspects, Regulation, and Educational Scenarios
-
3. UX Research Involving People with Disabilities: A Path Toward Inclusive Web Media
Downloads
References
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al. (2022). Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems, 35:23716–23736.
Albres, N. A. and Sousa, D. V. C. (2019). Termo de assentimento livre e esclarecido: uso de história em quadrinhos em pesquisas com crianças. Revista Sinalizar, 4.
ANDRADE, M. G.; RABELO, D. M. F.; MARTINS, R. de S.; VIANA, W. Investigating the accessibility of popular mobile Android apps: a prevalence, category, and language study. In: SIMPÓSIO BRASILEIRO DE SISTEMAS MULTIMÍDIA E WEB (WEBMEDIA), 30., 2024. Anais [...]. Porto Alegre: SBC, 2024. p. 400-404.
ARCANJO, L. D. S.; COELHO, L. F.; GUIMARÃES, S. J. F.; PATROCÍNIO JÚNIOR, Z. K. do; CARDOSO, L. V. Automatic time-aware recognition of Brazilian sign language based on dynamic time warping. In: SIMPÓSIO BRASILEIRO DE SISTEMAS MULTIMÍDIA E WEB (WEBMEDIA), 30., 2024. Anais [...]. Porto Alegre: SBC, 2024. p. 72-79.
ASSOCIAÇÃO BRASILEIRA DE NORMAS TÉCNICAS. ABNT NBR 17225:2025: acessibilidade em conteúdo e aplicações web — requisitos. Rio de Janeiro: ABNT, 2025. Disponível em: [link]. Acesso em: 12 maio 2026.
Baevski, A., Zhou, Y., Mohamed, A., and Auli, M. (2020). wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in neural information processing systems, 33:12449–12460.
Baltrušaitis, T., Ahuja, C., and Morency, L.-P. (2018). Multimodal machine learning: A survey and taxonomy. IEEE transactions on pattern analysis and machine intelligence, 41(2):423–443.
BARTALOTTI, Celina Camargo. Inclusão social das pessoas com deficiência: utopia ou possibilidade?. Paulus, 2006.
Be My Eyes (2023). Be My Eyes – Be My AI: Accessible visual assistance powered by GPT-4. [link]. Acesso em: out. 2025.
Belmont (1979). The Belmont Report: Ethical Principles and Guidelines for the Protection of Human Subjects of Research. The National Commission for the Protection of Human Subjects of Biomedical and Behavioral Research.
Bender, E., Gebru, T., McMillan-Major, A., and Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? pages 610–623.
BERSCH, Rita. Introdução à tecnologia assistiva. Porto Alegre: Assistiva Tecnologia e Educação, 2017.
BI, Tingting et al. Accessibility in software practice: A practitioner’s perspective. ACM Transactions on Software Engineering and Methodology (TOSEM), v. 31, n. 4, p. 1-26, 2022.
BIGDATACORP; MOVIMENTO WEB PARA TODOS. Panorama da acessibilidade digital no Brasil. São Paulo: BigDataCorp, 2024. Disponível em: [link]. Acesso em: 8 maio 2026.
Birhane, A., Prabhu, V. U., and Kahembwe, E. (2021). Multimodal datasets: misogyny, pornography, and malignant stereotypes.
Borth, D., Ji, R., Chen, T., Breuel, T., and Chang, S.-F. (2013). Largescale visual sentiment ontology and detectors using adjective noun pairs. In Proceedings of the 21st ACM International Conference on Multimedia, MM ’13, page 223–232, New York, NY, USA. Association for Computing Machinery.
BRASIL. Agenda 2030 para o Desenvolvimento Sustentável. Tradução do Centro de Informação das Nações Unidas para o Brasil (UNIC Rio). Brasília, DF, 2016. Disponível em: [link]. Acesso em: 8 maio 2026.
BRASIL. Decreto nº 5.296, de 2 de dezembro de 2004. Regulamenta as Leis nº 10.048, de 8 de novembro de 2000, e nº 10.098, de 19 de dezembro de 2000. Diário Oficial da União: Brasília, DF, 3 dez. 2004. Disponível em: [link]. Acesso em: 11 maio 2026.
BRASIL. Decreto nº 6.949, de 25 de agosto de 2009. Promulga a Convenção Internacional sobre os Direitos das Pessoas com Deficiência e seu Protocolo Facultativo, assinados em Nova York, em 30 de março de 2007. Diário Oficial da União: Brasília, DF, 26 ago. 2009. Disponível em: [link]. Acesso em: 11 maio 2026.
BRASIL. Lei nº 10.048, de 8 de novembro de 2000. Dá prioridade de atendimento às pessoas que especifica, e dá outras providências. Diário Oficial da União: Brasília, DF, 9 nov. 2000. Disponível em: [link]. Acesso em: 11 maio 2026.
BRASIL. Lei nº 10.098, de 19 de dezembro de 2000. Estabelece normas gerais e critérios básicos para a promoção da acessibilidade das pessoas portadoras de deficiência ou com mobilidade reduzida, e dá outras providências. Diário Oficial da União: Brasília, DF, 20 dez. 2000. Disponível em: [link]. Acesso em: 11 maio 2026.
BRASIL. Lei nº 13.146, de 6 de julho de 2015. Institui a Lei Brasileira de Inclusão da Pessoa com Deficiência (Estatuto da Pessoa com Deficiência). Diário Oficial da União: seção 1, Brasília, DF, 7 jul. 2015. Disponível em: [link]. Acesso em: 11 maio 2026.
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020). Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
BYERLEY, Suzanne L.; BETH CHAMBERS, Mary. Accessibility of Web-based library databases: the vendors’ perspectives. Library Hi Tech, v. 21, n. 3, p. 347-357, 2003.
Caffagni, D., Cocchi, F., Barsellotti, L., Moratelli, N., Sarto, S., Baraldi, L., Cornia, M., and Cucchiara, R. (2024). The revolution of multimodal large language models: a survey. arXiv preprint arXiv:2402.12451.
Carvalho, L. P. (2024). Ética Computacional no Brasil: uma análise metacientífica da prática acadêmico-científico computacional brasileira. Tese de doutorado, Universidade Federal do Rio de Janeiro, Rio de Janeiro.
Caspar, E. A. (2021). A novel experimental approach to study disobedience to authority. Scientific Reports, 11(1):22927.
CASTELLS, Manuel. A Galáxia Internet: reflexões sobre a Internet, negócios e a sociedade. Zahar, 2003.
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. (2020). A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607. PmLR.
Cliquet, L. O. B. V., Pimentel, M. d. G. C., Batistoni, S. S. T., Rodrigues, K. R. d. H., Zaine, I., and Cachioni, M. (2023). Use of smartphones by older adults: characteristics and reports of students enrolled at a University of the Third Age (U3A). PerCursos, 24:1–30.
CNS (1996). Resolução CNS no 196, de 10 de outubro de 1996. Conselho Nacional de Saúde (Brasil). Diretrizes e normas regulamentadoras de pesquisas envolvendo seres humanos.
CNS (2012). Resolução CNS no 466, de 12 de dezembro de 2012. Conselho Nacional de Saúde (Brasil). Diretrizes e normas regulamentadoras de pesquisas envolvendo seres humanos.
CNS (2016). Resolução CNS no 510, de 7 de abril de 2016. Conselho Nacional de Saúde (Brasil). Define as especificidades éticas das pesquisas em Ciências Humanas e Sociais.
CNS (2022). Resolução CNS no 674, de 6 de maio de 2022. Conselho Nacional de Saúde (Brasil). Atualiza o Sistema CEP/CONEP e orienta sobre pesquisa com dados pessoais e digitais.
CONEP (2022). Ofício Circular nº 17, de 5 de julho de 2022: Orientações acerca do artigo 1º da Resolução CNS nº 510, de 7 de abril de 2016. Comissão Nacional de Ética em Pesquisa (CONEP). Atualizado em 23 abr. 2025.
COSTA, Rafael Rodrigues; OMAIA, Diego; ARAÚJO, Thiago Moreira; PEREIRA, José Otávio; COUTINHO, André Silva; CRUZ, Marcelo Pereira; FILHO, Gilson Luiz da Silva. Acessibilidade na TV 3.0 brasileira a partir de mídias de legenda, glosa e áudio descrição. In: Simpósio Brasileiro de Sistemas Multimídia e Web, 29., 2023, Juiz de Fora. Anais do Simpósio Brasileiro de Sistemas Multimídia e Web (WebMedia 2023). Porto Alegre: Sociedade Brasileira de Computação, 2023. p. 123-129.
COSTA, V.; TOMASZEWSKI, J. P.; PEREIRA, L.; SANTANA, B. S.; CORRÊA, G. AIBased Approaches for Brazilian Sign Language Recognition: A Systematic Literature Review: Insights on Methods, Metrics, and Resources for LIBRAS Recognition. In: SIMPÓSIO BRASILEIRO DE SISTEMAS MULTIMÍDIA E WEB (WEBMEDIA), 31., 2025. Anais [...]. Porto Alegre: SBC, 2025. p. 585-597.
CRESWELL, John W. Qualitative inquiry & research design: choosing among five approaches. 2. ed. Thousand Oaks: Sage Publications, 2007.
Cunha, B. C. R., Rodrigues, K. R. D. H., Zaine, I., da Silva, E. A. N., Viel, C. C., and Pimentel, M. D. G. C. (2021). Experience Sampling and Programmed Intervention Method and System for Planning, Authoring, and Deploying Mobile Health Interventions: Design and Case Reports. J Med Internet Res, 23(7):e24278.
Cunha, B. C., Rodrigues, K. R., and Pimentel, M. d. G. C. (2019). Synthesizing guidelines for facilitating elderly-smartphone interaction. In Proceedings of the 25th Brazillian Symposium on Multimedia and the Web, pages 37–44.
DA COSTA NUNES, Erik Henrique et al. Digital accessibility at the brazilian symposium on human factors in computing systems (IHC): An updated systematic literature review. In: Proceedings of the XXII Brazilian Symposium on Human Factors in Computing Systems. 2023. p. 1-15.
da Silva, N. B., Harrison, J., Minetto, R., Delgado, M. R., Nassu, B. T., and Silva, T. H. (2025). Multimodal LLMs see sentiment. arXiv preprint arXiv:2508.16873.
Dai, W., Lee, N., Wang, B., Yang, Z., Liu, Z., Barker, J., Rintamaki, T., Shoeybi, M., Catanzaro, B., and Ping, W. (2024). NVLM: Open frontier-class multimodal LLMs. arXiv preprint arXiv:2409.11402.
DE FÁTIMA GRANATTO, Cleusa; PALLARO, Marynea AP; BIM, Sílvia Amélia. Digital accessibility: systematic review of papers from the Brazilian symposium on human factors in computer systems. In: Proceedings of the 15th Brazilian Symposium on Human Factors in Computing Systems. 2016. p. 1-10.
Decreto12651 (2025). Decreto no 12.651, de 7 de outubro de 2025 – Regulamenta a Lei no 14.874, de 28 de maio de 2024, que dispõe sobre a pesquisa com seres humanos e institui o Sistema Nacional de Ética em Pesquisa com Seres Humanos. Brasil. Publicado no Diário Oficial da União em 8 de outubro de 2025.
Deitke, M., Clark, C., Lee, S., Tripathi, R., Yang, Y., Park, J. S., Salehi, et al. (2025). Molmo and PixMo: Open weights and open data for state-of-the-art vision-language models. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 91–104.
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers), pages 4171–4186.
Dickerson, C. (2015a). The VA’s Broken Promise To Thousands Of Vets Exposed To Mustard Gas. National Public Radio.
Dickerson, C. (2015b). U.S. Troops Were Tested by Race in Secret World War II Chemical Experiments. National Public Radio.
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929.
Driess, D., Xia, F., Sajjadi, M. S. M., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., Huang,W., Chebotar, Y., Sermanet, P., Duckworth, D., Levine, S., Vanhoucke, V., Hausman, K., Toussaint, M., Greff, K., Zeng, A., Mordatch, I., and Florence, P. (2023). PaLM-E: An embodied multimodal language model.
Dzabraev, M., Kunitsyn, A., and Ivaniuta, A. (2024). VLRM: Vision-language models act as reward models for image captioning. arXiv preprint arXiv:2404.01911.
Elizalde, B., Deshmukh, S., Al Ismail, M., and Wang, H. (2023). CLAP learning audio concepts from natural language supervision. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5. IEEE.
Eusebio, J. and Pimentel, M. (2025). Avaliação Retrospectiva da Experiência de uma Coorte de Egressas de um Curso de Introdução à Programação para Alunas do Ensino Médio e Concluintes. In Anais do XIX Women in Information Technology, pages 275–286, Porto Alegre, RS, Brasil. SBC.
Fei, H., Wu, S., Ji, W., Zhang, H., Zhang, M., Lee, M.-L., and Hsu, W. (2024). Video-of-thought: Step-by-step video reasoning from perception to cognition. arXiv preprint arXiv:2501.03230.
FIGUEIRA, Emílio. As Pessoas Com Deficiência na História do Brasil: uma trajetória de silêncio e gritos!. Wak, 2021.
FRANÇA, Lúcio Cauper Freitas de; ANDRADE, Lisieux Marie M. dos S.; VIANA, Windson. Um estudo sobre o uso de prompts LLM para a avaliação da acessibilidade web dos portais das prefeituras da região dos Sertões dos Crateús. In: SIMPÓSIO BRASILEIRO DE SISTEMAS MULTIMÍDIA E WEB (WEBMEDIA), 31., 2025. Anais [...]. Porto Alegre: SBC, 2025. p. 516-525.
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D. (2022). GPTQ: Accurate post-training quantization for generative pre-trained transformers. arXiv preprint arXiv:2210.17323.
GARRETT, Jesse James. The elements of user experience: user-centered design for the web and beyond. 2. ed. Berkeley: New Riders, 2011.
Gaśińska, A. (2015). The contribution of women to radiobiology: Marie Curie and beyond. Reports of Practical Oncology and Radiotherapy, 21(3):250–258.
Goldman, O., Shaham, U., Malkin, D., Eiger, S., Hassidim, A., Matias, Y., Maynez, J., Gilady, A. M., Riesa, J., Rijhwani, S., et al. (2025). ECLeKTic: a novel challenge set for evaluation of cross-lingual knowledge transfer. arXiv preprint arXiv:2502.21228.
Gong, L., Yang, J., Han, S., and Ji, Y. (2025). MedBLIP: A multimodal method of medical question-answering based on fine-tuning large language model. Computers in Medical Imaging and Graphics, 124:102581.
GOODMAN, Elizabeth; KUNIAVSKY, Mike; MOED, Andrea. Observing the user experience: a practitioner’s guide to user research. 2. ed. Amsterdam; Boston: Morgan Kaufmann Publishers, 2012.
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D. (2017). Making the V in VQA matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE International Conference on Computer Vision (ICCV).
Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Vaughan, A., et al. (2024). The LLaMA 3 herd of models. arXiv preprint arXiv:2407.21783.
Han, S., Huang, W., Shi, H., Zhuo, L., Su, X., Zhang, S., Zhou, X., Qi, X., Liao, Y., and Liu, S. (2025). Videoespresso: A large-scale chain-of-thought dataset for fine-grained video reasoning via core frame selection. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 26181–26191.
HENRY, S. L.; ABOU-ZAHRA, S.; BREWER, J. The role of accessibility in a universal web. In: WEB FOR ALL CONFERENCE, 11., 2014, Seoul, Coreia. Proceedings [...]. New York: Association for Computing Machinery, 2014. Artigo 17. DOI: 10.1145/2596695.2596719. Acesso em: 9 fev. 2026.
Hinton, G., Vinyals, O., and Dean, J. (2015). Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531.
HORTON, Sarah; QUESENBERY, Whitney. Uma web para todos: projetando experiências de usuário acessíveis . Rosenfeld Media, 2014.
Hossain, M. D., Sohel, F., Shiratuddin, M. F., and Laga, H. (2019). A comprehensive survey of deep learning for image captioning. ACM Computing Surveys (CSUR), 51(6):1–36.
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al. (2022). Lora: Low-rank adaptation of large language models. ICLR, 1(2):3.
ISO. ISO 9241-210:2019: ergonomics of human-system interaction — Part 210: Human-centred design for interactive systems. Geneva: International Organization for Standardization, 2019. Disponível em: [link]. Acesso em: 12 maio 2026.
Jegham, N., Abdelatti, M., Elmoubarki, L., and Hendawi, A. (2025). How hungry is AI? benchmarking energy, water, and carbon footprint of LLM inference. arXiv preprint arXiv:2505.09598.
JÚNIOR, J. M. L. de M.; ARAÚJO, M. de C. C.; FAÇANHA, A. R.; VIANA, W.; SÁNCHEZ, J. Uma ferramenta para a customização de ambientes virtuais para práticas de orientação e mobilidade. In: SIMPÓSIO BRASILEIRO DE SISTEMAS MULTIMÍDIA E WEB (WEBMEDIA), 28., 2022. Anais [...]. Porto Alegre: SBC, 2022. p. 103-106.
JÚNIOR, Mário Cléber Martins Lanna (Ed.). História do movimento político das pessoas com deficiência no Brasil. Presidência da República, Secretaria de Direitos Humanos, Secretaria Nacional de Promocão dos Direitos da Pessoa com Deficiência, 2010.
Kalai, A. T., Nachum, O., Vempala, S. S., and Zhang, E. (2025). Why language models hallucinate. arXiv preprint arXiv:2509.04664.
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.-J., Shamma, D. A., Bernstein, M. S., and Fei-Fei, L. (2017). Visual genome: Connecting language and vision using crowdsourced dense image annotations. International Journal of Computer Vision (IJCV), 123(1):32–73.
Kuang, J., Shen, Y., Xie, J., Luo, H., Xu, Z., Li, R., Li, Y., Cheng, X., Lin, X., and Han, Y. (2025). Natural language understanding and inference with MLLM in visual question answering: A survey. ACM Computing Surveys, 57(8):1–36.
Kühne, T. (2018). Nazi Morality, chapter 13, pages 215–229. John Wiley & Sons, Ltd.
LangChain Inc. (2025a). LangChain. [link]. Acesso em: 16 out. 2025.
LangChain Inc. (2025b). LangGraph. [link]. Acesso em: 16 out. 2025.
LangChain Inc. (2025c). LangSmith. [link]. Acesso em: 16 out. 2025.
LAZAR, Jonathan; GOLDSTEIN, Daniel F.; TAYLOR, Anne. Ensuring digital accessibility through process and policy. Waltham: Morgan Kaufmann, 2015.MOURA, Filipa Raquel Teixeira. Avaliação da Acessibilidade dos Sítios web das PME’s Portuguesas. 2009.
Lei14874 (2024). Lei no 14874, de 28 de maio de 2024 – Dispõe sobre a pesquisa com seres humanos e institui o Sistema Nacional de Ética em Pesquisa com Seres Humanos. Brasil.
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in neural information processing systems, 33:9459–9474.
LGPD (2018). Lei no 13.709, de 14 de agosto de 2018 – Lei Geral de Proteção de Dados Pessoais (LGPD). Brasil. Redação dada pela Lei no 13.853, de 2019.
Li, J., Bercea, C. I., Müller, P., Felsner, L., Kim, S. H., Rueckert, D., Wiestler, B., and Schnabel, J. A. (2024). Multi-image visual question answering for unsupervised anomaly detection. In CoRR, volume abs/3sJvMF5m5D. Preprint on OpenReview.
Li, J., Li, D., Savarese, S., and Hoi, S. (2023). BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, pages 19730–19742. PMLR.
Lima, A. P. d., Dantas, L. A., Manzato, M. G., Pimentel, M., Orlandi, B., and Castro, P. (2021). An Interpretable Recommendation Model for Gerontological Care. In Proceedings of the 15th ACM Conference on Recommender Systems, RecSys ’21, page 620–626, New York, NY, USA. Association for Computing Machinery.
LIMA, Manuella Aschoff; FRANÇA, Daniel de; SILVA NETO, Antônio da; SOUZA, Daniel Faustino L. de; OMAIA, Diego; ARAÚJO, Tiago Maritan Ugulino de. LibrasDetec: um componente de detecção de movimentos para karaokê em Libras. In: SIMPÓSIO BRASILEIRO DE SISTEMAS MULTIMÍDIA E WEB (WEBMEDIA), 31., 2025. Anais [...]. Porto Alegre: SBC, 2025. p. 340-348.
Liu, H., Li, C., Wu, Q., and Lee, Y. J. (2023). Visual instruction tuning. Advances in neural information processing systems, 36:34892–34916.
Liu, Y., Liang, Z., Wang, Y., Wu, X., Tang, F., He, M., Li, J., Liu, Z., Yang, H., Lim, S., et al. (2025). Unveiling the ignorance of mLLMs: Seeing clearly, answering incorrectly. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 9087–9097.
Longo, L., Brcic, M., Cabitza, F., Choi, J., Confalonieri, R., Ser, J. D., Guidotti, R., Hayashi, Y., Herrera, F., Holzinger, A., Jiang, R., Khosravi, H., Lecue, F., Malgieri, G., Páez, A., Samek, W., Schneider, J., Speith, T., and Stumpf, S. (2024). Explainable artificial intelligence (XAI) 2.0: A manifesto of open challenges and interdisciplinary research directions. Information Fusion, 106:102301.
Lopes, C. R., Minetto, R., Delgado, M. R., and Silva, T. H. (2023). PerceptSent - exploring subjectivity in a novel dataset for visual sentiment analysis. IEEE Transactions on Affective Computing, 14(3).
LOPES, Renan; FAÇANHA, Agebson Rocha; VIANA, Windson. I can’t pay! Accessibility analysis of mobile banking apps. In: Proceedings of the Brazilian Symposium on Multimedia and the Web. 2022. p. 253-257.
MAGALHÃES, Gabriel; FONSECA, Davi; AMORIM, Glauco; VIANA, Windson; SANTOS, Joel dos. Designing a Multimodal Feedback Component for VR Applications. In: BRAZILIAN SYMPOSIUM ON MULTIMEDIA AND THE WEB (WEBMEDIA), 31., 2025, Rio de Janeiro, RJ. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2025. p. 194-201.
Mandal, J., Acharya, S., and Parija, S. C. (2011). Ethics in human research. Tropical Parasitology, 1(1):2–3.
MARTINS, Luisa; FERREIRA, Stênio; MARITAN, Tiago. Automatic 3D animation generation for Sign Language: A case study of Sign Language Machine Translation. In: Brazilian Symposium on Multimedia and the Web (WebMedia). SBC, 2025. p. 67-76.
Mellanby, K. (1947). Medical Experiments on Human Beings in Concentration Camps in Nazi Germany. British Medical Journal, 1(4490):148–150. Published 25 January 1947.
Mersha, M., Lam, K., Wood, J., AlShami, A. K., and Kalita, J. (2024). Explainable artificial intelligence: A survey of needs, techniques, applications, and future direction. Neurocomputing, 599:128111.
Milgram, S. (1963). Behavioral study of obedience. Journal of Abnormal and Social Psychology, 67(4):371–378.
MITTAL, Anant et al. Jod: Examining design and implementation of a videoconferencing platform for mixed hearing groups. In: Proceedings of the 25th International ACM SIGACCESS Conference on Computers and Accessibility. 2023. p. 1-18.
Mondshine, I., Paz-Argaman, T., and Tsarfaty, R. (2025). Beyond english: The impact of prompt translation strategies across languages and tasks in multilingual LLMs. arXiv preprint arXiv:2502.09331.
MONTEIRO, I. T. Accessibility by mediation dialogues: development and evaluation of a web navigation helper. 2011. Dissertação (Mestrado em Informática) — Pontifícia Universidade Católica do Rio de Janeiro, Rio de Janeiro, 2011.
MORAES JÚNIOR, José Martônio Lopes de; ARAÚJO, Maria do Carmo Cavalcante; FAÇANHA, Agebson Rocha; VIANA, Windson. Virtual Reality for People with Visual Impairments in the Orientation and Mobility Context. In: SIMPÓSIO BRASILEIRO DE SISTEMAS MULTIMÍDIA E WEB (WEBMEDIA), 30., 2024. Anais Estendidos [...]. Porto Alegre: SBC, 2024.
MORAES, J. M.; CARMO FILHO, P.; ARAÚJO, M. de C. C.; COSTA, B.; AMORIM, M.; FAÇANHA, A. R.; et al. Dashboards to support rehabilitation in orientation and mobility for people who are blind. In: SIMPÓSIO BRASILEIRO DE SISTEMAS MULTIMÍDIA E WEB (WEBMEDIA), 29., 2023. Anais [...]. Porto Alegre: SBC, 2023. p. 6-10.
MOURA, Filipa Raquel Teixeira. Avaliação da Acessibilidade dos Sítios web das PME’s Portuguesas. 2009.
Museum, U. S. H. M. (2025). Nazi medical experiments. Accessed Feb. 28, 2026.
NIC.br. Cartilha de acessibilidade na Web: fascículo IV: tornando o conteúdo Web acessível [livro eletrônico]. Coordenação de Reinaldo Ferraz. Ilustrações de Mônica Lopes. São Paulo: Comitê Gestor da Internet no Brasil, 2020. Disponível em: [link]. Acesso em: 9 maio 2026.
NIC.br. Pesquisa sobre o uso das tecnologias de informação e comunicação nos domicílios brasileiros: TIC Domicílios 2024 [livro eletrônico]. São Paulo: Comitê Gestor da Internet no Brasil, 2025. Disponível em: [link]. Acesso em: 8 maio 2026.
NORMAN, Don. The design of everyday things. Revised and expanded edition. New York: Basic Books, 2013.
Oliveira, W. B. d., Dorini, L. B., Minetto, R., and Silva, T. H. (2020). OutdoorSent: Sentiment analysis of urban outdoor images by using semantic and deep features. ACM Trans. on Information Systems, 38(3).
OpenAI (2023a). GPT-4 technical report. arXiv preprint arXiv:2303.08774.
OpenAI (2023b). GPT-4V(ision) system card. [link]. Acesso em: 3 dez. 2025.
OSTERWALDER, Alexander; PIGNEUR, Yves; BERNARDA, Gregory; SMITH, Alan; PAPADAKOS, Trish. Value proposition design: how to create products and services customers want. Hoboken: John Wiley & Sons, 2014.
Pal, R., Kar, S., Prasad, D. K., and Sekh, A. A. (2025). Multilingual visual question answering for visually impaired people. Discover Artificial Intelligence, 5.
PARDINI, Rafael; BÁRBARA, João; SCHEID, Henrique; PEREIRA, Ana Carolina; MEIRA JÚNIOR, Wagner; FERRAZ, Rogério; ROCHA, Bruno. Observatório da acessibilidade da web brasileira. In: SIMPÓSIO BRASILEIRO DE SISTEMAS MULTIMÍDIA E WEB (WEBMEDIA), 27., 2021. Anais Estendidos [...]. Porto Alegre: SBC, 2021. p. 71-74.
PARTHASARATHY, P. D. et al. Skill, will, or both? understanding digital inaccessibility from accessibility professionals’ viewpoint. In: Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems. 2025. p. 1-9.
Pedrosa, D., Pimentel, M. d. G., and Truong, K. N. (2015a). Filteryedping: A dwell-free eye typing technique. In Proceedings of the 33rd annual ACM Conference Extended Abstracts on Human Factors in Computing Systems, pages 303–306.
Pedrosa, D., Pimentel, M. D. G., Wright, A., and Truong, K. N. (2015b). Filteryedping: Design Challenges and User Performance of Dwell-Free Eye Typing. ACM Trans. Access. Comput., 6(1).
Pereira, A. P., Machado Neto, O. J., Elui, V. M. C., and Pimentel, M. d. G. C. (2025). Wearable Smartphone-Based Multisensory Feedback System for Torso Posture Correction: Iterative Design and Within-Subjects Study. JMIR Aging, 8:e55455.
Pimentel, M. d. G. C., Freire, A. P., Carvalho, L. P., and Rodrigues, K. R. d. H. (2023). Submissão de Projetos de PesquisaWeb e Multimídia a Comitês de Ética em Pesquisa. In Marques Neto, M. C., Pimentel, M. d. G., Willrich, R., Macedo, A. A., and Goularte, R., editors, Minicursos do XXIX Simpósio Brasileiro de Sistemas Multimídia e Web. Sociedade Brasileira de Computação, Porto Alegre, RS, Brasil.
Pimentel, M. d. G. C., Pereira, A. P., Machado Neto, O. J., Zimmermann, L. C., and Elui, V. M. C. (2025). Quantitative Evaluation of Postural Smart-Vest’s Multisensory Feedback for Affordable Smartphone-Based Post-Stroke Motor Rehabilitation. International Journal of Environmental Research and Public Health, 22(7).
PÓVOA, Marcello. Anatomia da Internet: investigações estratégicas sobre o universo digital. Casa da Palavra, 2000.
PREECE, Jennifer; ROGERS, Yvonne; SHARP, Helen. Interaction design: beyond humancomputer interaction. New York: John Wiley & Sons, 2002.
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. (2021). Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PmLR.
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al. (2018). Improving language understanding by generative pre-training. OpenAI blog.
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. (2019). Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
Raschka, S. (2024). Understanding Multimodal LLMs. [link]. Acesso em: 16 out. 2025.
RIBEIRO, Carlos Nery; OLIVEIRA, Cynthya Letícia Teles de; ARAUJO, Lucas Padilha Modesto de; RODRIGUES, Kamila Rios da Hora; MANZATO, Marcelo Garcia. ARIA e Interactive Access: projetando chatbots para idosos. In: CONCURSO DE TRABALHOS DE INICIAÇÃO CIENTÍFICA (CTIC), 4., 2024. Anais Estendidos [...]. Porto Alegre: SBC, 2024. p. 41-44.
Riedel, S. (2005). Edward Jenner and the History of Smallpox and Vaccination. Proceedings (Baylor University Medical Center), 18(1):21–25.
ROBERTSON, Toni; SIMONSEN, Jesper. Challenges and opportunities in contemporary participatory design. Design Issues, Cambridge, v. 28, n. 3, p. 3-9, 2012.
Rodrigues, K. R. d. H., Silva Junior, J. M., Nespule, L., Maia, M. S., Garcia, L. E., and Galati, H. B. (2020). Lume: Um jogo digital para incentivar a leitura de lendas do folclóre brasileiro. In Anais do XIX Simpósio Brasileiro de Jogos e Entretenimento Digital, pages 1240–1249.
Rodrigues, K. R., Cunha, B. C., Zaine, I., Viel, C. C., Scalco, L. F., and Pimentel, M. G. (2018). ESPIM system: interface evolution to enable authoring and interaction with multimedia intervention programs. In Proceedings of the 24th Brazilian Symposium on Multimedia and the Web, pages 125–132.
Rodrigues, K. R., Pereira, S. S., Quinelato, L. G., Melo, E. L., Neris, V. P., and Teixeira, C. A. (2011). Enhancing TV experience with additional content and multiple device interactivity. In Proceedings of the 17th Brazilian Symposium on Multimedia and the Web on Brazilian Symposium on Multimedia and the Web - Volume 1, WebMedia 2011, page 18–25, Porto Alegre, BRA. Brazilian Computer Society.
Rodrigues, K., Carvalho, L., Freire, A., and Pimentel, M. (2024a). GranDIHC-BR 2025-2035 - GC2: Ethics and Responsibility: Principles, Regulations, and Societal Implications of Human Participation in HCI Research. In Anais do XXIII Simpósio Brasileiro sobre Fatores Humanos em Sistemas Computacionais, pages 969–987, Porto Alegre, RS, Brasil. SBC.
RODRIGUES, Kamila Rios da Hora; SANTOS, Suzane Santos dos; GALLEGO, Daniele; MARTINS, Ketlen; MALPARTIDA, Katherin Felipa Carhuaz; VERHALEN, Aline Elias Cardoso; DEUS, João Pedro de. Práticas com smartphones para idosos: um projeto de extensão do ICMC/USP. In: WEBMEDIA FOR EVERYONE (W4E), 3., 2024. Anais Estendidos [...]. Porto Alegre: SBC, 2024. p. 239-243.
Rodrigues, S., Fortes, R., and Rodrigues, K. (2024b). Guidelines for designing iot applications for older adults. In Anais do XXIII Simpósio Brasileiro sobre Fatores Humanos em Sistemas Computacionais, pages 472–486, Porto Alegre, RS, Brasil. SBC.
ROHRER, Christian. When to Use Which User-Experience Research Methods. Fremont: Nielsen Norman Group, 2014.
Sanches, R., Ponti, M., and Rodrigues, K. (2022). Evasão universitária e estratégias para retenção de alunos com base em intervenções remotas. In Anais Estendidos do XXI Simpósio Brasileiro sobre Fatores Humanos em Sistemas Computacionais, pages 84–87, Porto Alegre, RS, Brasil. SBC.
SCHULER, Douglas; NAMIOKA, Aki (org.). Participatory design: principles and practices. Hillsdale: Lawrence Erlbaum Associates, 1993.
Shuster, E. (1997). Fifty years later: The significance of the nuremberg code. The New England Journal of Medicine, 337(20):1436–1440.
SILVA, Marcos André Bezerra da; LIMA, Manuella Aschoff C. B.; SILVA, Diego Ramon Bezerra da; SOUZA, Daniel Faustino L. de; SILVA, Samuel de Moura Moreira da; ARAÚJO, Tiago Maritan Ugulino de. Uma investigação sobre técnicas de data augmentation aplicadas à tradução automática Português-LIBRAS. In: SIMPÓSIO BRASILEIRO DE SISTEMAS MULTIMÍDIA E WEB (WEBMEDIA), 30., 2024. Anais [...]. Porto Alegre: SBC, 2024. p. 319-325.
SOCIEDADE BRASILEIRA DE COMPUTAÇÃO (SBC). Grandes Desafios da Computação no Brasil 2025–2035. Coordenação: André Luís de Medeiros Santos e Flávio Rech Wagner. Porto Alegre: Sociedade Brasileira de Computação, 2025.
STICKDORN, Marc; HORMESS, Markus; LAWRENCE, Adam; SCHNEIDER, Jakob. This is service design doing: applying service design thinking in the real world. Sebastopol: O’Reilly Media, 2018.
Tavares, Daniela Cardoso; Penha, Marcio Rogério; Borges, José Antonio dos Santos; Dias, Angélica Fonseca da Silva; Ferreira, Thiago de Melo; "GOOGLEVOX: UMA INTERFACE ADAPTATIVA PARA ALAVANCAR A INTERAÇÃO DE DEFICIENTES VISUAIS EM PESQUISAS NO GOOGLE", p-472-482. In: . São Paulo: Blucher, 2017.
TAVARES, Daniela Cardoso. Vídeos mediadores do conhecimento: uma abordagem para a comunicação acessível a pessoas com deficiência visual. 2019. Dissertação de Mestrado. Instituto Politecnico de Leiria (Portugal).
Team, G., Anil, R., Borgeaud, S., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., Millican, K., et al. (2023). Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805.
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al. (2023). LLaMA: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
Varkey, B. (2021). Principles of clinical ethics and their application to practice. Medical Principles and Practice, 30(1):17–28.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30.
Verhalen, A. E. C. and Rodrigues, K. R. d. H. (2024). The use of serious digital games to talk about grief and finitude with children: reports on the design and evaluation process of two games. Journal on Interactive Systems, 15(1):926–941.
VIEIRA, Alex de Souza; GUEDES, Álan Lívio V.; MORAES, Daniel de Sousa; MADEIRA, Lucas Ribeiro; COLCHER, Sérgio; SOARES NETO, Carlos de Souza. ListeningTV: Accessible Video using Interactive Audio Descriptions. In: Simpósio Brasileiro de Sistemas Multimídia e Web, 26., 2020, São Luís. Anais Estendidos do Simpósio Brasileiro de Sistemas Multimídia e Web (WebMedia 2020). Porto Alegre: Sociedade Brasileira de Computação, 2020. p. 71–74. DOI: 10.5753/webmedia_estendido.2020.13065.
Vinyals, O., Toshev, A., Bengio, S., and Erhan, D. (2015). Show and tell: A neural image caption generator. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 3156–3164.
W3C BRASIL. Cartilha acessibilidade na Web: fascículo 2: benefícios, legislação e diretrizes da acessibilidade na Web. São Paulo: Comitê Gestor da Internet no Brasil, 2015. Livro eletrônico. ISBN 978-85-5559-008-5. Disponível em: [link]. Acesso em: 11 maio 2026.
WACNIK, Peter; DALY, Shanna; VERMA, Aditi. Participatory design: a systematic review and insights for future practice. In: Proceedings of the Design Society. Cambridge: Cambridge University Press, 2023. v. 3, p. 2285–2294. DOI: 10.1017/pds.2023.229.
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D. (2022). Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al. (2022a). Emergent abilities of large language models. arXiv preprint arXiv:2206.07682.
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. (2022b). Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837.
WMA (2024). (World Medical Association) Declaration of Helsinki: Ethical Principles for Medical Research Involving Human Participants. World Medical Association.
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al. (2020). Transformers: State-of-theart natural language processing. In Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations, pages 38–45.
Wolfe, C. R. (2024). Decoder-only transformers: The workhorse of generative LLMs. [link]. Acesso em: out. 2025.
WORLD WIDE WEB CONSORTIUM (W3C). Diretrizes de Acessibilidade para Conteúdo Web (WCAG) 2.2. Recomendação do W3C, 12 dez. 2024. Disponível em: [link]. Acesso em: 11 maio 2026.
Wu, J., Gan, W., Chen, Z., Wan, S., and Yu, P. S. (2023). Multimodal large language models: A survey. In 2023 IEEE International Conference on Big Data (BigData), pages 2247–2256. IEEE.
XAVIER, L. E. O.; RABELO, D. M. F.; MARTINS, R. de S.; VIANA, W. Image-to-UI Accessibility: Assessing ChatGPT’s Effectiveness in Generating Accessible Android Screens from Figma Templates. In: SIMPÓSIO BRASILEIRO DE SISTEMAS MULTIMÍDIA E WEB (WEBMEDIA), 31., 2025. Anais [...]. Porto Alegre: SBC, 2025. p. 302-311.
Xiao, G., Lin, J., Seznec, M.,Wu, H., Demouth, J., and Han, S. (2023). SmoothQuant: Accurate and efficient post-training quantization for large language models. In International conference on machine learning, pages 38087–38099. PMLR.
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K. R., and Cao, Y. (2022). ReAct: Synergizing reasoning and acting in language models. In The eleventh international conference on learning representations.
Yin, S., Fu, C., Zhao, S., Li, K., Sun, X., Xu, T., and Chen, E. (2024). A survey on multimodal large language models. National Science Review, 11(12):nwae403.
Zaine, I., Frohlich, D. M., Rodrigues, K. R. D. H., Cunha, B. C. R., Orlando, A. F., Scalco, L. F., and Pimentel, M. D. G. C. (2019). Promoting Social Connection and Deepening Relations Among Older Adults: Design and Qualitative Evaluation of Media Parcels. J Med Internet Res, 21(10):e14112.
Zhang, H., Li, X., and Bing, L. (2023). Video-LLaMA: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858.
Zhang, X., Chen, Y., Yeh, S., and Li, S. (2025). MetaMind: Modeling human social thoughts with metacognitive multi-agent systems. arXiv preprint arXiv:2505.18943.
Zhu, D., Chen, J., Shen, X., Li, X., and Wang, X. (2023). MiniGPT-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592.

