Title of article
Envisioning Answers: Unleashing Deep Learning for Visual Question Answering in Artistic Images
Author/Authors
Zolghadriha ، Erfan Deep Learning Research Lab, Department of Computer Engineering - Faculty of Engineering, College of Farabi - University of Tehran , Fouladi-Ghaleh ، Kazim Department of Computer Engineering - Faculty of Engineering, College of Farabi - University of Tehran , Ardehkhani ، Pouya Deep Learning Research Lab, Department of Computer Engineering - Faculty of Engineering, College of Farabi - University of Tehran
From page
191
To page
202
Abstract
In specialized fields, the accurate answering of visual questions is crucial for practical applications, and this study focuses on improving a visual question-answering model for artistic images by utilizing a dataset with both visual and knowledge-based questions. The approach involves employing a pre-trained BERT model to understand query nature and using the iQAN model with MLB and MUTAN mechanisms for visual queries, along with an XLNet-based model for knowledge-based information. The results demonstrate a 78.92% accuracy for visual questions, 47.71% for knowledge-based questions, and an overall accuracy of 55.88% by combining both branches. Additionally, the research explores the impact of parameters like the number of glances and activation functions on the model’s performance.
Keywords
Art Pictures , Visual Question Answering (VQA) , Natural Language Processing (NLP) , Computer Vision , Attention
Journal title
AUT Journal of Electrical Engineering
Journal title
AUT Journal of Electrical Engineering
Record number
2773952
Link To Document