Ir directamente a la navegación principal Ir directamente a la búsqueda Ir directamente al contenido principal

Visual Question Answering Using Semantic Information from Image Descriptions

Producción científica: Conference articlerevisión exhaustiva

Resumen

In this work, we propose a deep neural architecture that uses an attention mechanism which utilizes region based image features, the natural language question asked, and semantic knowledge extracted from the regions of an image to produce open-ended answers for questions asked in a visual question answering (VQA) task. The combination of both region based features and region based textual information about the image bolsters a model to more accurately respond to questions and potentially do so with less required training data. We evaluate our proposed architecture on a VQA task against a strong baseline and show that our method achieves excellent results on this task.

Idioma originalEnglish
PublicaciónProceedings of the International Florida Artificial Intelligence Research Society Conference, FLAIRS
Volumen34
DOI
EstadoPublished - 2021
Evento34th International Florida Artificial Intelligence Research Society Conference, FLAIRS-34 2021 - North Miami Beach, United States
Duración: may 16 2021may 19 2021

Nota bibliográfica

Publisher Copyright:
© 2021by the authors. All rights reserved.

ASJC Scopus subject areas

  • Software
  • Artificial Intelligence

Huella

Profundice en los temas de investigación de 'Visual Question Answering Using Semantic Information from Image Descriptions'. En conjunto forman una huella única.

Citar esto