Scinovex
Physical Sciences → Computer Science → Computer Vision and Pattern Recognition

Multimodal Machine Learning Applications

This cluster of papers focuses on the development and improvement of visual question answering systems, image captioning techniques, and neural networks for understanding and generating descriptions of images and videos. The research involves semantic reasoning, multimodal fusion, scene graph generation, attention mechanisms, and deep learning approaches to bridge the gap between vision and language.

67.4K works worldwide704.7K citations
Visual Question AnsweringImage CaptioningNeural NetworksSemantic ReasoningMultimodal FusionScene Graph GenerationVideo DescriptionAttention MechanismLanguage UnderstandingDeep Learning

Journals publishing in this area

1Neurocomputing cover
Neurocomputing
ISSN 0925-2312781 articles in this topic
253h-index
2IEEE Transactions on Pattern Analysis and Machine Intelligence cover
547h-index
3Pattern Recognition cover
Pattern Recognition
ISSN 0031-3203606 articles in this topic
300h-index
4IEEE Transactions on Image Processing cover
IEEE Transactions on Image Processing
ISSN 1057-7149538 articles in this topic
391h-index
5
International Journal of Computer Vision
ISSN 0920-5691342 articles in this topic
287h-index
6Neural Networks cover
Neural Networks
ISSN 0893-6080253 articles in this topic
246h-index