Дослідження архітектур нейронних мереж для обробки мультимодальних даних
Main Article Content
Abstract
У роботі проведено дослідження сучасних архітектур нейронних мереж, призначених для оброблення мультимодальних даних, що включають текстові та візуальні компоненти. Розглянуто основні підходи до інтеграції модальностей, зокрема раннє, пізнє та гібридне злиття, а також контрастивне навчання та крос-модальну увагу. Виконано порівняльний аналіз моделей CLIP, ViLT, LXMERT та Perceiver IO за критеріями ефективністю, швидкістю навчання, масштабованістю та інтерпретованістю результатів. Результати дослідження демонструють переваги гібридних та контрастивних підходів у задачах мультимодальної класифікації та прогнозування, а також перспективи застосування таких моделей у медичній діагностиці, автономних системах та інтелектуальних системах рекомендацій.
Article Details
References
Wang, Z. SimVLM: Simple Visual Language Model Pretraining with Weak Supervision / Z. Wang, J. Yu, A. W. Yu [et al.] // arXiv. – 2021. https://doi.org/10.48550/arXiv.2108.10904.
Мінухін С.В. Дослідження моделі сегментації зображень з використанням розподілених режимів TensorFlow та згорткової нейронної мережі U-Net / С. В. Мінухін. // Системи обробки інформації. - 2020. - № 1(160). - C. 115-122. https://doi.org/10.30748/soi.2020.160.15. DOI: https://doi.org/10.30748/soi.2020.160.15
Bai, Y. Sequential Modeling Enables Scalable Learning for Large Vision Models / Y. Bai, X. Geng, K. Mangalam [et al.] // arXiv. – 2022. https://doi.org/10.48550/arXiv.2206.11389.
Dosovitskiy, A. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale / A. Dosovitskiy, L. Beyer, A. Kolesnikov [et al.] // International Conference on Learning Representations. – 2021. https://doi.org/10.48550/arXiv.2010.11929.
Garcia, N. A Survey on Vision-Language Pre-training: Evolution, Foundations and Applications / N. Garcia, C. Ye, Z. Liu [et al.] // IEEE Transactions on Pattern Analysis and Machine Intelligence. – 2023. – Vol. 45, No. 9. – P. 10781–10800. https://doi.org/10.1109/TPAMI.2023.3261989.
Kim, W. ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision / W. Kim, B. Son, I. Kim // Proceedings of the 38th International Conference on Machine Learning. – 2021. –
Vol. 139. – P. 5583–5594. https://doi.org/10.48550/arXiv.2102.03334.
Li, J. BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation / J. Li, D. Li,
C. Xiong, S. Hoi // Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. – 2022. – P. 12808–12817. https://doi.org/10.48550/arXiv.2201.12086.
Liu, Y. A Closer Look at Invariances in Self-Supervised Pre-training for Vision-Language Models / Y. Liu, Z. Wu, Z. Kira, Y. Wang // Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. – 2022. – P. 15378–15388. https://doi.org/10.48550/arXiv.2205.13135.
Zeng, Z. Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers / Z. Zeng, Z. Huang, B. Liu [et al.] // arXiv. – 2020. https://doi.org/10.48550/arXiv.2004.00849.
Radford, A. Learning Transferable Visual Models from Natural Language Supervision / A. Radford, J. W. Kim, C. Hallacy [et al.] // Proceedings of the 38th International Conference on Machine Learning. – 2021. – Vol. 139. – P. 8748–8763. https://doi.org/10.48550/arXiv.2103.00020.
Alayrac, J.-B. Flamingo: a Visual Language Model for Few-Shot Learning / J.-B. Alayrac, J. Donahue, P. Luc [et al.] // Advances in Neural Information Processing Systems. – 2022. – Vol. 35. – P. 23716–23736. https://doi.org/10.48550/arXiv.2204.14198.
Yu, J. CoCa: Contrastive Captioners are Image-Text Foundation Models / J. Yu, Z. Wang, V. Vasudevan [et al.] // Transactions on Machine Learning Research. – 2022. https://doi.org/10.48550/arXiv.2205.01917.
Zhang, Y. Contrastive Learning of Medical Visual Representations from Paired Images and Text / Y. Zhang, H. Jiang, Y. Miura [et al.] // Proceedings of the Conference on Health, Inference, and Learning. – 2022. – P. 2–25. https://doi.org/10.48550/arXiv.2110.07931.
Zhou, L. Unified Vision-Language Pre-Training for Image Captioning and VQA / L. Zhou, H. Palangi, L. Zhang [et al.] // Proceedings of the AAAI Conference on Artificial Intelligence. – 2020. – Vol. 34, No. 07. – P. 13041–13049. https://doi.org/10.48550/arXiv.1909.11059. DOI: https://doi.org/10.1609/aaai.v34i07.7005
