Mordal: Automated Pretrained Model Selection for Vision Language Models 文章

ArXiv CS.CV2026-06-17NEWSen作者: Shiqi He, Insu Jang, Mosharaf Chowdhury

详细信息

来源站点
ArXiv CS.CV
作者
Shiqi He, Insu Jang, Mosharaf Chowdhury
文章类型
NEWS
语言
en
发布日期
2026-06-17

摘要

arXiv:2502.00241v2 Announce Type: replace-cross Abstract: Incorporating multiple modalities into large language models (LLMs) is a powerful way to enhance their understanding of non-textual data, enabling them to perform multimodal tasks. Vision language models (VLMs) form the fastest growing category of multimodal models because of their many practical use cases, including in healthcare, robotics, and accessibility. Unfortunately, even though different VLMs in the literature demonstrate impressive visual capabilities in different benchmarks, they are handcrafted by human experts; there is no automated framework to create task-specific multimodal models. We introduce Mordal, an automated multimodal model search framework that efficiently finds the best VLM for a user-defined task without manual intervention. Mordal achieves this both by reducing the number of candidates to consider during the search process and by minimizing the time required to evaluate each remaining candidate.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据