Wang, Y., Azadani, M. N., Sedwards, S., & Czarnecki, K. (2025). Leo-mini: An efficient multimodal large language model using conditional token reduction and mixture of multi-modal experts Presented at the Conference on Empirical Methods in Natural Language Processing (EMNLP) conference. Suzhou, China. https://doi.org/10.18653/v1/2025.emnlp-main.368 (Original work published 2025)
Reference author: Mozhgan Azadani
First name
Mozhgan
Middle name
Nasr
Last name
Azadani
Azadani, M. N., Riddell, J., Sedwards, S., & Czarnecki, K. (2026). Rethinking the Mixture of Vision Encoders Paradigm for Enhanced Visual Understanding in Multimodal LLMs Transactions on Machine Learning Research. Retrieved from https://openreview.net/pdf?id=tgnTVmRybs (Original work published 2026)
Wang, Y., Azadani, M. N., Sedwards, S., & Czarnecki, K. (2025). Hawaii: Hierarchical Visual Knowledge Transfer for Efficient Vision-Language Models Presented at the Advances in Neural Information Processing Systems (NeurIPS) conference. San Diego, USA: Curran Associates, Inc. Retrieved from https://proceedings.neurips.cc/paper_files/paper/2025/file/018af723a54a6d3bc5706d3a3126abf0-Paper-Conference.pdf (Original work published 2025)