Multimodal NLP
Vision-language models that read and reason about memes, including hateful and propagandistic content across several languages.
Research Assistant · QCRI
محمد بيان قميناسي
I'm a Research Assistant at the Qatar Computing Research Institute (QCRI), where I work on natural language processing and multimodal models. Most of my research looks at how models can read and reason about memes, and how to detect harmful content such as hate speech and propaganda, with a particular focus on Arabic.
Lately I've been using reinforcement learning (GRPO) and chain-of-thought methods to help language and vision-language models reason better and explain their decisions. I finished my MSc in Computing at Qatar University in early 2026, after a BSc in Computer Engineering where I graduated first in my class. My work has appeared at NAACL, EMNLP, The Web Conference, and WISE.
Vision-language models that read and reason about memes, including hateful and propagandistic content across several languages.
Detecting harmful and propagandistic content in a way that explains the reasoning, not just the label.
Datasets, benchmarks, and multilingual models that make NLP work better for Arabic and other low-resource languages.
Using GRPO and chain-of-thought supervision to post-train models for stronger reasoning and explainability.
Selected papers below. The full, up-to-date list lives on my Google Scholar profile.
Preprint · ACL 2026 (under review) A multilingual, multimodal benchmark and models for understanding memes.
@article{shahroor2026memelens,
title = {MemeLens: Multilingual Multitask VLMs for Memes},
author = {Shahroor, Ali Ezzat and Kmainasi, Mohamed Bayan and Hasnat, Abul and Dimitrov, Dimitar and Da San Martino, Giovanni and Nakov, Preslav and Alam, Firoj},
journal = {arXiv preprint arXiv:2601.12539},
year = {2026}
}
WWW 2026 Companion Do reasoning ("thinking") models actually help for hateful-meme detection?
Preprint GRPO with chain-of-thought supervision for explainable multimodal detection.
Preprint · ACL 2026 demo (under review) A tool for critical digital literacy and resilience against misinformation.
NAACL 2025 Findings An instruction-tuned multilingual LLM for analyzing news and social-media content.
@inproceedings{kmainasi2025llamalens,
title = {LlamaLens: Specialized Multilingual LLM for Analyzing News and Social Media Content},
author = {Kmainasi, Mohamed Bayan and Shahroor, Ali Ezzat and Hasanain, Maram and Laskar, Sahinur Rahman and Hassan, Naeemul and Alam, Firoj},
booktitle = {Findings of the Association for Computational Linguistics: NAACL 2025},
pages = {5642--5664},
year = {2025},
url = {https://aclanthology.org/2025.findings-naacl.313/}
}
EMNLP 2025 Explainable multimodal detection, with a dedicated instruction dataset.
@inproceedings{kmainasi2025memeintel,
title = {MemeIntel: Explainable Detection of Propagandistic and Hateful Memes},
author = {Kmainasi, Mohamed Bayan and Hasnat, Abul and Hasan, Md. Arid and Shahroor, Ali Ezzat and Alam, Firoj},
booktitle = {Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing},
year = {2025},
url = {https://aclanthology.org/2025.emnlp-main.1539/}
}
EMNLP 2025 Findings Using LLMs to make propaganda detection explainable.
@inproceedings{hasanain2025propxplain,
title = {PropXplain: Can LLMs Enable Explainable Propaganda Detection?},
author = {Hasanain, Maram and Hasan, Md. Arid and Kmainasi, Mohamed Bayan and Sartori, Elisa and Shahroor, Ali Ezzat and Da San Martino, Giovanni and Nakov, Preslav and Alam, Firoj},
booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2025},
pages = {23855--23863},
year = {2025},
url = {https://aclanthology.org/2025.findings-emnlp.1296/}
}
Preprint · ICML 2026 (under review) Culturally grounded spoken visual question answering.
Preprint LLMs for Arabic legal judgment prediction.
WISE 2024 My most-cited paper: does prompting in your native language help?
@inproceedings{kmainasi2024native,
title = {Native vs Non-Native Language Prompting: A Comparative Analysis},
author = {Kmainasi, Mohamed Bayan and Khan, Rakif and Shahroor, Ali Ezzat and Bendou, Boushra and Hasanain, Maram and Alam, Firoj},
booktitle = {Web Information Systems Engineering -- WISE 2024},
pages = {406--420},
year = {2024},
url = {https://doi.org/10.1007/978-981-96-0576-7_30}
}
IEEE · JIBEC 2024 Machine learning for lung-cancer level detection from lifestyle data.
Work on the Fanar speech project training large multilingual speech LLMs, with papers at EMNLP 2025 and WWW 2026 and more under review. Also contributed to EduLLM on QA and testing.
First-author papers at WISE 2024 and NAACL 2025. Placed 2nd in the QCRI Programming Contest 2024.
Built an internal safety app for incident reporting and team communication.
Datasets, models, and systems I've released with my research. Most are public on GitHub and Hugging Face.
Specialized multilingual LLM with instruction datasets in Arabic, English, and Hindi for analyzing news and social media.
Multilingual, multimodal benchmark and vision-language models for understanding memes.
Explainable detection of propagandistic and hateful memes, released with the MemeXplain dataset.
Explainable propaganda detection with LLMs, released with the PropXplain dataset.
Code for the reinforcement-learning (GRPO) and chain-of-thought approach to explainable meme detection.
Contributor to QCRI's framework for benchmarking LLMs across many tasks and languages.
The fastest way to reach me is by email. I'm open to research collaborations, especially around multimodal NLP and Arabic language technology.
📍 Doha, Qatar · Qatar Computing Research Institute (QCRI)