MONETA: Multimodal Industry Classification through Geographic Information with Multi Agent Systems

Published in ACL 2026, 2026

MONETA introduces the first multimodal industry classification benchmark that combines text sources (website content, Wikipedia, Wikidata) with geospatial sources (OpenStreetMap and satellite imagery) to classify businesses by economic activity. The dataset comprises 1,000 European businesses labeled according to EU economic activity guidelines.

We evaluate multimodal large language models and multi-agent system designs as baselines, achieving performance ranging from 62.10% to 74.10%, with improvements of up to 22.80% through enhanced design techniques such as multi-turn interactions, enriched context, and explanation generation.

The benchmark and code are publicly available to support future research on multimodal and geographically grounded industry classification.

Recommended citation: Yüksel, A., Thiem, G., Walter, S., Felka, P., Alves Werb, G., & Habernal, I. (2026). "MONETA: Multimodal Industry Classification through Geographic Information with Multi Agent Systems." *Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)*.
Download Paper