{"quality_controlled":"1","acknowledgement":"This research was funded in parts by the FORTE program of the Austrian Research Promotion Agency (FFG) and the Federal Ministry of Agriculture, Regions and Tourism (BMLRT) as part of the AMMONIS project (grant no. 879705). The research was also supported by the Scientific Service Units (SSU) of IST Austria through resources provided by Scientific Computing (SciComp).","author":[{"last_name":"Lampert","full_name":"Lampert, Jasmin","first_name":"Jasmin"},{"last_name":"Lampert","full_name":"Lampert, Christoph","first_name":"Christoph","orcid":"0000-0002-4561-241X","id":"40C20FD2-F248-11E8-B48F-1D18A9856A87"}],"_id":"10752","external_id":{"isi":["000800559505036"]},"publication_identifier":{"isbn":["9781665439022"]},"date_created":"2022-02-10T14:08:23Z","month":"01","user_id":"317138e5-6ab7-11ef-aa6d-ffef3953e345","title":"Overcoming rare-language discrimination in multi-lingual sentiment analysis","article_processing_charge":"No","page":"5185-5192","doi":"10.1109/bigdata52589.2021.9672003","department":[{"_id":"ChLa"}],"publication_status":"published","year":"2022","date_published":"2022-01-13T00:00:00Z","isi":1,"oa_version":"None","citation":{"mla":"Lampert, Jasmin, and Christoph Lampert. “Overcoming Rare-Language Discrimination in Multi-Lingual Sentiment Analysis.” 2021 IEEE International Conference on Big Data, IEEE, 2022, pp. 5185–92, doi:10.1109/bigdata52589.2021.9672003.","short":"J. Lampert, C. Lampert, in:, 2021 IEEE International Conference on Big Data, IEEE, 2022, pp. 5185–5192.","ista":"Lampert J, Lampert C. 2022. Overcoming rare-language discrimination in multi-lingual sentiment analysis. 2021 IEEE International Conference on Big Data. Big Data: International Conference on Big Data, 5185–5192.","ieee":"J. Lampert and C. Lampert, “Overcoming rare-language discrimination in multi-lingual sentiment analysis,” in 2021 IEEE International Conference on Big Data, Orlando, FL, United States, 2022, pp. 5185–5192.","ama":"Lampert J, Lampert C. Overcoming rare-language discrimination in multi-lingual sentiment analysis. In: 2021 IEEE International Conference on Big Data. IEEE; 2022:5185-5192. doi:10.1109/bigdata52589.2021.9672003","chicago":"Lampert, Jasmin, and Christoph Lampert. “Overcoming Rare-Language Discrimination in Multi-Lingual Sentiment Analysis.” In 2021 IEEE International Conference on Big Data, 5185–92. IEEE, 2022. https://doi.org/10.1109/bigdata52589.2021.9672003.","apa":"Lampert, J., & Lampert, C. (2022). Overcoming rare-language discrimination in multi-lingual sentiment analysis. In 2021 IEEE International Conference on Big Data (pp. 5185–5192). Orlando, FL, United States: IEEE. https://doi.org/10.1109/bigdata52589.2021.9672003"},"abstract":[{"text":"The digitalization of almost all aspects of our everyday lives has led to unprecedented amounts of data being freely available on the Internet. In particular social media platforms provide rich sources of user-generated data, though typically in unstructured form, and with high diversity, such as written in many different languages. Automatically identifying meaningful information in such big data resources and extracting it efficiently is one of the ongoing challenges of our time. A common step for this is sentiment analysis, which forms the foundation for tasks such as opinion mining or trend prediction. Unfortunately, publicly available tools for this task are almost exclusively available for English-language texts. Consequently, a large fraction of the Internet users, who do not communicate in English, are ignored in automatized studies, a phenomenon called rare-language discrimination.In this work we propose a technique to overcome this problem by a truly multi-lingual model, which can be trained automatically without linguistic knowledge or even the ability to read the many target languages. The main step is to combine self-annotation, specifically the use of emoticons as a proxy for labels, with multi-lingual sentence representations.To evaluate our method we curated several large datasets from data obtained via the free Twitter streaming API. The results show that our proposed multi-lingual training is able to achieve sentiment predictions at the same quality level for rare languages as for frequent ones, and in particular clearly better than what mono-lingual training achieves on the same data. ","lang":"eng"}],"type":"conference","language":[{"iso":"eng"}],"scopus_import":"1","conference":{"name":"Big Data: International Conference on Big Data","end_date":"2021-12-18","start_date":"2021-12-15","location":"Orlando, FL, United States"},"acknowledged_ssus":[{"_id":"ScienComp"}],"publication":"2021 IEEE International Conference on Big Data","date_updated":"2024-10-21T06:01:53Z","corr_author":"1","publisher":"IEEE","status":"public","day":"13"}