Enhanced geocoding precision for location inference of tweet text using spaCy, Nominatim and Google Maps. A comparative analysis of the influence of data selection-Reference-Cited by-同舟云学术

Enhanced geocoding precision for location inference of tweet text using spaCy, Nominatim and Google Maps. A comparative analysis of the influence of data selection

Published:2023-03-15 Issue:3 Volume:18 Page:e0282942
ISSN:1932-6203
Container-title:PLOS ONE
language:en
Short-container-title:PLoS ONE

Author:

Serere Helen Ngonidzashe^ORCID,Resch Bernd,Havas Clemens Rudolf

Abstract

Twitter location inference methods are developed with the purpose of increasing the percentage of geotagged tweets by inferring locations on a non-geotagged dataset. For validation of proposed approaches, these location inference methods are developed on a fully geotagged dataset on which the attached Global Navigation Satellite System coordinates are used as ground truth data. Whilst a substantial number of location inference methods have been developed to date, questions arise pertaining the generalizability of the developed location inference models on a non-geotagged dataset. This paper proposes a high precision location inference method for inferring tweets’ point of origin based on location mentions within the tweet text. We investigate the influence of data selection by comparing the model performance on two datasets. For the first dataset, we use a proportionate sample of tweet sources of a geotagged dataset. For the second dataset, we use a modelled distribution of tweet sources following a non-geotagged dataset. Our results showed that the distribution of tweet sources influences the performance of location inference models. Using the first dataset we outweighed state-of-the-art location extraction models by inferring 61.9%, 86.1% and 92.1% of the extracted locations within 1 km, 10 km and 50 km radius values, respectively. However, using the second dataset our precision values dropped to 45.3%, 73.1% and 81.0% for the same radius values.

Funder

Austria Research Promotion Agency

Publisher

Public Library of Science (PLoS)

Subject

Multidisciplinary

Reference47 articles.

1. Citizen-centric urban planning through extracting emotion information from twitter in an interdisciplinary space-time-linguistics algorithm;B Resch;Urban Plan,2016

2. Event relatedness assessment of Twitter messages for emergency response.;F Laylavi;Inf Process Manag,2017

3. CIME: Context-aware geolocation of emergency-related posts.;G Scalia;Geoinformatica,2021

4. Combining machine-learning topic models and spatiotemporal analysis of social media data for disaster footprint and damage assessment;B Resch;Cartogr Geogr Inf Sci,2018

5. Urchs S, Wendlinger L, Mitrović J, Granitzer M. MMoveT15: A Twitter Dataset for Extracting and Analysing Migration-Movement Data of the European Migration Crisis 2015. Proc. - 2019 IEEE 28th Int. Conf. Enabling Technol. Infrastruct. Collab. Enterp. WETICE 2019, 2019, p. 146–9. https://doi.org/10.1109/WETICE.2019.00039.

Cited by 10 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Leveraging Crowdsourcing for Mapping Mobility Restrictions in Data-Limited Regions;Smart Cities;2024-09-07

2. DLRGeoTweet: A comprehensive social media geocoding corpus featuring fine-grained places;Information Processing & Management;2024-07

3. Understanding the impact of geotagging on location inference models for accurate generalization to non-geotagged datasets;Geomatica;2024-07

4. Scientific production in sexual and reproductive health and rights research according to gender and affiliation: An analysis of publications from 1972 to 2021;PLOS ONE;2024-06-26

5. Addressing bias in preterm birth research: The role of advanced imputation techniques for missing race and ethnicity in perinatal health data;Annals of Epidemiology;2024-06