Downloads · 30 days
328
2% of all-time downloads
k4tel/geo-bert-multilingual
geo-bert-multilingual is a machine learning model from k4tel. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers.
This model predicts the geolocation of short texts (less than 500 words) in a form of two-dimensional distributions also referenced as the Gaussian Mixture Model (GMM).
Downloads · 30 days
328
2% of all-time downloads
All-time downloads
15.4K
Public
Repo size
1.4 GB
Likes
7
Public
Click a slice to open those files.
.bin712 MB · 100%
From the Hugging Face model README
This model predicts the geolocation of short texts (less than 500 words) in a form of two-dimensional distributions also referenced as the Gaussian Mixture Model (GMM).
Number of predicted points: 5 Custom transformers pipeline and result visualization: https://github.com/K4TEL/geo-twitter/tree/predict
This project was aimed to solve the tweet/user geolocation prediction task and provide a flexible methodology for the geotagging of textual big data. The suggested approach implements BERT-based neural networks for NLP to estimate the location in a form of two-dimensional GMMs (longitude, latitude, weight, covariance). The base model has been finetuned on a Twitter dataset containing text content and metadata context of the tweets.
Geo-tagging of Big data
Per-tweet geolocation prediction
Per-tweet geolocation prediction without "user" metadata is expected to show lower accuracy of predictions.
Risk for unethical use on the basis of data that is not publicly available.
The limitation of text length is dictated by the BERT-based model's capacity of 500 tokens (words).
Use the code below to get started with the model:
https://github.com/K4TEL/geo-twitter/tree/predict
A short startup guide is given in the repository branch description.
The Twitter dataset contained tweets with their text content, metadata ("user" and "place") context, and geolocation coordinates.
Information about the model training on the user-defined data could be found in the GitHub repository: https://github.com/K4TEL/geo-twitter
All performance metrics and results are demonstrated in the Results section of the article pre-print: https://arxiv.org/pdf/2303.07865.pdf
Worldwide dataset of tweets with TEXT-ONLY and NON-GEO features
Spatial metrics: mean and median Simple Accuracy Error (SAE), Acc@161 Probabilistic metrics: mean and median Cumulative Accuracy Error (CAE), mean and median Prediction Area Region (PRA) for 95% density area, Coverage of PRA
Tweet geolocation prediction task
User home geolocation prediction task
Implemented wrapper layer of liner regression with a custom number of output variables that operates with classification token generated by the base BERT model.
NVIDIA GeForce GTX 1080 Ti
Python IDE