Downloads · 30 days
0
CCRss/topic_modeling_top2vec_scientific-texts
topic_modeling_top2vec_scientific-texts is a machine learning model from CCRss. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
Downloads · 30 days
0
Access
Public
Updated Apr 8, 2024
Repo size
1.9 GB
Likes
0
Public
Click a slice to open those files.
Other1.9 GB · 100%
From the Hugging Face model README

This repository hosts the top2vec_scientific_texts model, a specialized Top2Vec model trained on scientific texts for topic modeling and semantic search.
The top2vec_scientific_texts model is built for analyzing scientific literature. It leverages the Universal Sentence Encoder for embedding texts and uses Top2Vec for topic modeling.
To use the model, you need to install the following dependencies:
pip install top2vec
pip install top2vec[sentence_encoders]
pip install tensorflow==2.8.0
pip install tensorflow-probability==0.16.0
The entire process of model training, dataset creation, and visualization is documented in the main.ipynb Jupyter notebook. To explore the code and replicate the results:
main.ipynb notebook in Jupyter Lab or Jupyter Notebook.For more details, please refer to the main.ipynb notebook in this repository.
Here's an example of how to use the model for topic modeling:
from top2vec import Top2Vec
# Load your documents
docs = ["Document 1 text", "Document 2 text", ...]
# Initialize the Top2Vec model
model = Top2Vec(
documents=docs,
speed='learn',
workers=80,
embedding_model='universal-sentence-encoder',
umap_args={'n_neighbors': 15, 'n_components': 5, 'metric': 'cosine', 'min_dist': 0.0, 'random_state': 42},
hdbscan_args={'min_cluster_size': 15, 'metric': 'euclidean', 'cluster_selection_method': 'eom'}
)
model.save('top2vec_scientific_texts_model')
The model was trained on a dataset of scientific abstracts sourced from arXiv. The dataset covers a range of topics within the field of computer science from 2010 to 2024.
You can access the dataset arxiv_papers_cs.
The top2vec_scientific_texts model can be used for various purposes, including:
Here are some examples of the model's output for the thematic group "UAV in Disasters and Emergency":

This graph shows the trend of interest in the use of UAVs in disaster and emergency situations over time.
Analysis for Thematic Group: Disasters & Emergency
| Year | Number of Publications | Growth Acceleration | Change in Number of Publications | Relative Growth |
|---|---|---|---|---|
| 2010 | 19 | 0 | 0 | 0.0% |
| 2011 | 15 | -4 | -4 | -21.05% |
| 2012 | 28 | 17 | 13 | 86.67% |
| 2013 | 38 | -3 | 10 | 35.71% |
| 2014 | 28 | -20 | -10 | -26.32% |
| 2015 | 47 | 29 | 19 | 67.86% |
| 2016 | 63 | -3 | 16 | 34.04% |
| 2017 | 94 | 15 | 31 | 49.21% |
| 2018 | 173 | 48 | 79 | 84.04% |
| 2019 | 266 | 14 | 93 | 53.76% |
| 2020 | 337 | -22 | 71 | 26.69% |
| 2021 | 380 | -28 | 43 | 12.76% |
| 2022 | 453 | 30 | 73 | 19.21% |
| 2023 | 509 | -17 | 56 | 12.36% |
We welcome contributions to the top2vec_scientific_texts model. If you have suggestions, improvements, or encounter any issues, please feel free to open an issue or submit a pull request.
This project is licensed under the MIT License