Downloads · 30 days
0
ManuelCadena/mananeras-word2vec
mananeras-word2vec is a machine learning model from ManuelCadena. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Spanish skip-gram Word2Vec embeddings trained on the presidential interventions of Mexican morning press conferences ("mañaneras") from December 2018 to August 2025.
Downloads · 30 days
0
Access
Public
Updated Sep 3, 2026
Repo size
27.8 MB
Likes
1
Public
Click a slice to open those files.
.txt27.8 MB · 100%
From the Hugging Face model README
Spanish skip-gram Word2Vec embeddings trained on the presidential interventions of Mexican morning press conferences ("mañaneras") from December 2018 to August 2025.
conferencias_matutinas_amlo and conferencias_matutinas_sheinbaum repositories, preprocessed with spaCy es_core_news_md.This is a research artifact for the CSCI E-89B final project “What Mexico Hears vs. What Mexico Believes” (Harvard Extension, Fall 2026). It is intended for:
It is not a general-purpose Spanish embedding model and should not be used as a drop-in replacement for larger pre-trained embeddings.
corpus_*.csv.gz from final_project/data/amlo/ and final_project/data/sheinbaum/.es_presidente == 1.es_core_news_md.gensim.models.Word2Vec skip-gram model.mananeras-word2vec.gensim) and the plain text vectors (mananeras-word2vec.txt).The exact training script is in final_project/scripts/train_word2vec.py.
@misc{cadena2026mananeras,
title={What Mexico Hears vs. What Mexico Believes: Structural Topic Modeling and Transformer-Based Sentiment of Mexican Presidential Morning Conferences},
author={Manuel Cadena Ortiz de Montellano},
year={2026},
note={CSCI E-89B Final Project, Harvard Extension School}
}
The source transcripts are public Mexican government records compiled by NOSTRODATA. This derived model is released for academic use; cite NOSTRODATA for the raw corpus.