Downloads · 30 days
11
11% of all-time downloads
HamidBekam/MarketBERT
MarketBERT is a machine learning model from HamidBekam. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
This repo builds and trains Market2Vec from trademark data using a sequence-of-products view:
Downloads · 30 days
11
11% of all-time downloads
All-time downloads
102
Public
Parameters
37M
190 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors148 MB · 99%
From the Hugging Face model README
This repo builds and trains Market2Vec from trademark data using a sequence-of-products view:
We support two versions of the MLM objective:
Purpose: learn to forecast the last product’s items from the firm’s earlier product history.
[APP] and the next [APP_END] (or [SEP])ITEM_* tokens inside that last event (mask probability = 1.0)This version is best when your downstream use-case is “given past trademark products, predict items in the most recent/next product”.
Purpose: learn general co-occurrence/semantic structure of items in firm timelines (classic MLM).
p (e.g., 15%)This version is best when you want broad embeddings capturing item relationships and temporal context without specifically focusing on forecasting the last event.
A packed firm sequence looks like:
[CLS] DATE_YYYY_MM [APP] NICE_* ... ITEM_* ... [APP_END] DATE_YYYY_MM [APP] NICE_* ... ITEM_* ... [APP_END] ... DATE_YYYY_MM [APP] NICE_* ... ITEM_* ... [APP_END] <-- last event [SEP]
ITEM_* in the last [APP]..[APP_END] segment (forecasting target)Validation reports:
ITEM_* tokenFor forecasting, Item@K is the main metric because it directly measures how well the model predicts items in the last product basket.
A4_full_fixed_alpha_optionA_h512_h323.64330.59960.66510.6944from transformers import AutoTokenizer, AutoModel
tok = AutoTokenizer.from_pretrained("HamidBekam/MarketBERT")
model = AutoModel.from_pretrained("HamidBekam/MarketBERT")