Downloads · 30 days
3
13% of all-time downloads
chili-lab/Emoji-ByteLM
Emoji-ByteLM is a text generation model from chili-lab. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
This model is a byte-level language model trained with a byte-level Universal transformer (UT) architecture for translating text descriptions to emojis.
Downloads · 30 days
3
13% of all-time downloads
All-time downloads
23
Public
Repo size
1.9 GB
Likes
2
Public
Click a slice to open those files.
.pth1.9 GB · 100%
From the Hugging Face model README
This model is a byte-level language model trained with a byte-level Universal transformer (UT) architecture for translating text descriptions to emojis.
Looped Transformer Architecture:
Model Dimensions:
A looped transformer applies the same transformer layers multiple times in an iterative refinement process. This is particularly effective for translation tasks as it allows the model to:
In this model, 24 layers are applied 8 times with residual connections between loops.
This model is designed to translate text descriptions into appropriate emojis.
Example Usage:
Input: "I love pizza"
Output: "🍕❤️"
The model was trained on the KomeijiForce/Text2Emoji dataset, which contains over 500,000 text-emoji pairs.
This repository contains:
consolidated.pth: PyTorch model weightsparams.json: Complete model and training configurationtrain_state_*.json: Training state information from checkpointTo use this model, you'll need the original BFlowNet/loopedLM codebase to load the architecture:
import torch
import json
# Load model parameters
with open('params.json', 'r') as f:
params = json.load(f)
# Load model weights
checkpoint = torch.load('consolidated.pth', map_location='cpu')
# Initialize model with your BFlowNet loopedLM architecture
# from apps.loopedLM import LoopedTransformer
# model = LoopedTransformer(**params['model'])
# model.load_state_dict(checkpoint)
For best results, use:
This model was trained using the BFlowNet framework with looped transformer architecture.
Dataset: KomeijiForce/Text2Emoji
Apache 2.0