Downloads · 30 days
0
q494081672/diversity_selection
diversity_selection is a machine learning model from q494081672. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This folder contains the helper scripts for selecting high-quality and diverse Tulu samples with:
Downloads · 30 days
0
Access
Public
Updated Aug 4, 2026
Repo size
7.5 GB
Likes
0
Public
Click a slice to open those files.
.pkl6.5 GB · 82%
From the Hugging Face model README
This folder contains the helper scripts for selecting high-quality and diverse Tulu samples with:
final_score_continuous as the quality score.BAAI/bge-large-en-v1.5 as the semantic embedding model.prepare_deita_diversity_inputs.py: builds a top-N candidate JSON from the
original parquet and score CSV.embed_with_bge.py: creates a Deita-compatible embedding pickle with BGE.run_deita_diversity_selection.sh: runs BGE embedding and Deita filtering.top_50k_by_final_score.json: default quality-ranked candidate pool.top_50k_by_final_score_for_embed.json: Deita-style conversation copy kept for
compatibility/reference.From the project root:
conda activate tokenclean
bash diversity_selection/run_deita_diversity_selection.sh
Common options:
GPU=0 THRESHOLD=0.85 DATA_SIZE=10000 BGE_BATCH_SIZE=128 \
bash diversity_selection/run_deita_diversity_selection.sh
Regenerate the top-50k candidate pool:
python diversity_selection/prepare_deita_diversity_inputs.py
Force embedding regeneration:
FORCE_REEMBED=1 bash diversity_selection/run_deita_diversity_selection.sh
The default output is:
diversity_selection/top_10k_by_final_score_diverse.json
The default BGE embedding cache is:
diversity_selection/top_50k_by_final_score_bge_embeddings.pkl