Downloads ยท 30 days
0
Shuche/cfg3f_dataset
cfg3f_dataset is a machine learning model from Shuche. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This repository contains three versions of the CFG3f synthetic dataset, generated using Context-Free Grammar rules for language modeling research.
Downloads ยท 30 days
0
Access
Public
Updated Aug 31, 2025
Repo size
815 MB
Likes
0
Public
Click a slice to open those files.
.bin815 MB ยท 100%
From the Hugging Face model README
This repository contains three versions of the CFG3f synthetic dataset, generated using Context-Free Grammar rules for language modeling research.
| Version | Description | Train Size | Val Size | Total Size | Avg Tokens/Sample |
|---|---|---|---|---|---|
| Balanced | Equal probability for all grammar rules | 800K samples | 80K samples | 437MB | ~261 tokens |
| Moderate | Moderate imbalance (10:1 weight ratio) | 800K samples | 80K samples | 192MB | ~114 tokens |
| Extreme | Extreme imbalance (100:1 weight ratio) | 800K samples | 80K samples | 150MB | ~89 tokens |
cfg3f_datasets/
โโโ balanced/ # Balanced generation (all rules equal probability)
โ โโโ train/ # 800K training samples
โ โโโ val/ # 80K validation samples
โโโ moderate/ # Moderate imbalance (10:1 ratio)
โ โโโ train/ # 800K training samples
โ โโโ val/ # 80K validation samples
โโโ extreme/ # Extreme imbalance (100:1 ratio)
โโโ train/ # 800K training samples
โโโ val/ # 80K validation samples
The dataset is generated using a hierarchical context-free grammar with 4 levels:
22 -> 20 21 | 20 19 21 | 21 19 19 | 20 20
21 -> 18 17 | 17 16 | 16 17 18 | 16 18
20 -> 16 17 | 17 16 18
19 -> 18 16 18 | 17 18 | 18 18
18 -> 14 15 13 | 15 13 13 | 13 15
17 -> 14 15 | 15 14 | 15 14 13
16 -> 13 15 13 | 15 15 | 14 13 | 14 14
15 -> 10 11 11 | 11 11 10 | 10 10 | 12 12 11
14 -> 12 10 12 | 10 12 12 | 12 11 | 10 12
13 -> 11 12 | 10 12 11 | 12 11 12
12 -> 7 9 7 | 9 8 | 8 8 9
11 -> 8 8 | 9 7 | 9 7 7
10 -> 8 9 9 | 9 7 9 | 7 9 9
9 -> '3' '3' | '1' '1' | '1' '2' '1'
8 -> '3' '3' '1' | '3' '1' '1' | '1' '2'
7 -> '3' '2' | '3' '2' '2' | '3' '1' '2' | '2' '2' '1'
MODERATE_WEIGHTS = {
'9': [1, 10, 1], # Prefer "1 1" over others
'8': [1, 1, 10], # Prefer "1 2" over others
'7': [10, 1, 1, 1], # Prefer "3 2" over others
# ... similar patterns for all rules
}
EXTREME_WEIGHTS = {
'9': [1, 100, 1], # Heavily prefer "1 1"
'8': [1, 1, 100], # Heavily prefer "1 2"
'7': [100, 1, 1, 1], # Heavily prefer "3 2"
# ... similar patterns for all rules
}
'1' โ 0'2' โ 1'3' โ 2'[BOS]' โ 3'[EOS]' โ 4| Aspect | Balanced | Moderate | Extreme |
|---|---|---|---|
| Rule Selection | Uniform | 10:1 bias | 100:1 bias |
| Min Branch Prob | 25-33% | 8.33% | 0.97% |
| Sentence Complexity | High | Medium | Low |
| Rare Patterns | Common | Present | Very Rare |
| Research Focus | Baseline | Moderate imbalance | Long-tail distribution |
import numpy as np
def read_bin_file(filename):
with open(filename, 'rb') as f:
# Read header
header = np.frombuffer(f.read(256 * 4), dtype=np.int32)
magic, version, num_tokens = header[0], header[1], header[2]
# Read tokens
tokens = np.frombuffer(f.read(), dtype=np.uint16)
return tokens
# Load training data
train_tokens = read_bin_file("balanced/train/cfg3f_train_000000.bin")
print(f"Loaded {len(train_tokens)} tokens")
If you use this dataset in your research, please cite:
@dataset{cfg3f_dataset_2024,
title={CFG3f: Context-Free Grammar Dataset with Controlled Imbalance},
author={Shuche Wang},
year={2024},
url={https://huggingface.co/datasets/Shuche/cfg3f_dataset}
}
This dataset is part of research on:
MIT License - Feel free to use for research and educational purposes.