Downloads · 30 days
32
100% of all-time downloads
MARS-Retokenization/olmo2-7b-instruct-inv-reference-mixed-klrev
olmo2-7b-instruct-inv-reference-mixed-klrev is a machine learning model from MARS-Retokenization. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Research checkpoint from a study of tokenization (reader) invariance and its effect on robustness to adversarial re-tokenization. Fine-tuned from allenai/OLMo-2-1124-7B-Instruct.
Downloads · 30 days
32
100% of all-time downloads
All-time downloads
32
Public
Parameters
7.3B
14.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors14.6 GB · 100%
From the Hugging Face model README
Research checkpoint from a study of tokenization (reader) invariance and its
effect on robustness to adversarial re-tokenization. Fine-tuned from
allenai/OLMo-2-1124-7B-Instruct.
These are research artifacts, not products. Numbers below are measured on a 200-prompt AdvBench holdout that was excluded from training by construction.
| metric | value |
|---|---|
| — | not yet measured |
AdvTok ASR is attack success rate under adversarial tokenization
(Geh et al., arXiv:2503.02174), Llama-Guard-3-8B judged, greedy decoding.
XSTest over-refusal is the refusal rate on safe prompts — the cost side.
| field | value |
|---|---|
mode | reference |
objective | invariance |
kl_direction | reverse |
ce_weighting | uniform |
ema_beta | None |
harmful_mix | mixed |
harmful_fraction | 0.286 |
num_encodings | 8 |
cvar_quantile | 0.25 |
max_steps | 700 |
learning_rate | 1e-05 |
grad_accum | 8 |
seed | 42 |
reference_model | allenai/OLMo-2-1124-7B-Instruct |
max_new_tokens | 128 |
prefix_tokens | 8 |
| group | relative L2 |
|---|---|
| attn | 0.01202 |
| embed_tokens | 0.00096 |
| lm_head | 0.00605 |
| mlp | 0.01320 |
| norm | 0.00101 |