Downloads · 30 days
0
berkbirkan/aiagent-security-bpe-tokenizer
aiagent-security-bpe-tokenizer is a machine learning model from berkbirkan. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
This repository contains a fast ByteLevel BPE tokenizer trained on a focused AI-agent-security corpus. It is a tokenizer artifact, not a trained language model.
Downloads · 30 days
0
Access
Public
Updated Jul 21, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.json482 KB · 99%
From the Hugging Face model README
This repository contains a fast ByteLevel BPE tokenizer trained on a focused AI-agent-security corpus. It is a tokenizer artifact, not a trained language model.
| Property | Value |
|---|---|
| Updated | 21 July 2026 |
| Training corpus | AI-agent-security paper titles and abstracts |
| Training corpus size | 85,724 characters |
| Algorithm | Hugging Face tokenizers ByteLevel BPE |
| Target vocabulary size | 8,192 (2**13) |
| Actual vocabulary size | 7,027 |
| Special tokens | <unk>, <pad>, <bos>, <eos> |
| Prefix-space behavior | add_prefix_space=False |
The actual vocabulary is smaller than the requested target because training stopped after the corpus could not provide additional BPE merges.
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained(
"berkbirkan/aiagent-security-bpe-tokenizer"
)
text = "Prompt injection attacks target tool-using LLM agents."
token_ids = tokenizer.encode(text, add_special_tokens=False)
print(tokenizer.convert_ids_to_tokens(token_ids))
print(token_ids)
print(tokenizer.decode(token_ids))
The tokenizer was compared with Qwen2.5, Gemma 2, DeepSeek-V3, and SecureBERT. All counts exclude automatically added special tokens. Vocabulary sizes are not matched, so the benchmark measures practical segmentation rather than an algorithm-only comparison.
| Sentence | This tokenizer | Qwen2.5 | Gemma 2 | DeepSeek-V3 | SecureBERT |
|---|---|---|---|---|---|
| Prompt injection attacks target tool-using LLM agents and autonomous coding agents. | 16 | 14 | 15 | 15 | 17 |
| Agent red-teaming reveals vulnerabilities in skill recommendation systems. | 11 | 11 | 12 | 11 | 12 |
| Security guardrails reduce prompt injection risks in AI agent workflows. | 11 | 12 | 12 | 12 | 13 |
| Secret scanner agents extract credentials and access context from unstructured documents. | 13 | 13 | 12 | 12 | 12 |
| Tokenizer | Total tokens | Characters/token |
|---|---|---|
| This tokenizer | 15,103 | 5.676 |
| Qwen2.5 | 16,522 | 5.188 |
| Gemma 2 | 16,521 | 5.189 |
| DeepSeek-V3 | 16,093 | 5.327 |
| SecureBERT | 16,027 | 5.349 |
The tokenizer produces the lowest token count on its own training corpus. This is expected for a domain-trained tokenizer and does not imply the same ranking on general-purpose or out-of-domain text.
| Term | This tokenizer | Qwen2.5 | Gemma 2 | DeepSeek-V3 | SecureBERT |
|---|---|---|---|---|---|
| prompt injection | 4 | 2 | 2 | 3 | 3 |
| LLM agents | 2 | 3 | 3 | 3 | 3 |
| tool-using agents | 4 | 3 | 4 | 4 | 4 |
| autonomous agents | 3 | 3 | 2 | 3 | 2 |
| coding agents | 3 | 2 | 2 | 2 | 2 |
| agent red-teaming | 4 | 4 | 5 | 4 | 5 |
| skill recommendation | 3 | 2 | 2 | 2 | 2 |
| security guardrails | 3 | 3 | 3 | 3 | 4 |
| secret scanner | 4 | 2 | 2 | 2 | 2 |
| vulnerability detection | 4 | 4 | 3 | 4 | 2 |
| agentic workflows | 3 | 3 | 3 | 3 | 3 |
| supply-chain attacks | 6 | 3 | 4 | 4 | 5 |
The tokenizer backend uses tokenizers.models.BPE, a ByteLevel pre-tokenizer,
a ByteLevel decoder, and the complete ByteLevel initial alphabet. It is wrapped
with PreTrainedTokenizerFast for direct use through Transformers and
AutoTokenizer.
Training code, corpus-building steps, reproducible benchmark outputs, and charts are available in the security-bpe-tokenizer GitHub repository.