Downloads · 30 days
18
51% of all-time downloads
ash3241/Amplitude-based-tokenizer
Amplitude-based-tokenizer is a machine learning model from ash3241. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
This repository contains the trained weights and vocabulary for a proof-of-concept toy model utilizing the novel Continuous Amplitude Tokenization (CAT) architecture.
Downloads · 30 days
18
51% of all-time downloads
All-time downloads
35
Public
Parameters
4.6M
18.6 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors18.6 MB · 98%
From the Hugging Face model README
This repository contains the trained weights and vocabulary for a proof-of-concept toy model utilizing the novel Continuous Amplitude Tokenization (CAT) architecture.
Traditional LLMs rely on discrete BPE tokenization, forcing them to treat intensity modifiers ("very", "extremely", "slightly") as independent, arbitrary tokens. The CAT architecture splits language into two streams:
1.0, 2.0, 5.0.This toy model processes sentences like "The food was incredibly bad" as a single compressed token mapping: [Concept: "bad", Amplitude: 4.0], completely eliminating the need for combinatorial adjective tokens and physically shortening sequence lengths.
During benchmarking against a standard NanoGPT model on the exact same text, this CAT model achieved:
Amplitude: 10.0 to force intensity scaling during generation).To use these weights, you must use the custom CAT Model code from the official GitHub repository.
git clone https://github.com/Yash943/CAT-Concept-Amplitude-Tokenization-.git
cd CAT-Concept-Amplitude-Tokenization-
pip install -e .
import torch
from cat.model import CATModel, CATConfig
from huggingface_hub import hf_hub_download
# Initialize the architecture
config = CATConfig(
vocab_size=15000,
hidden_size=128,
num_hidden_layers=4,
num_attention_heads=4,
intermediate_size=512,
max_position_embeddings=64
)
model = CATModel(config)
# Load the weights
weights_path = hf_hub_download(repo_id="your-repo-id/here", filename="cat_toy_model.pt")
model.load_state_dict(torch.load(weights_path))
print("CAT Model loaded successfully!")
This is an experimental "toy" model designed strictly to prove the mathematical viability of Fourier-FiLM embeddings and AP-RMSNorm at reducing token fertility. It is not intended for production chat or generation tasks.
For full mathematical breakdowns, empirical verification scripts, and architectural source code, visit the official GitHub repository.