Downloads · 30 days
0
Amitkumar001/Law_Slm
Law_Slm is a text generation model from Amitkumar001. Use it when you need the model to write or continue text. It is set up for slm. The card lists the license as mit.
An industrial-grade, decoder-only Small Language Model (SLM) ecosystem built completely from scratch in pure Python and PyTorch primitive operations.
Downloads · 30 days
0
Access
Public
Updated Aug 8, 2026
Repo size
57.3 MB
Likes
0
Public
Click a slice to open those files.
.node57.3 MB · 36%
From the Hugging Face model README
An industrial-grade, decoder-only Small Language Model (SLM) ecosystem built completely from scratch in pure Python and PyTorch primitive operations.
This repository contains NO reliance on Hugging Face (transformers, tokenizers, datasets, accelerate), pre-built GPT/Llama models, or third-party attention implementations:
<pad>, <unk>, <s>, </s>, <mask >), serialization, and encoding/decoding completely from scratch.CustomAdamW with decoupled weight decay and CustomLion.| Configuration Profile | Vocabulary Size | Hidden Dim ($d_{\text{model}}$) | Attention Heads | Transformer Layers | Feed-Forward Dim ($d_{\text{ff}}$) | Max Sequence Length | Parameter Count |
|---|---|---|---|---|---|---|---|
Nano SLM (configs/nano_config.yaml) | 2,000 | 128 | 4 | 2 | 512 | 256 | ~0.5M |
Standard SLM (configs/default_config.yaml) | 32,000 | 512 | 8 | 8 | 2,048 | 1,024 | ~28M |
| Medium SLM (Custom Config) | 32,000 | 1,024 | 16 | 16 | 4,096 | 2,048 | ~140M |
Given key/query projection vectors $x \in \mathbb{R}^{d_{\text{head}}}$, RoPE rotates token pair slices using complex position angle frequencies: $$\mathbf{R}{\Theta, m}^{d} \mathbf{x}{m} = \begin{pmatrix} x_1 \cos m\theta_1 - x_2 \sin m\theta_1 \ x_1 \sin m\theta_1 + x_2 \cos m\theta_1 \ \vdots \end{pmatrix}$$ where $\theta_i = 10000^{-2(i-1)/d}$.
Instead of centering by mean, RMSNorm scales vector activations directly by their root-mean-square amplitude: $$\text{RMSNorm}(\mathbf{x}) = \frac{\mathbf{x}}{\sqrt{\frac{1}{d}\sum_{i=1}^d x_i^2 + \epsilon}} \odot \gamma$$
SwiGLU applies Swish gating across parallel linear weight projections ($W_g, W_v, W_o$): $$\text{SwiGLU}(\mathbf{x}) = \left( \left( \mathbf{x} W_g \right) \cdot \sigma\left( \mathbf{x} W_g \right) \right) \odot \left( \mathbf{x} W_v \right) W_o$$
Autoregressive attention prevents future token leakage using lower-triangular causal masks: $$\text{Attention}(Q, K, V) = \text{softmax}\left( \frac{Q K^T}{\sqrt{d_k}} + M_{\text{causal}} \right) V$$ where $M_{\text{causal}}[i, j] = 0$ for $i \ge j$ and $-\infty$ for $i < j$.
To reduce parameters and improve generalization, token embedding weights $W_{\text{emb}} \in \mathbb{R}^{V \times d}$ are tied directly to the final Causal LM output head projection layer: $$W_{\text{lm_head}} = W_{\text{emb}}^T$$
graph TD
User[User / Application Request] --> Ingestion[1. Dataset Ingestion & Preprocessing]
Ingestion --> Tokenizer[2. Custom BPE Subword Tokenizer]
Tokenizer --> DatasetLoader[3. Causal LM Dataset & DataLoader]
DatasetLoader --> Transformer[4. SLM Decoder Transformer Core]
subgraph Transformer Block Architecture
Transformer --> Embed[Token Embeddings + RoPE]
Embed --> RMSNorm1[RMSNorm Layer 1]
RMSNorm1 --> MHA[Multi-Head Causal Attention]
MHA --> Res1[Residual Connection 1]
Res1 --> RMSNorm2[RMSNorm Layer 2]
RMSNorm2 --> SwiGLU[SwiGLU FFN Layer]
SwiGLU --> Res2[Residual Connection 2]
end
Res2 --> LMHead[5. Tied LM Output Projection Head]
LMHead --> LossOrSampler{6. Execution Mode}
LossOrSampler -- Training --> Trainer[PyTorch Autograd / AMP Trainer]
LossOrSampler -- Generation --> SamplerEngine[Top-K / Top-P / Temp Sampling Engine]
Trainer --> Checkpoints[(Save Model Checkpoints)]
SamplerEngine --> REPL[Interactive Chat REPL Interface]
SamplerEngine --> REST[FastAPI REST Web Server]
Raw Documents (.txt / .json / .csv / .pdf)
│
▼
[TextCleaner] normalize_unicode() ➔ strip_tags() ➔ deduplicate()
│
▼
[BPETokenizer] train_on_texts() ➔ learn merge rules ➔ save vocab.json
│
▼
[CausalLMDataset] Token windowing ➔ Shift Targets (Input: x_0..x_T-1, Target: x_1..x_T)
│
▼
[DataLoader] Dynamic Batching & Shuffling ➔ GPU Tensor Tensors
================================================================================
SLM INTERACTIVE CHAT REPL (Pure Python Decoder Model)
================================================================================
Model Config: d_model=512 | n_heads=8 | n_layers=8 | vocab=32000 | device=cuda
System Mode : Autoregressive Causal Sampling (Temp=0.7, Top-K=40, Top-P=0.9)
================================================================================
Type your prompt/question below. Type 'exit', 'quit', or 'q' to terminate.
--------------------------------------------------------------------------------
User > What is the primary function of a Small Language Model?
SLM > A Small Language Model (SLM) is a compact transformer network designed
for high-efficiency local inference and rapid domain pre-training
without needing massive datacenter compute.
--------------------------------------------------------------------------------
User > Explain Rotary Position Embeddings (RoPE).
SLM > RoPE rotates query and key vectors in 2D vector slices by position-
dependent angles, enabling natural relative position decay.
--------------------------------------------------------------------------------
User > q
[Chat Session Terminated. Model memory freed.]
================================================================================
┌──────────────────────────────────────────────────────────────────────────────┐
│ FASTAPI REST API DASHBOARD — Small Language Model Service (v1.0.0) │
│ Base URL: http://localhost:8000 │
├──────────────────────────────────────────────────────────────────────────────┤
│ ENDPOINTS: │
│ [GET] /health ➜ Server & Hardware Status Check │
│ [GET] /info ➜ Model Architecture & Parameter Metadata │
│ [POST] /generate ➜ Autoregressive Text Generation │
│ [POST] /tokenizer/encode ➜ Tokenize Raw Text into Integer Sequence │
│ [POST] /tokenizer/decode ➜ Decode Integer IDs back into String │
├──────────────────────────────────────────────────────────────────────────────┤
│ SAMPLE POST /generate REQUEST PAYLOAD: │
│ { │
│ "prompt": "User: What is Causal Attention?\nSLM:", │
│ "max_new_tokens": 128, │
│ "temperature": 0.7, │
│ "top_k": 40, │
│ "top_p": 0.9 │
│ } │
└──────────────────────────────────────────────────────────────────────────────┘
Clone the repository and install in editable mode:
cd lawslm
pip install -e .
Dependencies required: torch, numpy, pyyaml, fastapi, uvicorn, pydantic, pytest.
You can download open-source training text datasets directly using the included downloader script:
# Download WikiText-2 (Wikipedia articles)
python scripts/download_sample_dataset.py --name wikitext2
# Download TinyStories (Synthetic clean story corpus)
python scripts/download_sample_dataset.py --name tinystories
# Download TinyShakespeare (Shakespeare corpus)
python scripts/download_sample_dataset.py --name tinyshakespeare
| Dataset Name | Best Use Case | Source Link |
|---|---|---|
| TinyStories | Top Choice for SLMs: Synthetically generated clean stories to teach small parameter models grammar & reasoning fast. | roneneldan/TinyStories |
| WikiText-103 & WikiText-2 | Standard high-quality Wikipedia corpus formatted in clean plain .txt. | Salesforce WikiText |
| Project Gutenberg | Over 70,000 free public domain books (literature, science, history). | gutenberg.org |
| Legal & Court Datasets | Free Law Project / CourtListener public court opinions and statutes. | courtlistener.com |
Preprocess raw text files (.txt, .json, .jsonl, .csv, .md) to normalize Unicode, strip HTML tags, split sentences, and deduplicate lines:
python scripts/preprocess_data.py --input data/wikitext2_train.txt --output data/cleaned_dataset.txt
To start training with nano_config.yaml or default_config.yaml:
python -m slm.cli.main train --config configs/nano_config.yaml --dataset data/cleaned_dataset.txt
Execute the training script directly:
python scripts/train_run.py
checkpoints/ (weights, optimizer, scheduler, tokenizer, and random seeds).Launch a real-time interactive chat session:
python scripts/chat_run.py checkpoints/best_model.pt
Run via the CLI chat command:
python -m slm.cli.main chat --checkpoint checkpoints/best_model.pt
To generate text continuation for a single prompt:
python -m slm.cli.main generate \
--prompt "The primary goal of language modeling is" \
--checkpoint checkpoints/best_model.pt \
--max_tokens 100 \
--temperature 0.8 \
--top_k 40 \
--top_p 0.9
Launch the FastAPI production web server:
uvicorn slm.api.app:app --host 0.0.0.0 --port 8000 --reload
Interactive OpenAPI Documentation is available at: http://localhost:8000/docs
cURLcurl -X POST "http://localhost:8000/generate" \
-H "Content-Type: application/json" \
-d '{
"prompt": "User: What is deep learning?\nSLM:",
"max_new_tokens": 64,
"temperature": 0.7,
"top_k": 40,
"top_p": 0.9
}'
requestsimport requests
response = requests.post(
"http://localhost:8000/generate",
json={
"prompt": "User: Explain self-attention in simple terms.\nSLM:",
"max_new_tokens": 100,
"temperature": 0.7
}
)
data = response.json()
print("Model Reply:", data["generated_text"])
Benchmark forward pass latency, token throughput (tokens/sec), and VRAM memory stats:
python scripts/benchmark_run.py
Or via CLI:
python -m slm.cli.main benchmark --batch_size 4 --seq_len 256 --device auto
Execute PyTest to verify tokenization, causal attention masking, RMSNorm, custom AdamW optimizer, and model forward pass:
pytest -v tests/
README.md and MODEL_CARD.md to resolve the YAML Metadata Warning: empty or missing yaml metadata in repo card..gitignore and uploader rules excluding node_modules/, web/node_modules/, dist/, and binary build artifacts to eliminate security scanner flags.Push your model weights, configs, tokenizer, and repo card automatically using push_to_hf.py:
python scripts/push_to_hf.py --repo_id Amit123103/Law_model_slm --token hf_YOUR_HF_ACCESS_TOKEN
$env:HF_TOKEN="hf_YOUR_HF_ACCESS_TOKEN"
python scripts/push_to_hf.py --repo_id Amit123103/Law_model_slm
(Note: Replace hf_YOUR_HF_ACCESS_TOKEN with your personal Hugging Face access token from https://huggingface.co/settings/tokens with Write permissions).
c:/Users/amita/myprojects/lawslm/
├── configs/ # YAML Model & Train configurations
│ ├── default_config.yaml
│ └── nano_config.yaml
├── docker/ # Dockerfile & Docker Compose
│ ├── Dockerfile
│ └── docker-compose.yml
├── docs/ # Comprehensive user guides
│ ├── API.md
│ ├── ARCHITECTURE.md
│ ├── INFERENCE.md
│ └── TRAINING.md
├── scripts/ # Executable training, benchmark, chat, and preprocessing scripts
│ ├── benchmark_run.py
│ ├── chat_run.py
│ ├── download_sample_dataset.py
│ ├── preprocess_data.py
│ └── train_run.py
├── slm/ # Core package source
│ ├── api/ # FastAPI web server (/generate, /info, /health)
│ ├── attention/ # Multi-Head Causal Attention with RoPE
│ ├── checkpoint/ # State serialization & rotation manager
│ ├── cli/ # Command line interface (train, generate, chat, benchmark)
│ ├── config/ # Model and Training configuration schemas
│ ├── dataset/ # Multi-format dataset ingestion & CausalLMDataset
│ ├── embeddings/ # Token Embeddings & RoPE/Sinusoidal/Learned position encodings
│ ├── evaluation/ # Metrics (BLEU, ROUGE, PPL) & Hardware benchmark suite
│ ├── feedforward/ # SwiGLU & GELU Feed-Forward Networks
│ ├── model/ # SLMForCausalLM Decoder Transformer
│ ├── normalization/ # RMSNorm & CustomLayerNorm
│ ├── optimizer/ # Custom AdamW & Lion optimizers from scratch
│ ├── sampling/ # Text generator & sampling strategies
│ ├── scheduler/ # Cosine with Warmup LR schedulers
│ ├── tokenizer/ # Custom BPE Tokenizer built from scratch
│ ├── transformer/ # Pre-Norm Causal Transformer Block
│ ├── training/ # Industrial Trainer with AMP and grad accumulation
│ └── utils/ # Logging, device resolution, memory stats
├── tests/ # Comprehensive PyTest suite
├── pyproject.toml
└── README.md
MIT License. Built for industrial AI research and custom Small Language Model development completely from total zero.