Downloads · 30 days
0
kunjcr2/MedAssistGPT
MedAssistGPT is a machine learning model from kunjcr2. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Experimental medical-domain LLM pretraining project. ⚠️ Research-only. Not for clinical, diagnostic, or production use.
Downloads · 30 days
0
Access
Public
Updated Jan 3, 2026
Repo size
18.9 GB
Likes
0
Public
Click a slice to open those files.
.pt8.6 GB · 100%
From the Hugging Face model README
Experimental medical-domain LLM pretraining project.
⚠️ Research-only. Not for clinical, diagnostic, or production use.
This repository contains multiple pretraining checkpoints of the MedAssistGPT architecture, released in two parameter scales:
Both variants:
The purpose of this repository is to document architecture choices, data pipelines, and large-scale training behavior, rather than to present a fully converged or production-ready medical language model.
All models are decoder-only Transformers implemented from scratch in PyTorch.
tiktoken p50k_base| Variant | Parameters | d_model | Heads | GQA (KV heads) | Blocks |
|---|---|---|---|---|---|
| 303M | ~303M | 1024 | 16 | 4 | 24 |
| 401M | ~401M | 1024 | 32 | 4 | 24 |
Both variants use the same architectural template; the 401M model increases attention width while preserving GQA.
| Item | Value |
|---|---|
| Dataset | japhba/pubmed_simple |
| Text field | abstract |
| Domain | Biomedical / medical research |
| Cleaning | Minimal (raw abstracts) |
| Sequence length | 1,024 |
| Sliding window stride | 512 |
| Item | Value |
|---|---|
| Objective | Causal language modeling (next-token prediction) |
| Optimizer | AdamW |
| Betas | (0.9, 0.95) |
| Precision | bf16 |
| Gradient accumulation | Enabled |
| Gradient clipping | 1.0 |
| Effective batch size | 128 |
The checkpoints/ directory contains multiple snapshots of the same model variants at different training stages.
Examples:
checkpoint_step_25000.pt (303M) → ~2.5B tokens seen⚠️ Important:
All released checkpoints are early-stage pretraining snapshots.
At ~2.5B tokens (~8× tokens/parameter for 303M), the models are undertrained and should not be treated as finished base models.
They are provided to:
from transformers import AutoTokenizer, AutoModelForCausalLM
repo_id = "kunjcr2/MedAssistGPT-303M" # or 401M repo
tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(repo_id, trust_remote_code=True)
prompt = "A patient was admitted with severe headache. Initial assessment revealed"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(
**inputs,
max_new_tokens=100,
temperature=0.7,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
This repository is intended for:
🚫 Not intended for clinical or production medical use.
Apache 2.0