Downloads · 30 days
4
2% of all-time downloads
Japhari/cds-maternal-4b
cds-maternal-4b is a text generation model from Japhari. Use it when you need the model to write or continue text. It is set up for peft.
A LoRA adapter fine-tuned on google/medgemma-4b-it for structured clinical data extraction from maternal health case narratives. Given a free-text case story, the model outputs a structured JSON object capturing triag…
Downloads · 30 days
4
2% of all-time downloads
All-time downloads
213
Public
Repo size
219 MB
Likes
0
Public
Click a slice to open those files.
.pt120 MB · 55%
From the Hugging Face model README
A LoRA adapter fine-tuned on google/medgemma-4b-it for structured clinical data extraction from maternal health case narratives. Given a free-text case story, the model outputs a structured JSON object capturing triage-relevant fields.
| Property | Value |
|---|---|
| Base model | google/medgemma-4b-it |
| Adapter type | LoRA (PEFT 0.19.1) |
| Task | Causal LM — clinical JSON extraction |
| LoRA rank (r) | 16 |
| LoRA alpha | 32 |
| Dropout | 0.05 |
| Target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Setting | Value |
|---|---|
| Epochs | 1 |
| Steps | 2 125 |
| Batch size | 2 |
| Peak learning rate | ~2 × 10⁻⁴ (linear warmup + cosine decay) |
| Total tokens seen | ~5.7 M |
Loss dropped rapidly from 3.20 → 0.13 in the first ~100 steps, then continued a slow, stable descent to ~0.118 at the end of training, with no signs of divergence or overfitting.
Evaluated on a held-out split at the end of epoch 1:
| Metric | Value |
|---|---|
| Eval loss | 0.1176 |
| Eval token accuracy | 94.74 % |
| Eval entropy | 0.1172 |
Token accuracy measures how often the model predicts the correct next token in the structured JSON output. At 94.74 % the adapter reliably reproduces field names, delimiters, and clinical values in the expected schema.
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel
import torch
base_model_id = "google/medgemma-4b-it"
adapter_id = "Japhari/cds-maternal-4b"
tokenizer = AutoTokenizer.from_pretrained(adapter_id)
base_model = AutoModelForCausalLM.from_pretrained(
base_model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
model = PeftModel.from_pretrained(base_model, adapter_id)
model.eval()
case = """
A 28-year-old woman at 36 weeks gestation presents with severe headache,
visual disturbances, and blood pressure of 160/110 mmHg. She has +3 proteinuria
on dipstick. No seizures reported.
"""
prompt = f"Extract the triage fields as JSON:\n{case.strip()}"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
output = model.generate(**inputs, max_new_tokens=512, do_sample=False)
print(tokenizer.decode(output[0], skip_special_tokens=True))
Inherits the license of the base model (google/medgemma-4b-it). Check Google's MedGemma terms before deployment.