Downloads · 30 days
0
SEBK4C/Fluffy-LoRA
Fluffy-LoRA is a feature extraction model from SEBK4C. Use it when you need embeddings to search or compare text. It is set up for peft. The card lists the license as gemma.
One body, three heads. Text, image, audio — one embedding space.
Downloads · 30 days
0
Access
Public
Updated Jul 12, 2026
Repo size
131 MB
Likes
0
Public
Click a slice to open those files.
.safetensors131 MB · 100%
From the Hugging Face model README
One body, three heads. Text, image, audio — one embedding space.
Named for Fluffy, Hagrid's three-headed dog (Cerberus's better-tempered nephew). Each head is an input modality; the body is a single shared embedding space.
No bolted-on encoders — no CLIP/CLAP/Whisper glued together with projectors.
google/gemma-4-12b-it natively ingests all three modalities into one backbone;
we QLoRA the language tower (r=8, α=16, ~32.8M params), freeze the native towers,
and pool the last token into one L2-normed embedding. Trained with symmetric
InfoNCE (τ=0.02) against frozen, ratcheted retrieval evals.
The honest scoreboard and all code live in the GitHub repo: → https://github.com/SEBK4C/Fluffy-LoRA
fluffy-text-v0/ — alpha artifact from the abandoned v1 text-only runThis adapter is published as a research artifact, not because it passed the ratchet — the ratchet rejected every checkpoint of the run it came from. Full disclosure:
gemma-4-12b-it, trained with symmetric
InfoNCE (τ=0.02, in-batch negatives, last-token pooling, L2 norm,
max_length 512) on 40,941 fully self-synthetic text cards (no scraped or
third-party data — rights-clean).attention_mask.sum(1)-1 (a right-padding
assumption) — so every sequence that wasn't the longest in its batch was
pooled at a padding position. Two unrelated texts pooled this way sit at
cos 0.96 (their true last-token embeddings: cos 0.75). The adapter was
trained on a mostly-collapsed embedding function.nDCG@10 for retrieval, Spearman for STS. "fixed pool" = corrected last-token
pooling (h[:, -1] under left padding). Reference = Qwen3-Embedding-8B with
its own card protocol — it reproduces its published MTEB scores on this
harness, which validates the pipeline.
| Task | base gemma | base + this adapter | base (fixed pool) | + adapter (fixed pool) | reference |
|---|---|---|---|---|---|
| SciFact | 0.000 | 0.004 | 0.002 | 0.002 | 0.788 |
| NFCorpus | 0.013 | 0.010 | 0.009 | — | 0.414 |
| FiQA2018 | 0.000 | 0.000 | — | — | 0.612 |
| STSBenchmark | 0.021 | 0.034 | 0.035 | — | 0.935 |
| STS17 (en-en) | 0.357 | 0.295 | 0.315 | — | 0.957 |
Two honest take-aways: (1) the adapter moved no external metric — train it as we did and you learn the training data's domain, not transferable embedding geometry; (2) raw decoder last-token (or mean) embeddings of gemma-4-12b-it are near-unusable for retrieval without contrastive adaptation — the reference model's scores on the same harness show what a properly-trained embedder does. Full analysis and raw results: LEARNINGS-V1.md.
Treat it as a checkpoint of scientific interest (what ~1.4k steps of contrastive warmup on a broken pooling fn does to a decoder LM's embedding geometry), not as a useful embedding model.
import torch, torch.nn.functional as F
from transformers import AutoModel, AutoTokenizer
from peft import PeftModel
base = AutoModel.from_pretrained("google/gemma-4-12b-it", torch_dtype=torch.bfloat16)
if hasattr(base, "language_model"):
base = base.language_model # text tower only, as trained
model = PeftModel.from_pretrained(base, "SEBK4C/Fluffy-LoRA", subfolder="fluffy-text-v0").eval()
tok = AutoTokenizer.from_pretrained("google/gemma-4-12b-it")
enc = tok(["a query"], padding=True, truncation=True, max_length=512, return_tensors="pt")
h = model(**enc).last_hidden_state
idx = enc["attention_mask"].sum(1) - 1 # last real token
emb = F.normalize(h[torch.arange(h.shape[0]), idx].float(), dim=-1)
v1 (text-only) is over; v2 (three-lane multimodal) is being built in the open in the GitHub repo. Future checkpoints land here only when they beat the frozen evals by more than ε — the ratchet rejects everything else.