Downloads · 30 days
50
61% of all-time downloads
AdityaPS/SpaceLLM_Single_Turn_QA
SpaceLLM_Single_Turn_QA is a text generation model from AdityaPS. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as apache-2.0.
A LoRA adapter for openai/gpt-oss-20b, fine-tuned for single-turn question answering over space-agency-derived text. This repository contains adapter weights only; the base model must be loaded separately.
Downloads · 30 days
50
61% of all-time downloads
All-time downloads
82
Public
Repo size
59.7 MB
Likes
0
Public
Click a slice to open those files.
.safetensors31.9 MB · 53%
From the Hugging Face model README
A LoRA adapter for openai/gpt-oss-20b, fine-tuned for single-turn question answering over space-agency-derived text. This repository contains adapter weights only; the base model must be loaded separately.
CAUSAL_LM)openai/gpt-oss-20b, revision 6cee5e81ee83917806bbde320786a8fb61efebeePlease note
- This adapter uses no retrieval. There is no RAG pipeline, no index, and nothing looked up at inference time.
- LoRA targets attention projections only —
q_proj,k_proj,v_proj,o_proj— notlm_head.- Training and evaluation references were generated by a teacher model and have not been independently fact-checked. Reported scores measure similarity to those references, not verified factual accuracy.
- This is a distinct artifact from the older
AdityaPS/SpaceLLM_v1checkpoint, which used a different configuration (lm_head-only, rank 32) and different data/evaluation. IfSpaceLLM_v1remains public, treat it as a separate, earlier experiment — its card and weights should not be conflated with this one.
Intended:
Out of scope:
Expert verification is recommended before relying on any specific technical claim this model produces.
The example loads the pinned base revision and this adapter, applies the Harmony chat template, forces the assistant final channel, and extracts only that channel's text — the same approach used in evaluation. Relying on skip_special_tokens=True alone is avoided because it can mix analysis-channel and final-channel text.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, Mxfp4Config
from peft import PeftModel
BASE_MODEL = "openai/gpt-oss-20b"
BASE_REVISION = "6cee5e81ee83917806bbde320786a8fb61efebee"
ADAPTER_REPO = "AdityaPS/SpaceLLM_Single_Turn_QA"
tokenizer = AutoTokenizer.from_pretrained(ADAPTER_REPO)
base_model = AutoModelForCausalLM.from_pretrained(
BASE_MODEL,
revision=BASE_REVISION,
torch_dtype=torch.bfloat16,
quantization_config=Mxfp4Config(dequantize=True),
device_map="auto",
)
model = PeftModel.from_pretrained(base_model, ADAPTER_REPO).eval()
question = "What was the primary objective of the Mars Science Laboratory mission?"
prompt = tokenizer.apply_chat_template(
[{"role": "user", "content": question}],
tokenize=False,
add_generation_prompt=True,
)
prompt += "<|channel|>final<|message|>" # force the final channel
inputs = tokenizer(prompt, return_tensors="pt", add_special_tokens=False).to(model.device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=512, # inference-example setting; evaluation used 2048 (see below)
do_sample=False,
repetition_penalty=1.05,
no_repeat_ngram_size=6,
)
generation = tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=False)
for stop_token in ("<|return|>", "<|end|>"):
generation = generation.split(stop_token)[0]
print("Answer:", generation.strip())
Note: this example uses max_new_tokens=512 for a quick, practical demo. This differs from the 2,048-token limit used during evaluation (see Evaluation) — treat this as an inference-example setting, not the evaluated protocol.
Data was built offline, with no retrieval stage used anywhere in training or inference:
mistral:7b teacher model.| Split | Usable examples |
|---|---|
| Train | 14,395 |
| Validation | 1,695 |
| Test (evaluated) | 845 |
These are the only counts supported for this artifact. A previously circulated figure of "3,428 documents / 16,831 training samples" does not match this adapter's training/evaluation split counts or file hashes and should not be cited for this repository.
| Setting | Value |
|---|---|
| LoRA target modules | q_proj, k_proj, v_proj, o_proj |
| LoRA rank ($r$) / alpha ($\alpha$) | 16 / 32 |
| LoRA dropout | 0.05 |
| Task type | CAUSAL_LM |
| Max sequence length | 2,048 tokens |
| Base weights | MXFP4, dequantized at load time |
| Compute precision | BF16 |
The base model is the published MXFP4 gpt-oss-20b checkpoint, loaded with MXFP4 dequantization and run in BF16 compute — this is a loading/compute configuration, not a separately distributed "BF16 version" of the model.
Test set: 845 single-turn records.
Test file SHA-256: 337a80d2375a053ab7d8e9e1c5ee2ec7b975f74bfe16db13bf0d1c5a60aa91f6
Adapter config SHA-256: ef641c55f5f961a86be20f364b39b80f7777053959caff1523b2348f00432748
Protocol: Base and adapter used identical tokenized prompts. All 845 records produced a final answer; none were skipped. Generation was deterministic (do_sample=False) with max_seq_len=2048, max_new_tokens=2048, repetition_penalty=1.05, and no_repeat_ngram_size=6, applied to both base and adapter. Only text from the Harmony final channel was scored; analysis-channel text was excluded.
Metrics: BERTScore computed with roberta-large; token F1 computed on lowercased whitespace tokens; exact match computed after stripping surrounding whitespace. A missing answer would score zero on all metrics while remaining in the denominator — no answers were missing for this test set.
References were produced by the same teacher-based generation process used for training data and have not received independent factuality verification.
| Model | BERTScore P | BERTScore R | BERTScore F1 | Token F1 | Exact match | Hit 2,048-token cap |
|---|---|---|---|---|---|---|
Base (gpt-oss-20b) | 0.781081 | 0.869184 | 0.822320 | 0.122433 | 0 / 845 | 3 / 845 |
| + SpaceLLM Single-Turn Adapter | 0.905112 | 0.890707 | 0.897506 | 0.371768 | 2 / 845 | 1 / 845 |
BERTScore F1 change: an absolute gain of 0.075186. This is not a "6% improvement," and it is not evidence of improved factual accuracy or domain understanding — it reflects greater semantic/stylistic similarity to teacher-generated references under this specific protocol.
An earlier unguarded diagnostic run at the same 2,048-token maximum, without repetition controls, produced outputs that fell into exact phrase loops. The final evaluation protocol added repetition_penalty=1.05 and 6-token n-gram blocking (no_repeat_ngram_size=6) to both base and adapter generations to address this.
| Model | Unguarded diagnostic (cap events) | Guarded final run (cap events) | Reduction |
|---|---|---|---|
| Base | 35 | 3 | 91.4% |
| Adapter | 28 | 1 | 96.4% |
This is reported as an engineering diagnostic, not a controlled ablation: prompt hashes changed across rebuilt evaluation images between the unguarded and guarded runs, so the reduction cannot be attributed to the decoding settings alone. Additionally, the guarded settings removed the conspicuous exact-repetition loops but did not eliminate semantic degeneration generally — a small number of guarded outputs still reached the token cap while producing drifting numerical lists or otherwise unrelated text.
mistral:7b teacher process used for training data; higher similarity scores may partly reflect the adapter learning to imitate the teacher's style, length, and any of its errors, rather than confirmed improvements in correctness.no_repeat_ngram_size=6 can suppress legitimate repeated technical phrases (e.g., a mission name or unit repeated for clarity), which may slightly affect fluency on some answers.openai/gpt-oss-20b also apply.337a80d2375a053ab7d8e9e1c5ee2ec7b975f74bfe16db13bf0d1c5a60aa91f6ef641c55f5f961a86be20f364b39b80f7777053959caff1523b2348f00432748adapter_model.safetensors, adapter_config.json, tokenizer files, chat_template.jinja, README.md.
Apache-2.0, matching the base model (openai/gpt-oss-20b). Note: this covers the adapter weights and code; it does not by itself establish licensing terms for the scraped source text or the teacher-generated training data used to produce this adapter.