Downloads · 30 days
81
10% of all-time downloads
mjpsm/activity-generation-model-v1.1
activity-generation-model-v1.1 is a text generation model from mjpsm. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
mjpsm/activity-generation-model-v1.1 is a fine-tuned activity-generation model designed for the MyVillage learning workflow. It generates one small next learning activity from a learner's current context.
Downloads · 30 days
81
10% of all-time downloads
All-time downloads
809
Public
Parameters
494M
1000 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors988 MB · 99%
From the Hugging Face model README
mjpsm/activity-generation-model-v1.1 is a fine-tuned activity-generation model designed for the MyVillage learning workflow. It generates one small next learning activity from a learner's current context.
The model takes three pieces of context:
It returns exactly three activity fields:
{
"title": "string",
"description": "string",
"instructions": "string"
}
Version 1.1 was trained to emphasize micro-progression: rather than turning each next activity into a large assignment or multi-step project, it should generate the smallest meaningful next step that builds on demonstrated knowledge.
| Property | Value |
|---|---|
| Model | mjpsm/activity-generation-model-v1.1 |
| Base model | Qwen/Qwen2.5-0.5B-Instruct |
| Previous version | mjpsm/activity-generation-model-v1 |
| Task | Conditional educational activity generation |
| Output | JSON containing title, description, and instructions |
| Fine-tuning approach | Supervised fine-tuning with LoRA |
| Language | English |
The LoRA adapter was merged back into the base model for the published standalone model. Users therefore do not need PEFT or the original adapter to run this repository.
The model is intended to generate a short next activity after a learner completes an activity and submits a description of what they learned.
A typical flow is:
Village goal
+
Previous activity title
+
Knowledge submission
↓
activity-generation-model-v1.1
↓
Next micro-activity
The model is best suited for learning systems where activities should progress incrementally rather than assigning a large project after every submission.
V1.1 was compared with mjpsm/activity-generation-model-v1 using the same fixed set of 20 benchmark cases, identical prompts, and deterministic decoding.
| Metric | V1 | V1.1 | Change |
|---|---|---|---|
| Valid JSON | 100% | 100% | Maintained |
| Exact output schema | 100% | 100% | Maintained |
| Single-sentence instructions | 40% | 95% | +55 pp |
| Micro-activity heuristic pass | 65% | 90% | +25 pp |
| Expected-focus hit | 80% | 95% | +15 pp |
| Sequencing-marker rate | 20% | 5% | -15 pp |
| Unnecessary-setup marker rate | 5% | 0% | -5 pp |
| Large-scope marker rate | 0% | 0% | Maintained |
| Instructions over 30 words | 25% | 5% | -20 pp |
| Forbidden-pattern hit rate | 0% | 0% | Maintained |
| Average instruction length | 25.1 words | 16.2 words | -35.5% |
| Average description length | 18.05 words | 12.15 words | -32.7% |
| Average generated tokens | 70.8 | 49.35 | -30.3% |
| Average benchmark latency | 2.18 s | 1.56 s | -28.6% |
The strongest measured change is activity simplicity. V1.1 produced single-sentence instructions in 95% of benchmark cases versus 40% for V1 and passed the benchmark's micro-activity heuristic in 90% of cases versus 65%.
V1.1 also preserved 100% valid-JSON and exact-schema compliance while producing substantially shorter outputs.
The micro-activity heuristic checks whether an output:
The expected-focus diagnostic checks whether an output contains one of the case-specific concepts associated with the learner's stated gap.
These are behavioral diagnostics rather than complete measures of educational quality.
A paired human-review sheet was generated for relevance, progression quality, Village-goal alignment, micro-step quality, unsupported assumptions, and clarity. No human scores were entered in the supplied benchmark results, so this model card does not report human-evaluation claims.
Village goal:
Understand how to train, evaluate, and improve machine learning models.
Previous activity title:
Train a Linear Regression Model
Knowledge submission:
I trained a linear regression model and got an R-squared value of 0.7, but I am not sure what that score means for the quality of my model.
{
"title": "Interpret Your R-Squared Score",
"description": "Learn what your R-squared score says about your model.",
"instructions": "Research what R-squared measures and write one sentence explaining what your score means."
}
The exact generated activity can vary. Applications should validate the returned JSON before storing or displaying it.
pip install torch transformers accelerate
torchao is not required.
import json
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
MODEL_ID = "mjpsm/activity-generation-model-v1.1"
SYSTEM_PROMPT = """You are an educational activity generator for MyVillage.
Given:
1. the Village goal,
2. the previous activity title, and
3. the student's knowledge submission,
generate exactly one small, realistic next learning activity.
The activity must:
- directly build on what the student demonstrated,
- move the student toward the Village goal,
- represent the smallest meaningful next step,
- stay short and focused,
- avoid unnecessary setup,
- avoid large or multi-step projects,
- avoid unsupported named tools, datasets, APIs, people, files, platforms, or requirements,
- allow the same activity title to appear for different students when the same next activity is appropriate.
Return valid JSON only with exactly these fields:
- title
- description
- instructions
Do not include markdown, commentary, activityType, estimatedMinutes, or extra fields.
"""
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
dtype = torch.float16 if torch.cuda.is_available() else torch.float32
model = AutoModelForCausalLM.from_pretrained(
MODEL_ID,
torch_dtype=dtype,
device_map="auto" if torch.cuda.is_available() else None,
)
if not torch.cuda.is_available():
model = model.to("cpu")
model.eval()
village_goal = (
"Develop foundational 3D modeling skills and learn how to create "
"detailed objects and environments using professional 3D software."
)
previous_activity_title = "Introduction to Basic 3D Modeling"
knowledge_submission = (
"I learned how to create basic 3D objects using cubes, spheres, and "
"cylinders. I still need practice combining shapes into more complex models."
)
user_prompt = f"""Village goal:
{village_goal}
Previous activity title:
{previous_activity_title}
Knowledge submission:
{knowledge_submission}
Generate the next activity."""
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": user_prompt},
]
prompt = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
inputs = tokenizer(
prompt,
return_tensors="pt",
).to(model.device)
with torch.inference_mode():
output = model.generate(
**inputs,
max_new_tokens=180,
do_sample=False,
repetition_penalty=1.05,
pad_token_id=tokenizer.eos_token_id,
eos_token_id=tokenizer.eos_token_id,
)
generated_tokens = output[0, inputs["input_ids"].shape[1]:]
response = tokenizer.decode(
generated_tokens,
skip_special_tokens=True,
).strip()
try:
activity = json.loads(response)
print(json.dumps(activity, indent=2))
except json.JSONDecodeError:
print("Model returned non-JSON output:")
print(response)
Because the model is based on Qwen2.5-0.5B, it can also be loaded on CPU. CPU inference will generally be slower than GPU inference.
For production use, deterministic or low-temperature generation is recommended because the model is intended to generate concise structured activities.
Applications should:
title, description, and instructions,Both model versions were evaluated with:
do_sample=False), andThis controls the inference setup so the comparison focuses on model behavior rather than different prompting strategies.
Latency measurements are included for completeness but are hardware- and runtime-dependent and should not be treated as a universal model-speed benchmark.
Previous activity-generation model used as the baseline for the V1.1 evaluation.
This model generates educational activity suggestions. Outputs should be treated as generated recommendations rather than guaranteed pedagogically optimal activities. Systems using the model should validate outputs and apply appropriate human oversight for their learning environment.