Downloads · 30 days
38
18% of all-time downloads
Ma7ee7/MeetInstruct-0.6B-v1.0
MeetInstruct-0.6B-v1.0 is a text generation model from Ma7ee7. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
MeetInstruct-0.6B-v1.0 is the first finished release in the MeetInstruct series of small, general-purpose instruction-tuned language models by Ma7ee7.
Downloads · 30 days
38
18% of all-time downloads
All-time downloads
207
Public
Parameters
596M
1.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.2 GB · 99%
From the Hugging Face model README
MeetInstruct-0.6B-v1.0 is the first finished release in the MeetInstruct series of small, general-purpose instruction-tuned language models by Ma7ee7.
It is based on:
Qwen/Qwen3-0.6B-Base
Hugging Face repository:
Ma7ee7/MeetInstruct-0.6B-v1.0
MeetInstruct is an experiment in how much useful assistant behavior can be extracted from a relatively small pretrained language model through carefully designed post-training.
The goal is not to turn a 0.6B model into a benchmark-specialized reasoning system.
The goal is to make a small model that is simply pleasant and useful to interact with.
MeetInstruct is intended to be a general-purpose instruct/chat model.
The series focuses on behaviors such as:
A major inspiration for the series is the kind of behavioral quality commonly associated with models such as GPT-4.5:
natural language, flexible tone, creativity, nuance, good wording, and responses that feel less mechanical.
This does not mean MeetInstruct attempts to reproduce GPT-4.5's capabilities.
At 0.6B parameters, the difference in raw capability is enormous.
The inspiration is instead about the direction of the post-training: making the model communicate well rather than optimizing it around one narrow benchmark or task family.
| Property | Value |
|---|---|
| Model | Ma7ee7/MeetInstruct-0.6B-v1.0 |
| Series | MeetInstruct |
| Version | v1.0 |
| Base model | Qwen/Qwen3-0.6B-Base |
| Parameters | ~0.6B |
| Model type | General-purpose instruct/chat model |
| Training method | Full-parameter supervised fine-tuning |
| Primary SFT context | 4,096 tokens |
| Long-context polish | 8,192 tokens |
| Explicit reasoning training | No |
| Visible chain-of-thought training | No |
| DPO / preference optimization | No |
| Primary language | English |
MeetInstruct-0.6B-v1.0 was built around a relatively simple idea:
The pretrained model already knows language. Post-training should primarily teach it how to behave.
Instead of performing an extremely large instruction-tuning run, v1.0 uses a relatively small and targeted post-training budget.
The intention was to move the base model toward assistant behavior without unnecessarily overwriting the representations learned during pretraining.
The pipeline therefore prioritizes:
MeetInstruct-0.6B-v1.0 uses a two-stage supervised fine-tuning pipeline.
The first stage performs the majority of the behavioral adaptation.
Approximate configuration:
| Setting | Value |
|---|---|
| Context length | 4,096 |
| Training budget | ~32M nominal tokens |
| Peak learning rate | 1.5e-5 |
| Scheduler | Cosine |
| Warmup | 3% |
| Weight decay | 0.1 |
| Training | Full-parameter |
| Loss | Assistant-only causal LM loss |
| Packing | Enabled |
This stage is responsible for most of the model's transition from a pretrained base model into a conversational assistant.
The target was broad usefulness rather than specialization.
The second stage is much smaller.
Its purpose is to polish behavior while exposing the instruction-tuned model to longer conversations.
Approximate configuration:
| Setting | Value |
|---|---|
| Context length | 8,192 |
| Training budget | ~6M nominal tokens |
| Peak learning rate | 4e-6 |
| Scheduler | Cosine |
| Warmup | 5% |
| Weight decay | 0.05 |
| Training | Full-parameter |
| Loss | Assistant-only causal LM loss |
Examples for this stage were preferentially selected from higher-quality and longer conversations.
This stage was intentionally kept small.
It was not meant to relearn assistant behavior from scratch, but rather to refine the Stage 1 checkpoint.
The complete v1.0 supervised post-training run targeted approximately:
| Stage | Context | Nominal training tokens |
|---|---|---|
| Stage 1 | 4,096 | ~32M |
| Stage 2 | 8,192 | ~6M |
| Total | — | ~38M |
These numbers refer to the approximate training-token budget, not necessarily unique tokens.
The relatively small budget was intentional.
MeetInstruct-0.6B-v1.0 uses a mixture of several instruction and conversational datasets.
The primary sources were:
HuggingFaceTB/smol-smoltalkUsed as the largest component of the general instruction mixture.
It provides broad assistant-oriented examples suitable for relatively small language models.
argilla/magpie-ultra-v1.0Used for additional diversity across instructions, general questions, writing, coding, editing, and other assistant tasks.
Reasoning-oriented examples were filtered where possible.
HuggingFaceH4/no_robotsUsed as a source of human-written instruction and response demonstrations.
This is valuable because much modern instruction data is synthetic.
Human-written examples provide a useful counterweight to model-generated response styles.
OpenAssistant/oasst2Used primarily for genuine multi-turn conversation.
OASST2 is structured as a conversation tree rather than a simple prompt-response dataset.
For MeetInstruct, conversational branches were reconstructed from the message tree to produce usable multi-turn examples.
This helps teach behavior such as:
The preprocessing pool used approximately:
| Dataset | Target pool size |
|---|---|
| Smol-SmolTalk | ~55,000 |
| Magpie Ultra | ~25,000 |
| No Robots | ~9,500 |
| OpenAssistant 2 | ~15,000 |
The full pool was larger than the actual amount of data consumed during training.
Training duration was controlled primarily by a token-derived step budget, rather than simply performing multiple epochs over the complete dataset.
This was done to make the amount of post-training more predictable.
MeetInstruct-0.6B-v1.0 was trained using assistant-only supervision.
Conceptually:
System message → ignored by loss
User message → ignored by loss
Assistant response → trained
The user and system messages remain part of the model's context, but gradient loss is concentrated on the tokens the assistant is expected to generate.
This makes instruction tuning more directly about learning the desired response behavior.
The training pipeline does not assume that longer responses are automatically better.
Very short examples are intentionally retained.
For example:
User: 17 * 24?
Assistant: 408
is a perfectly useful instruction-tuning example.
This matters for tasks such as:
One of the goals of MeetInstruct is to avoid teaching the model that every request deserves a large answer.
MeetInstruct-0.6B-v1.0 is not a reasoning-specialized model.
The training pipeline explicitly filters visible reasoning patterns such as:
<think>
...
</think>
as well as obvious chain-of-thought-style response structures.
The goal is not to prevent the model from solving problems.
It can still perform normal inference, calculations, explanations, and problem solving.
The distinction is that the model was not intentionally trained to make long visible reasoning traces part of its normal response format.
MeetInstruct v1.0 is intended to behave more like:
User:
Why does ice float?
Assistant:
Ice floats because its crystal structure makes solid water less dense than liquid water.
rather than automatically producing a long hidden-thought-style transcript before every answer.
Reasoning-specialized variants may be explored separately in the future.
The underlying Qwen3-0.6B architecture supports substantially more context than the main instruction-training length.
MeetInstruct v1.0 was primarily post-trained at:
This was a deliberate compute tradeoff.
Training the entire post-training corpus at extremely long context lengths would have substantially increased compute cost while providing relatively little benefit for most everyday assistant conversations.
The smaller 8K second stage provides some longer-context exposure without making long sequences dominate the training budget.
Users should not interpret the base architecture's maximum supported context length as a guarantee that v1.0 will maintain equal quality across the entire window.
MeetInstruct-0.6B-v1.0 is intended primarily for experimentation with small conversational language models.
Potential uses include:
Because of its relatively small parameter count, it may also be useful as a starting point for specialized downstream variants.
MeetInstruct-0.6B-v1.0 is not intended to be:
The goal of this release is intentionally broader:
Make a small model into a competent, natural general assistant.
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_id = "Ma7ee7/MeetInstruct-0.6B-v1.0"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
)
messages = [
{
"role": "user",
"content": "Explain what RAM does in a computer in two sentences.",
}
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
inputs = tokenizer(
text,
return_tensors="pt",
).to(model.device)
with torch.no_grad():
output = model.generate(
**inputs,
max_new_tokens=256,
do_sample=True,
temperature=0.7,
top_p=0.8,
top_k=20,
repetition_penalty=1.05,
)
generated = output[0, inputs["input_ids"].shape[1]:]
print(
tokenizer.decode(
generated,
skip_special_tokens=True,
)
)
A reasonable starting point for ordinary chat:
do_sample = True
temperature = 0.7
top_p = 0.8
top_k = 20
repetition_penalty = 1.05
For tasks where deterministic output matters more:
do_sample = False
Generation settings are task-dependent, so these should be treated as starting points rather than universal defaults.
MeetInstruct-0.6B-v1.0 is still a 0.6B parameter model.
Its size places significant limits on its capabilities.
It may struggle with:
The model may confidently produce incorrect information.
It can also misunderstand prompts, lose track of conversational details, repeat itself, or produce responses that are less nuanced than larger models.
The objective of MeetInstruct is to make efficient use of a small model, not to pretend those size limitations do not exist.
MeetInstruct is an ongoing small-model post-training project by Ma7ee7.
The series explores how much general assistant quality can be achieved through:
MeetInstruct-0.6B-v1.0 is the first completed release and establishes the baseline for the series.
Future versions may alter the training recipe substantially rather than simply adding more data.
Development after v1.0 focuses on MeetInstruct-0.6B-v1.5.
The goal for v1.5 is not merely to train v1.0 for longer.
The post-training pipeline is being reconsidered from the ground up, including:
The central goal remains the same:
A small, general-purpose instruct/chat model that communicates naturally and flexibly.
Particular attention is being given to qualities such as:
v1.5 is intended to improve the behavioral quality of the series rather than simply chase higher benchmark scores.
MeetInstruct-0.6B-v1.0 is derived from:
Qwen/Qwen3-0.6B-Base
Please refer to the original Qwen3 model card for details about the base architecture, pretraining, tokenizer, licensing, and base-model limitations.
Apache License 2.0
Use of the model should also respect the licenses and terms associated with the original base model and the datasets used during post-training.
MeetInstruct-0.6B-v1.0 is an experimental language model.
Its outputs may be incorrect, misleading, biased, inappropriate, or otherwise unreliable.
Important information should be independently verified before being relied upon.