Downloads · 30 days
19
22% of all-time downloads
choco800/qwen3-4b-agent-v15
qwen3-4b-agent-v15 is a text generation model from choco800. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
This repository provides a fully merged model fine-tuned from Qwen/Qwen3-4B-Instruct-2507 using Unsloth.
Downloads · 30 days
19
22% of all-time downloads
All-time downloads
88
Public
Parameters
4B
8.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8 GB · 100%
From the Hugging Face model README
This repository provides a fully merged model fine-tuned from Qwen/Qwen3-4B-Instruct-2507 using Unsloth.
Unlike standard adapter repositories, this repository contains the merged weights, meaning you do not need to load the base model separately.
This model is trained to improve multi-turn agent task performance on ALFWorld (household tasks).
Loss is applied to all assistant turns in the multi-turn trajectory, enabling the model to learn environment observation, action selection, tool use, and recovery from errors.
train_on_responses_only was applied to <|im_start|>assistant\n).from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "choco800/qwen3-4b-agent-v15"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
Training data:
Dataset License: MIT License. These datasets are used and distributed under the terms of the MIT License. Compliance: Users must comply with the dataset licenses and the base model's original terms of use (Apache 2.0).