Downloads · 30 days
1.4K
5% of all-time downloads
Sweaterdog/MindCraft-LLM-tuning
MindCraft-LLM-tuning is a machine learning model from Sweaterdog. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
- Developed by: Sweaterdog - License: apache-2.0 - Finetuned from model : unsloth/Qwen2.5-7B-bnb-4bit and unsloth/Llama-3.2-3B-Instruct
Downloads · 30 days
1.4K
5% of all-time downloads
All-time downloads
27.3K
Public
Repo size
136 GB
Likes
9
Public
Click a slice to open those files.
.gguf67.3 GB · 100%
From the Hugging Face model README
The MindCraft LLM tuning CSV file can be found here, this can be tweaked as needed. MindCraft-LLM
This model is built and designed to play Minecraft via the extension named "MindCraft" Which allows language models, like the ones provided in the files section, to play Minecraft.
Why a new model?
While, yes, models that aren't fine tuned to play Minecraft Can play Minecraft, most are slow, innaccurate, and not as smart, in the fine tuning, it expands reasoning, conversation examples, and command (tool) usage.
What kind of Dataset was used?
I'm deeming the first generation of this model, Hermesv1, for future generations, they will be named "Andy" based from the actual MindCraft plugin's default character. it was trained for reasoning by using examples of in-game "Vision" as well as examples of spatial reasoning, for expanding thinking, I also added puzzle examples where the model broke down the process step by step to reach the goal.
Why choose Qwen2.5 for the base model?
During testing, to find the best local LLM for playing Minecraft, I came across two, Gemma 2, and Qwen2.5, these two were by far the best at playing Minecraft before fine-tuning, and I knew, once tuned, it would become better.
If Gemma 2 and Qwen 2.5 are the best before fine tuning, why include Llama 3.2, especially the lower intelligence, 3B parameter version?
That is a great question, I know since Llama 3.2 3b has low amounts of parameters, it is dumb, and doesn't play minecraft well without fine tuning, but, it is a lot smaller than other models which are for people with less powerful computers, and the hope is, once the model is tuned, it will become much better at minecraft.
NOTE: Andy-3.5-mini performs better than Andy-v2-qwen, and is only 1.5B parameters, instead of 7B parameters!
Why is it taking so long to release more tuned models?
Well, you see, I do not have the most powerful computer, and Unsloth, the thing I'm using for fine tuning, has a google colab set up, so I am waiting for GPU time to tune the models, but they will be released ASAP, I promise.
Will there ever be vision fine tuning?
Yes! In MindCraft there will be vision support for VLM's (vision language models), Most likely, the model will be Qwen2-VL-7b, or LLaMa3.2-11b-vision since they are relatively new, yes, I am still holding out hope for llama3.2
In order to use this model, A, download the GGUF file of the version you want, either a Qwen, or Llama model, and then the Modelfile, after you download both, in the Modelfile, change the directory of the model, to your model. Here is a simple guide if needed for the rest:
Download the .gguf Model u want. For this example it is in the standard Windows "Download" Folder
Download the Modelfile
Open the Modelfile with / in notepad, or you can rename it to Modelfile.txt, and change the GGUF path, for example, this is my PATH "C:\Users\SweaterDog\OneDrive\Documents\Raw GGUF Files\Hermes-1.0\Hermes-1.Q8_0.gguf"
Safe + Close Modelfile
Rename "Modelfile.txt" into "Modelfile" if you changed it before-hand
Open CMD and type in "ollama create Hermes1 -f Modelfile" (You can change the name to anything you'd like, for this example, I am just using the same name as the GGUF)
Wait until finished
In the CMD window, type "ollama run Hermes1" (replace the name with whatever you called it)
(Optional, needed for versions after the 11/15/24 update) If you downloaded a model that was tuned from Qwen, and in the model name you kept Qwen, you need to go into the file "prompter.js" and remove the qwen section, if you named it something that doesn't include qwen in the name, you can skip this step.
@echo off
setlocal enabledelayedexpansion
:loop
node main.js
timeout /t 10 /nobreak
echo Restarting...
goto loop
For Anybody who is wondering what the context length is, the Qwen version, has a length of 64000 tokens, for the Llama version, it has a 128000 token context window. (NOTE Any model can support a longer context, but these are the supported values in training)
UPDATE The Qwen and Llama models are out, with the expanded dataset! I have found the llama models are incredibly dumb, but changing the Modelfile may provide better results, With the Qwen version of Andy, the Q4_K_M, it took 2 minutes to craft a wooden pickaxe, collected stone after that, took 5 minutes. UPDATE DO NOT USE THE MODELS FROM THIS REPO, IT TOOK ~3:00.00 TO GET A STONE PICKAXE, MUCH FASTER THAN ANDY-V2-QWEN