Downloads Β· 30 days
0
Matrix-Corp/TouchGrass-7b
TouchGrass-7b is a text generation model from Matrix-Corp. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
A Lightweight Music AI Assistant Fine-Tuned from Qwen3.5
Downloads Β· 30 days
0
Access
Public
Updated Mar 10, 2026
Repo size
β
Likes
0
Public
Click a slice to open those files.
.py409 KB Β· 91%
From the Hugging Face model README
A Lightweight Music AI Assistant Fine-Tuned from Qwen3.5
Touch Grass is a specialized music AI assistant built by fine-tuning Qwen3.5 models (3B and 7B variants) with music-specific capabilities. It understands guitar, piano, drums, vocals, music theory, ear training, songwriting, and productionβwith emotional intelligence to help musicians through frustration.
TouchGrass/
βββ configs/ # Model configurations
β βββ touchgrass_3b_config.py # 3B variant config
β βββ touchgrass_7b_config.py # 7B variant config
β βββ training_config.py # Training hyperparameters
βββ tokenizer/
β βββ music_token_extension.py # Extends Qwen tokenizer with music tokens
βββ models/ # Specialized music modules
β βββ tab_chord_module.py # Guitar tabs and chords
β βββ music_theory_module.py # Theory knowledge
β βββ ear_training_module.py # Ear training exercises
β βββ eq_adapter.py # Emotional intelligence
β βββ songwriting_module.py # Song creation assistance
βββ data/
β βββ music_qa_generator.py # Synthetic dataset generator
β βββ chat_formatter.py # Qwen chat format converter
β βββ dataset_loader.py # PyTorch dataset
βββ training/
β βββ losses.py # Multi-task loss functions
β βββ trainer.py # LoRA-aware trainer
β βββ train.py # Main training entry point
βββ inference/
β βββ inference.py # Unified inference with context
βββ benchmarks/
β βββ evaluate_music_modules.py # Module-level benchmarks
β βββ evaluate_inference.py # End-to-end inference benchmarks
βββ tests/ # Comprehensive test suite
β βββ test_*.py # Unit tests for each module
β βββ conftest.py # Pytest fixtures
β βββ run_tests.py # Test runner
βββ configuration_touchgrass.py # HuggingFace config class
βββ tokenization_touchgrass.py # HuggingFace tokenizer wrapper
βββ ollama_3b_modelfile # Ollama config for 3B
βββ ollama_7b_modelfile # Ollama config for 7B
βββ train.py # Main training script
# Clone the repository
cd TouchGrass
# Install dependencies
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cpu
pip install transformers peft datasets accelerate tqdm pytest
# Optional: For GPU support
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu118
python -c "
from TouchGrass.data.music_qa_generator import MusicQAGenerator
from TouchGrass.data.chat_formatter import ChatFormatter
# Generate synthetic dataset
generator = MusicQAGenerator(seed=42)
dataset = generator.generate_dataset(num_samples=1000, output_path='data/music_qa.jsonl')
# Format for Qwen
formatter = ChatFormatter()
formatted = formatter.format_dataset(dataset)
train_data, val_data = formatter.create_splits(formatted, val_size=0.1)
formatter.save_dataset(train_data, 'data/train.jsonl')
formatter.save_dataset(val_data, 'data/val.jsonl')
"
# Train 3B variant
python train.py \
--base_model Qwen/Qwen3.5-3B-Instruct \
--train_data data/train.jsonl \
--val_data data/val.jsonl \
--output_dir checkpoints/touchgrass-3b \
--lora_r 16 \
--lora_alpha 32 \
--batch_size 4 \
--gradient_accumulation_steps 4 \
--learning_rate 2e-4 \
--num_epochs 3 \
--mixed_precision fp16
# Train 7B variant (requires GPU with 16GB+ VRAM)
python train.py \
--base_model Qwen/Qwen3.5-7B-Instruct \
--train_data data/train.jsonl \
--val_data data/val.jsonl \
--output_dir checkpoints/touchgrass-7b \
--lora_r 16 \
--lora_alpha 32 \
--batch_size 2 \
--gradient_accumulation_steps 8 \
--learning_rate 1e-4 \
--num_epochs 3 \
--mixed_precision bf16
from TouchGrass.inference.inference import TouchGrassInference
# Load model
model = TouchGrassInference(
model_path="checkpoints/touchgrass-3b",
device="cpu" # or "cuda"
)
# Single query with instrument context
response = model.generate(
prompt="How do I play a G major chord?",
instrument="guitar",
skill_level="beginner",
max_new_tokens=200
)
print(response)
# Interactive mode
model.chat(instrument="piano")
# Create modelfile from provided template
cat ollama_3b_modelfile > Modelfile
# Build and run
ollama create touchgrass-3b -f Modelfile
ollama run touchgrass-3b "How do I play a G major chord on guitar?"
from transformers import AutoModelForCausalLM, AutoTokenizer
# Load with custom config and tokenizer
config = TouchGrassConfig.from_pretrained("checkpoints/touchgrass-3b")
tokenizer = TouchGrassTokenizer.from_pretrained("checkpoints/touchgrass-3b")
model = AutoModelForCausalLM.from_pretrained(
"checkpoints/touchgrass-3b",
config=config,
device_map="auto"
)
# Generate
inputs = tokenizer("system\nYou are a music assistant.\nuser\nHow do I play a G major chord?\nassistant\n", return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=200)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Run the comprehensive test suite:
# Run all tests
python tests/run_tests.py
# Run with coverage
python tests/run_tests.py --coverage
# Run specific test categories
pytest tests/test_music_theory_module.py -v
pytest tests/test_tokenizer.py -v
pytest tests/test_eq_adapter.py -v
# Skip slow tests
pytest -m "not slow"
Evaluate model performance on music-specific tasks:
# Evaluate music modules
python benchmarks/evaluate_music_modules.py --device cpu --d_model 768
# Run inference benchmarks
python benchmarks/evaluate_inference.py --model_path checkpoints/touchgrass-3b --device cpu
Edit configs/training_config.py to customize:
lm_loss_weight=1.0 (primary language modeling)eq_loss_weight=0.1 (emotional intelligence)music_module_loss_weight=0.05 (specialized modules)The tokenizer extension adds these special tokens:
Domain tokens: [GUITAR], [PIANO], [DRUMS], [VOCALS], [THEORY], [PRODUCTION]
Emotion tokens: [FRUSTRATED], [CONFUSED], [EXCITED], [CONFIDENT]
Difficulty tokens: [EASY], [MEDIUM], [HARD]
Function tokens: [TAB], [CHORD], [SCALE], [INTERVAL], [PROGRESSION]
EQ tokens: [SIMPLIFY], [ENCOURAGE]
Music notation: All note names (C, C#, D, etc.), chord types (m, dim, aug, 7, maj7, etc.)
The EQ Adapter detects user frustration and adapts responses:
from TouchGrass.data.music_qa_generator import MusicQAGenerator
# Create custom templates
custom_templates = {
"guitar": [
{
"system": "You are a {instrument} specialist.",
"user": "How do I play {chord}?",
"assistant": "Place your fingers: {fingering}"
}
]
}
generator = MusicQAGenerator(templates=custom_templates, seed=123)
dataset = generator.generate_dataset(num_samples=500)
from TouchGrass.inference.inference import TouchGrassInference
model = TouchGrassInference(model_path="checkpoints/touchgrass-3b")
# Switch between instruments seamlessly
guitar_response = model.generate("How do I palm mute?", instrument="guitar")
piano_response = model.generate("What are the scales in C major?", instrument="piano")
theory_response = model.generate("Explain the circle of fifths", instrument="theory")
from transformers import LoraConfig
lora_config = LoraConfig(
task_type=TaskType.CAUSAL_LM,
r=32, # Rank (higher = more parameters)
lora_alpha=64, # Alpha (typically 2Γr)
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"], # Qwen attention modules
lora_dropout=0.1,
bias="none"
)
tab_validator: Confidence score [0, 1] for tab validitydifficulty: 3-class classification (easy/medium/hard)get_scale_from_key(key, mode): Returns scale notesdetect_chord_function(root, chord_type, key): Returns Roman numeralget_circle_of_fifths(): Returns 12-key circleconstruct_chord(root, chord_type): Returns chord notesanalyze_progression(progression, key): Returns functional analysisfp16 for NVIDIA, bf16 for newer GPUspython tests/run_tests.py)MIT License - see LICENSE file for details.
Made with β€οΈ for musicians everywhere.
Touch Grass - because even AI needs to remember to make music, not just talk about it.