Downloads · 30 days
48
1% of all-time downloads
lgdevlop/Macaron-V1-Tall-l2-coding
Macaron-V1-Tall-l2-coding is a machine learning model from lgdevlop. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This repository contains the GGUF-quantized, pre-merged specialist models for Macaron-V1-Tall, optimized for local deployment on consumer hardware.
Downloads · 30 days
48
1% of all-time downloads
All-time downloads
5K
Public
Repo size
21.2 GB
Likes
0
Public
Click a slice to open those files.
.gguf21.2 GB · 100%
From the Hugging Face model README
This repository contains the GGUF-quantized, pre-merged specialist models for Macaron-V1-Tall, optimized for local deployment on consumer hardware.
All foundational research, training, and architecture design belong to the original authors. This repository solely provides a community-driven quantization and deployment format.
The Macaron-V1-Tall architecture utilizes a Mixture-of-LoRA (MoE) pattern over a Grouped-Query Attention (GQA) base model. Currently, the standard llama.cpp ecosystem encounters tensor reshaping limitations (NotImplementedError) when attempting to dynamically apply these specific LoRA adapters at runtime.
To bypass this mathematical constraint and enable seamless local execution, this repository utilizes a Pre-Merge Strategy. Each LoRA specialist has been physically fused into the 35B base model using CPU-RAM computation (bypassing VRAM bottlenecks) prior to GGUF conversion.
The Result: Four independent, fully fused GGUF models. Instead of dynamically swapping LoRAs in VRAM, developers can leverage operating system Page Caching to swiftly alternate between these ~21GB files in RAM, enabling rapid intent-based routing without out-of-memory (OOM) errors.
The model have been quantized to Q4_K_M to balance perplexity and memory footprint, reducing the required storage from ~70GB (FP16) to approximately 21GB per specialist.
macaron-l2-coding-35b-q4_k_m.gguf (Coding Specialist)Note: The experimental
mtp_num_hidden_layers(Multi-Token Prediction) metadata has been sanitized from the config to ensure strict compatibility with thellama.cpploader.
These models are heavily optimized for hybrid VRAM/RAM offloading. If you are running a consumer GPU (e.g., 16GB VRAM) backed by substantial system RAM, you must carefully balance the layers to prevent VRAM saturation.
llama-server)The following command demonstrates how to load the model while explicitly offloading the heaviest MoE layers to the CPU, freeing up your VRAM for the KV cache and attention heads.
./llama-server \
-m "macaron-l2-coding-35b-q4_k_m.gguf" \
--host 127.0.0.1 \
--port 9091 \
-ngl 40 \
-ncmoe 10 \
--ctx-size 16384 \
-b 512 \
-fa on \
-ctk q4_0 \
-ctv q4_0 \
-t 16 \
--reasoning-preserve \
--no-context-shift \
-sps 0.0 \
--ctx-checkpoints 0 \
--models-max 1 \
--parallel 1 \
--no-mmap \
--metrics