Downloads · 30 days
410
100% of all-time downloads
mindlab-research/Macaron-V1.1
Macaron-V1.1 is a text generation model from mindlab-research. Use it when you need the model to write or continue text. The card lists the license as mit.
<div align="center" <img src="assets/mindlablogo.svg" width="32%" alt="MindLab logo"/ </div
Downloads · 30 days
410
100% of all-time downloads
All-time downloads
410
Public
Parameters
753B
1.6 TB on disk
Likes
10
Public
Click a slice to open those files.
.safetensors1.6 TB · 100%
How the weights are stored.
BF16753B · 100%
From the Hugging Face model README
Macaron-V1.1 is a 752B-parameter agent model from MindLab Research, post-trained on GLM-5.3. It combines a 744B base model with four 2B LoRA specialists for Chat, Agent, Coding, and Generative UI.
Built with Mint Recursive, Mind Lab's serverless training and inference platform, Macaron-V1.1 follows Macaron-V1, which was based on GLM-5.2. The iteration from V1 to V1.1 was completed in two weeks, using a shared workflow for data preparation, experiments, checkpoint evaluation, and deployment.
| Field | Value |
|---|---|
| Model name | Macaron-V1.1 |
| Organization | MindLab Research |
| Base model | GLM-5.3 |
| Architecture | GLM-5.3 base + Mixture of LoRA (MoL) specialists |
| Parameter footprint | 752B release label: 744B base + four 2B LoRA specialists |
| Specialists | Chat, Agent, Coding, Generative UI |
| Post-training platform | Mint Recursive |
| Primary domains | Chat, personal-agent and office workflows, coding, Generative UI |
| License | MIT |
| Specialist | Release-labeled size | Focus |
|---|---|---|
| Chat | 2B | Task progress and interaction quality in open-ended conversations |
| Agent | 2B | Tool selection and execution in complex workflows |
| Coding | 2B | Long-horizon software engineering and terminal tasks |
| Generative UI | 2B | UI delivery, functionality, and visual design |
The specialists share a 744B GLM-5.3 base.

| Category | Benchmark | Macaron V1.1 | Macaron V1 | GLM 5.3 | DeepSeek V4 Pro 0813 | Qwen 3.8 Max | Kimi K3 | Claude Opus 5 |
|---|---|---|---|---|---|---|---|---|
| Chat | ChatBench v2 | 66.7 | 63.6 | 65.8 | 65.7 | 62.2 | 57.7 | 66.2 |
| Agent | AutomationBench | 53.7 | 31.8 | 48.2* | 43.2* | 39.8* | 46.7* | 50.3* |
| Agent | Toolathlon-Verified | 76.0 | 63.0 | 73.0* | 74.1* | 72.5* | 73.2* | 80.6* |
| Coding | DeepSWE v1.1 | 71.7 | 58.4 | 66.9* | 62.7* | 57.0* | 69.0* | 68.8* |
| Coding | SWE-Marathon | 50.0 | 15.0 | 42.5* | 10.6* | — | 48.1* | 50.0* |
| Coding | Terminal-Bench 3.0 | 31.4 | 5.7 | 28.7 | 11.8* | 29.0* | 17.7* | 42.7* |
| GenUI | UI4ABench v2 | 83.2 | 75.8 | 77.6 | 77.7 | 79.3 | 80.6 | 81.4 |
Higher is better. Bold marks the highest displayed score per row; * marks an externally sourced benchmark result; — means unavailable. The table reproduces the announcement chart, including its V1 comparison column. External results retain their original protocols and may differ in harness, task subset, execution budget, and aggregation. The table does not establish a controlled overall model ranking.
Macaron-V1.1 exceeds or matches the displayed Opus 5 scores on ChatBench v2, AutomationBench, DeepSWE v1.1, SWE-Marathon, and UI4ABench v2. It scores below Opus 5 on Toolathlon-Verified and Terminal-Bench 3.0. These comparisons retain the protocol qualifications above.
| Benchmark | Reported setup and metric |
|---|---|
| ChatBench v2 | Three independent runs per model–case pair using the production system prompt, user persona, and relevant conversation history. A privately deployed GLM-5.2 judge rates responses on a 1–5 scale; ratings are averaged over runs and cases and multiplied by 20. |
| AutomationBench | Public 600-task split of v1.0.6 across six business domains, using the API toolset with at most 50 model-response steps per task. |
| Toolathlon-Verified | Official Toolathlon evaluation service; pass@1. |
| DeepSWE v1.1 | Claude Code agent harness; pass@3, with a task solved if any of up to three attempts succeeds. External official leaderboard results use mini-swe-agent. |
| SWE-Marathon | pass@2. External leaderboard scores retain their respective published settings. |
| Terminal-Bench 3.0 | Claude Code v2.1.207; pass@3. Client-reported output limit of 32,000 tokens per response. Agent timeouts are 10× each task's configured timeout, corresponding to 5–80 hours. |
| UI4ABench v2 | 60 real-world-inspired UI-generation tasks of varying difficulty. Delivery measures first-attempt success; functional completeness is assessed through execution tests and visual quality through rendered screenshots. |
The announcement chart identifies these external sources: DeepSWE, SWE-Marathon, Toolathlon results, the Claude Opus 5 System Card for its Toolathlon score, BenchLM for its Terminal-Bench score, and the DeepSeek blog for Kimi K3 and DeepSeek V4 Pro Terminal-Bench scores. These are source attributions from the supplied chart, not independently revalidated leaderboard snapshots.
The Toolathlon-Verified, DeepSWE v1.1, and Terminal-Bench 3.0 results are also registered in .eval_results/ against Hugging Face benchmark datasets, so Hugging Face can link them to their benchmark leaderboards. These files do not include HF Jobs verifyToken values; they should therefore be displayed as self-reported rather than verified results. ChatBench v2, AutomationBench, SWE-Marathon, and UI4ABench v2 remain model-card self-reported results because their corresponding HF benchmark registrations are not available for this release.
Reproducibility details pending: checkpoint and adapter revisions, task-set revisions where unspecified, evaluation dates, complete generation settings, raw traces, confidence intervals, and train/evaluation overlap analysis. The ChatBench judge belongs to the same GLM family as the base model; cross-family judge calibration is not provided in the announcement.
Macaron-V1.1 was trained with Mint Recursive, Mind Lab's serverless model-training platform. The workflow integrates data preparation, experimentation, checkpoint evaluation, and deployment, enabling the iteration from Macaron-V1 to V1.1 to be completed in two weeks.
Training focused on four capability areas:
Consult the platform documentation for available model IDs, authentication, pricing, and rate limits.
This repository includes the base checkpoint at the repository root and four LoRA specialists under loras/L0 through loras/L3. Each LoRA directory contains adapter_config.json and adapter_model.safetensors, following the Macaron-V1-Venti release layout.
This repository is released under the MIT License. Users should also respect any requirements inherited from the GLM-5.3 base model and from dependencies used by the serving harness.