Downloads · 30 days
1.3K
10% of all-time downloads
TorpedoSoftware/Luau-Devstral-24B-Instruct-v0.2
Luau-Devstral-24B-Instruct-v0.2 is a text generation model from TorpedoSoftware. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
State-of-the-art Luau code generation through reinforcement learning post-training
Downloads · 30 days
1.3K
10% of all-time downloads
All-time downloads
13K
Public
Parameters
23.6B
301 GB on disk
Likes
6
Public
Click a slice to open those files.
.gguf199 GB · 81%
From the Hugging Face model README
State-of-the-art Luau code generation through reinforcement learning post-training
A refined version of Luau-Devstral-24B-Instruct-v0.1, enhanced with Dr. GRPO (Zichen Liu et al., 2025) to deliver superior Luau programming capabilities for Roblox development.
This model represents a significant advancement in specialized code generation for Luau, building upon continuous pretraining with targeted reinforcement learning to achieve exceptional code quality.
Key Achievements:
Evaluated on the test split of TorpedoSoftware/LuauLeetcode containing 226 challenges, with results averaged across 3 runs per challenge.
Base Models:
Competitive Benchmarks:
Note: OpenAI models utilize reasoning tokens as complete disabling of thinking is not available.
Measures problem-solving accuracy and correctness

Result: 4th place overall, demonstrating solid problem-solving capabilities while outperforming OpenAI models.
Evaluates fundamental code quality

Result: State-of-the-art performance with the lowest error rate by a significant margin.
Assesses non-critical code quality issues

Result: State-of-the-art performance in minimizing code warnings.
Strict mode typechecking compliance

Result: 2nd place, closely trailing Claude Opus 4.1. Our model favors explicit type definitions for enhanced code clarity, which creates more opportunities for mistakes compared to Claude's reliance on inferred types.
Edit distance from Stylua's standard format

Result: State-of-the-art performance with exceptional adherence to standard formatting conventions.
Average response size (excluding reasoning tokens)

Result: Most concise responses among all models, delivering direct solutions without unnecessary preamble. This efficiency suggests potential for further improvements in problem solving through explicit problem decomposition or reasoning.
Primary Source: TorpedoSoftware/LuauLeetcode
Curriculum Learning Approach:
Easy Difficulty Phase
Medium Difficulty Phase
Hard Difficulty Phase
Technical Configuration:
The model was optimized using four complementary reward signals:



Custom importance matrix computed using 5.73MB of specialized text data:
Calibration Sources:
This calibration ensures optimal performance for Luau/Roblox tasks while maintaining general intelligence. The imatrix.gguf file is included in the repository for custom quantization needs.
Carbon emissions estimated using the Machine Learning Impact calculator (Lacoste et al., 2019):