Downloads · 30 days
121
34% of all-time downloads
mlx-community/Tiny-Pickle-v3-Coder-4bit
Tiny-Pickle-v3-Coder-4bit is a text generation model from mlx-community. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as apache-2.0.
Tiny Pickle v3 Coder is a coding-focused adaptation of Qwen/Qwen3-Coder-30B-A3B-Instruct, converted to MLX and quantized for Apple Silicon.
Downloads · 30 days
121
34% of all-time downloads
All-time downloads
358
Public
Parameters
30.5B
17.2 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors17.2 GB · 100%
How the weights are stored.
U3230.5B · 100%
From the Hugging Face model README
Tiny Pickle v3 Coder is a coding-focused adaptation of
Qwen/Qwen3-Coder-30B-A3B-Instruct, converted to MLX and quantized for
Apple Silicon.
Qwen/Qwen3-Coder-30B-A3B-Instructvsan/tiny-pickle-v3-coder-LoRAvsan/tiny-pickle-v3-coderpip install -U mlx-lm
mlx_lm.chat --model mlx-community/Tiny-Pickle-v3-Coder-4bit
mlx_lm.generate \
--model mlx-community/Tiny-Pickle-v3-Coder-4bit \
--prompt "Write a tested Python implementation of an LRU cache." \
--max-tokens 800
Code generation, debugging, code review, implementation planning, test generation, and local software-engineering assistance on Apple Silicon.
This release uses MLX 4-bit affine quantization with group size 64. Quantization reduces storage and unified-memory requirements but may alter outputs or reduce quality relative to the merged BF16 model.
Tiny Pickle v3 Coder is experimental and has not yet been independently demonstrated to outperform its base model. Generated code may be incorrect, insecure, incomplete, or non-functional and must be reviewed and tested.
The following result is a single local inference measurement, not a standardized benchmark.
| Property | Result |
|---|---|
| Hardware | Apple M1 Max |
| Unified memory | 64 GB |
| Model format | MLX 4-bit affine |
| Quantization group size | 64 |
| Prompt length | 118 tokens |
| Prompt processing speed | 119.328 tokens/s |
| Generated length | 748 tokens |
| Generation speed | 63.481 tokens/s |
| Peak unified memory | 17.393 GB |
You are reviewing a Python async web crawler. Implement a complete, production-quality crawler that uses asyncio and aiohttp; limits global concurrency to 20; limits each domain to 2 concurrent requests; respects robots.txt; retries HTTP 429 and 5xx responses with exponential backoff and jitter; avoids duplicate URLs; normalizes relative links; restricts crawling to the starting domain; supports cancellation; records failures without stopping the crawl; and includes pytest tests using mocked HTTP responses. Return one self-contained Python module followed by the tests.
Results can vary with the macOS version, MLX-LM version, background processes, context length, sampling configuration, and thermal state.