Downloads · 30 days
0
cudo528/Darwin-9B-Opus-mlx-4bit
Darwin-9B-Opus-mlx-4bit is a text generation model from cudo528. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as apache-2.0.
This repository provides an MLX-optimized 4-bit quantized version of FINAL-Bench/Darwin-9B-Opus, an ultra-efficient advanced reasoning model.
Downloads · 30 days
0
Access
Public
Updated Aug 8, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.md3.7 KB · 71%
From the Hugging Face model README
This repository provides an MLX-optimized 4-bit quantized version of FINAL-Bench/Darwin-9B-Opus, an ultra-efficient advanced reasoning model.
진화적 머지(Evolutionary Merge) 및 레이어별 머지(Layer-wise Merge) 기법이 적용된 9B 체급의 강력한 모델을 Apple Silicon(Mac) 환경에서 하드웨어 리소스를 최소화하며 구동할 수 있도록 4-bit로 양자화 및 포맷 변환을 완료한 버전입니다.
⚠️ Hugging Face View Note: 허깅페이스 시스템이 MLX 4비트 패킹 포맷(
U32/BF16)을 파싱하는 과정에서 파라미터 수 및 텐서 타입이 일시적으로 다르게 표기될 수 있으나, 실제 4비트 양자화 가중치가 정상 적용된 파일입니다. 맥북 로컬 환경에서 정상적인 메모리 점유율로 구동됩니다.
<thought>)을 통한 고난도 논리 추론이 가능합니다.이 모델을 로컬 환경에서 구동하려면 최신 버전의 mlx-lm 라이브러리가 필요합니다.
pip install -U mlx-lm
from mlx_lm import load, generate
# 본인의 허깅페이스 저장소 아이디/이름으로 수정하여 사용하세요.
model, tokenizer = load("cudo528/Darwin-9B-Opus-mlx-4bit")
prompt = "수학적 귀납법의 원리를 중학생도 이해할 수 있게 쉬운 비유를 들어 한국어로 설명해줘."
response = generate(
model,
tokenizer,
prompt=prompt,
max_tokens=2048,
verbose=True # 모델의 심층 사고 과정(<thought>)을 터미널에서 실시간으로 보려면 True
)
print(response)
mlx_lm.generate --model cudo528/Darwin-9B-Opus-mlx-4bit --prompt "Write a Python script for a simple custom neural network layer." --max-tokens 1024
.safetensors packed)