Downloads Β· 30 days
75
13% of all-time downloads
Kwai-Keye/Keye-VL-671B-A37B
Keye-VL-671B-A37B is a video-text-to-text model from Kwai-Keye. Use it for the video-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
<div align="center" <img src="figures/keyelogo2.png" width="100%" alt="Kwai Keye-VL Logo" </div
Downloads Β· 30 days
75
13% of all-time downloads
All-time downloads
563
Public
Parameters
672B
674 GB on disk
Likes
20
Public
Click a slice to open those files.
.safetensors674 GB Β· 100%
How the weights are stored.
F8_E4M3669B Β· 100%
From the Hugging Face model README
<font size=3><div align='center' > [π X] [π¬ Discord] [π Home Page] [π Technical Report] [π» GitHub Repository] [π Keye-VL-8B-Preview ] [π Keye-VL-1.5-8B ] [π Demo]
</div></font>Meet Keye-VL-671B-A37B β the most powerful multi-modal language model in the Keye series to date.
As one of the largest and most capable MLLMs currently in existence, Keye-VL-671B-A37B demonstrates top-tier and in some cases even leading performance in text understanding and generation, complex visual perception and reasoning, comprehensive video understanding, and Olympic-level mathematical reasoning.
Efficient Perception Building with Limited Compute: We employ VisionEncoder from Keye-VL-1.5 and rigorously processed high-quality data to cost-effectively build the modelβs core perceptual capabilities, ensuring strong visual understanding without excessive computational overhead.
Multi-Modal Data Curation: We implement a automated data pipeline that performs strict filtering, re-sampling, and large-scale synthesis of structured VQA data, including OCR, charts, and tables. This end-to-end process significantly enhances the modelβs perception quality and generalization.
Reasoning Sustainment via Synthetic CoT Data: During the continual pretrain phase, we incorporate a diverse set of synthetically generated chain-of-thought (CoT) data. This ensures the model maintains its complex reasoning skills while progressing in perceptual pre-training.


docker run -it --gpus all lmsysorg/sglang:v0.5.2
# make sure each node use the following commands to install the custom SGLang branch
git clone -b keye-dpsk-infer-fp8-release https://github.com/Kwai-Keye/sglang.git sglang
pip install -e sglang/python[all]
MODEL_PATH=/path/to/Keye-VL-671B-A37B
DIST_INIT_ADDR="MASTER_NODE_IP:29500" # e.g. 192.168.1.100:29500
PORT=30000 # listening port on each node
python3 -m sglang.launch_server \
--model-path $MODEL_PATH \
--host 0.0.0.0 \
--port $PORT \
--tp-size 16 \
--nnodes 2 \
--node-rank 0 \
--dist-init-addr $DIST_INIT_ADDR \
--trust-remote-code \
--mm-attention-backend fa3 \
--attention-backend fa3 \
--disable-radix-cache \
--mem-fraction-static 0.8 \
--cuda-graph-max-bs 64 \
--model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 32}'
MODEL_PATH=/path/to/Keye-VL-671B-A37B
DIST_INIT_ADDR="MASTER_NODE_IP:29500" # e.g. 192.168.1.100:29500
PORT=30000 # listening port on each node
python3 -m sglang.launch_server \
--model-path $MODEL_PATH \
--host 0.0.0.0 \
--port $PORT \
--tp-size 16 \
--nnodes 2 \
--node-rank 1 \
--dist-init-addr $DIST_INIT_ADDR \
--trust-remote-code \
--mm-attention-backend fa3 \
--attention-backend fa3 \
--disable-radix-cache \
--mem-fraction-static 0.8 \
--cuda-graph-max-bs 64 \
--model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 32}'
For more deployment details, please refer to the Keye-VL-671B-A37B Deployment Tutorial.
import json
import requests
BASE_URL = "http://MASTER_NODE_IP:30000"
def generate(messages):
payload = {
"model": "",
"messages": messages,
"n": 1,
"temperature": 0.0,
"max_tokens": 256,
"top_k": 1,
"ignore_eos": False,
"skip_special_tokens": True,
}
resp = requests.post(
f"{BASE_URL}/v1/chat/completions",
headers={"Content-Type": "application/json"},
data=json.dumps(payload),
timeout=1800,
)
resp.raise_for_status()
return resp.json()
# Example: image + text
messages = [
{
"role": "user",
"content": [
{
"type": "image_url",
"image_url": {"url": "https://raw.githubusercontent.com/sgl-project/sglang/main/assets/logo.png"},
},
{"type": "text", "text": "Describe this image in detail./think"},
],
}
]
result = generate(messages)
print(result["choices"][0]["message"]["content"])