Downloads · 30 days
296
79% of all-time downloads
amkyawdev/Ling-2.6-1T
Ling-2.6-1T is a text generation model from amkyawdev. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
<p align="center" <img src="https://mdn.alipayobjects.com/huameiqa8qxu/afts/img/A4QxcQrBlTiAAAAAAQXAAAAgAemJ7AQ/original" width="100"/ </p <p align="center"🤗 <a href="https://huggingface.co/inclusionAI"Hugging Face</…
Downloads · 30 days
296
79% of all-time downloads
All-time downloads
375
Public
Parameters
1T
1 TB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1 TB · 100%
How the weights are stored.
F8_E4M31T · 98%
From the Hugging Face model README
Today, we are thrilled to open-source Ling–2.6–1T from the Ling family.
Tailored for real–world, complex scenarios, this trillion–parameter model introduces targeted optimizations across inference efficiency, token overhead, and agentic capabilities, making it highly effective for coding and daily workflows.
Key upgrades in Ling–2.6–1T include:
On Artificial Analysis, Ling-2.6-1T achieved an Intelligence Index of 34 with approximately 16M output tokens, representing a significant generational leap over the previous Ling-1T. This positioning underscores its ability to deliver high-tier intelligence with optimized token consumption.
<p align="center"> <img src="https://mdn.alipayobjects.com/huamei_fst7or/afts/img/48cCTY8XJgUAAAAAZvAAAAgADpRXAQJr/original" /> </p> <p align="center"> <img src="https://mdn.alipayobjects.com/huamei_fst7or/afts/img/AmTNT5tQHDYAAAAAaSAAAAgADpRXAQJr/original " width="48%"/> <img src="https://mdn.alipayobjects.com/huamei_fst7or/afts/img/Wv_8Toxbl7IAAAAAaRAAAAgADpRXAQJr/original" width="48%"/> </p>Ling-2.6-1T demonstrates balanced excellence across reasoning, coding, and tool-calling, achieving open-source SOTA status on multiple execution-heavy benchmarks:
Note: If you are interested in the previous version, please visit the past model collections on Huggingface or ModelScope.
https://openrouter.ai/inclusionai/ling-2.6-1t:free
https://zenmux.ai/inclusionai/ling-2.6-1t
pip install uv
uv venv ~/my_ling_env
source ~/my_ling_env/bin/activate
# uv pip "sglang-kernel>=0.4.1"
uv pip install "sglang[all]>=0.5.10.post1" --prerelease=allow
Here is the example to run Ling-1T with 8 GPUs, where the server port is ${PORT}:
Server
1. Standard Inference (Without MTP)
sglang serve \
--model-path inclusionAI/Ling-2.6-1T \
--tp-size 8 \
--max-running-requests 32 \
--mem-fraction-static 0.92 \
--chunked-prefill-size 8192 \
--context-length 262144 \
--trust-remote-code \
--model-loader-extra-config '{"enable_multithread_load":"true","num_threads":64}' \
--tool-call-parser qwen25
2. Inference with MTP (Multi-Token Prediction)
The current official SGLang implementation of MTP contains a bug. For better inference performance, we recommend installing our patched version. Our fix is currently under review and is expected to be merged into the official SGLang library shortly.
Install our SGLang
git clone -b ling_2_6 [email protected]:antgroup/sglang.git
cd sglang
pip install --upgrade pip
pip install -e "python"
Start server
sglang serve \
--model-path inclusionAI/Ling-2.6-1T \
--tp-size 8 \
--max-running-requests 32 \
--mem-fraction-static 0.92 \
--chunked-prefill-size 8192 \
--context-length 262144 \
--trust-remote-code \
--speculative-algorithm EAGLE \
--speculative-num-steps 3 \
--speculative-eagle-topk 1 \
--speculative-num-draft-tokens 4 \
--mamba-scheduler-strategy extra_buffer \
--mamba-full-memory-ratio 1.4 \
--model-loader-extra-config '{"enable_multithread_load":"true","num_threads":64}' \
--tool-call-parser qwen25
Client
curl -s http://${MASTER_IP}:${PORT}/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "auto", "messages": [{"role": "user", "content": "What is the capital of France?"}]}'
More usage can be found here
pip install uv
uv venv ~/my_ling_env
source ~/my_ling_env/bin/activate
git clone https://github.com/vllm-project/vllm.git
cd vllm
VLLM_USE_PRECOMPILED=1 uv pip install --editable . --torch-backend=auto
Server
vllm serve $MODEL_PATH \
--port $PORT \
--served-model-name my_model \
--trust-remote-code --tensor-parallel-size 8 \
--gpu-memory-utilization 0.85
Client
curl -s http://${MASTER_IP}:${PORT}/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "auto", "messages": [{"role": "user", "content": "What is the capital of France?"}]}'
While Ling-2.6-1T excels in reasoning and agentic efficiency, our future development will focus on:
We remain committed to pushing the boundaries of model performance to enhance delivery efficiency across all complex scenarios.
This code repository is licensed under the MIT License.