Downloads · 30 days
294
53% of all-time downloads
internlm/Intern-S2-Mobius-FP8
Intern-S2-Mobius-FP8 is a image-text-to-text model from internlm. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
<div align="center" <img src="./figs/title.png" /
Downloads · 30 days
294
53% of all-time downloads
All-time downloads
554
Public
Parameters
36B
103 GB on disk
Likes
9
Public
Click a slice to open those files.
.safetensors38.5 GB · 100%
How the weights are stored.
F8_E4M334.5B · 96%
From the Hugging Face model README
💻Github Repo • 🤗Model Collections • 🌳Arch Space
</div>We introduce Intern-S2-Mobius, a 35B foundation model built on the Mobius-v0 architecture realized by Xtuner and LMDeploy. Instead of binding knowledge storage and reasoning computation layer by layer as in conventional Transformer models, Mobius organizes knowledge into a globally shared Memory and lets multiple Reasoners iteratively query and refine hidden states against this shared repository.
This knowledge-reasoning separation gives Intern-S2-Mobius two native capabilities: Backward Residual Connection, where reasoning stages can access knowledge beyond their local layer hierarchy, and Dynamic Latent Reasoning, where deliberation, refinement, and multi-token prediction are internalized into high-density continuous states. Continual-pretrained from Qwen3.5-35B and further post-trained with SFT and RL, Intern-S2-Mobius preserves strong downstream capability while achieving substantially higher end-to-end inference efficiency, with nearly 4x speedup reported in the technical report.
Knowledge-reasoning decoupled architecture. Intern-S2-Mobius separates knowledge vectors from reasoning operators by replacing layer-bound FFN knowledge storage with a globally shared Memory. This gives each Reasoner access to a broader knowledge space and improves knowledge compression compared with a standard Transformer layout.
Backward Residual Connection. Through shared Memory, shallow and deep reasoning stages can access knowledge across the model rather than relying only on forward layer-wise information flow. This enables more flexible cross-layer knowledge composition and helps the model synthesize useful information in fewer reasoning steps.
Dynamic Latent Reasoning. Mobius refines continuous hidden states through recurrent latent iteration before decoding. This internalizes part of the deliberation process, reduces reliance on long visible chain-of-thought, and dynamically allocates computation to different tokens.
Higher inference efficiency with concise reasoning. On reasoning benchmarks, Intern-S2-Mobius reaches comparable or stronger scores than the Qwen3.5-35B baseline while producing markedly shorter reasoning traces and higher request throughput, leading to nearly 4x end-to-end inference speedup in the reported evaluation.
Strong general and scientific performance. Intern-S2-Mobius improves the reported average score over Qwen3.5-35B on general reasoning benchmarks, and shows large gains on scientific tasks such as Biology-Instructions, Mol-Instructions, and MolecularIQ.
We evaluate the Intern-S2-Mobius on various benchmarks, including general datasets and scientific datasets. We report the performance comparison with Qwen3.5-35B below. We use the OpenCompass to evaluate all models. For text benchmarks, Intern-S2-Mobius is evaluated with a maximum inference length of 64K tokens on MMLU Pro, SimpleQA, and HLE, and 128K tokens on the remaining text benchmarks.
<figure> <img src="./figs/performance.png" alt="performance"> <figcaption>Fig3: Performance comparison across general and scientific benchmarks. The higher score in each row is highlighted in <strong>bold</strong>.</figcaption> </figure> <figure> <img src="./figs/case-study.png" alt="case study"> <figcaption>Fig4: Step-aligned comparison between Intern-S2-Mobius-35B and Qwen3.5-35B on a linear-algebra multiple-choice question. Both models select the correct answer (Option C). Token counts are computed using the Qwen3.5-35B tokenizer. Mobius completes the same reasoning steps with fewer tokens, which mainly benefits from the model's elimination of repeated derivation and checks.</figcaption> </figure>The Intern-S2-Mobius release is a 35B model stored in bfloat16 weight format. This guide provides deployment examples for the following configurations:
NOTE: The commands below are reference configurations. Inference frameworks are under active development, so use the latest framework documentation and your local validation results when tuning production deployments.
Intern-S2-Mobius can be deployed using any of the following LLM inference frameworks:
We recommend using the following hyperparameters to ensure better results
top_p = 1
top_k = 50
min_p = 0.0
temperature = 0.8
Use the latest LMDeploy with Intern-S2-Mobius support. The examples below use single-GPU serving.
lmdeploy serve api_server \
internlm/Intern-S2-Mobius \
--trust-remote-code \
--backend pytorch \
--tp 1 \
--speculative-algorithm qwen3_5_mtp \
--speculative-num-draft-tokens 4 \
--dtype bfloat16 \
--max-batch-size 64
lmdeploy serve api_server \
internlm/Intern-S2-Mobius \
--trust-remote-code \
--backend pytorch \
--dtype bfloat16 \
--tp 1
Use a recent Transformers version with remote-code loading enabled.
import torch
from transformers import AutoModelForImageTextToText, AutoTokenizer
model_path = "internlm/Intern-S2-Mobius"
tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
model = AutoModelForImageTextToText.from_pretrained(
model_path,
trust_remote_code=True,
torch_dtype=torch.bfloat16,
device_map="auto",
).eval()
messages = [
{"role": "user", "content": "Give me a short introduction to Intern-S2-Mobius."}
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
with torch.no_grad():
output_ids = model.generate(
**inputs,
max_new_tokens=512,
do_sample=True,
temperature=0.8,
top_p=1,
)
response_ids = output_ids[0][inputs["input_ids"].shape[-1]:]
print(tokenizer.decode(response_ids, skip_special_tokens=True))
Use the latest vLLM Docker image or source build with Intern-S2-Mobius support.
vllm serve \
internlm/Intern-S2-Mobius \
--trust-remote-code \
--tensor-parallel-size 2 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--spec-method mtp \
--spec-tokens 4
vllm serve \
internlm/Intern-S2-Mobius \
--trust-remote-code \
--tensor-parallel-size 2 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder
Use the lmsysorg/sglang:dev Docker image or a recent source build with Intern-S2-Mobius support. See the SGLang cookbook for Docker commands, verified deployment recipes, benchmarks, and usage examples.
sglang serve \
--model-path internlm/Intern-S2-Mobius \
--trust-remote-code \
--tp 2 \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder \
--speculative-algorithm NEXTN \
--speculative-num-steps 3 \
--speculative-eagle-topk 1 \
--speculative-num-draft-tokens 4
sglang serve \
--model-path internlm/Intern-S2-Mobius \
--trust-remote-code \
--tp 2 \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder