Downloads · 30 days
829
56% of all-time downloads
ArchSpace-Collection/NCP_ArchPreview_dolma3_8.9B_Stage1
NCP_ArchPreview_dolma3_8.9B_Stage1 is a text generation model from ArchSpace-Collection. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Model collection | Technical report (arXiv) | HF Papers | Training code (coming soon) | Evaluation code
Downloads · 30 days
829
56% of all-time downloads
All-time downloads
1.5K
Public
Parameters
8.9B
250 GB on disk
Likes
9
Public
Click a slice to open those files.
.safetensors17.9 GB · 100%
From the Hugging Face model README
Model collection | Technical report (arXiv) | HF Papers | Training code (coming soon) | Evaluation code
NCP-ArchPreview is a latent-space autoregressive language model developed by The NCP Team at Shanghai AI Lab and LUMIA Lab, Shanghai Jiao Tong University. It learns to predict both the next token and the next concept: a representation spanning a short group of tokens in a learned latent space. Concept predictions guide the token decoder, while generation retains the standard next-token interface.
This is the Stage 1 base-model release, following large-scale pretraining on Dolma 3 Mix. The architecture follows the OLMo 3 7B token-level design and adds a Concept Module, a product-quantized concept vocabulary, and hierarchical residual connections, bringing the total parameter count to approximately 8.94B.
The convergence comparison measures tokens required to reach a reference loss; it does not measure wall-clock training speed or inference throughput.
NCP-ArchPreview processes text through three modules:
Intra-module residual connections mix states across depths. Cross-module residual connections connect Encoder to Concept Module, Encoder to Decoder, and Concept Module to Decoder. Concept feedback is shifted and repeated at token resolution to preserve causality.
tokens -> Token Encoder -> mean pooling -> Concept Module -> concept prediction
| |
+--------------> Token Decoder <-----------------+
|
next-token logits
| Property | Configuration |
|---|---|
| Hugging Face architecture | NCPOlmo3ForCausalLM |
| Total parameters | Approximately 8.94B |
| Encoder / Concept Module / Decoder | 16 / 8 / 16 causal Transformer layers |
| Hidden size | 4,096 |
| FFN intermediate size | 11,008 |
| Attention heads / KV groups | 32 / 32 |
| Attention head dimension | 128 |
| Vocabulary size | 100,278 |
| Maximum training context | 8,192 tokens |
| Token-level attention | 4,096-token local window; full attention every fourth layer |
| Position encoding | RoPE, base 500,000 |
| Activation / normalization | SwiGLU / RMSNorm; layer-wise QK RMSNorm |
| Concept compression | 4 token states per concept |
| Product quantization | 32 codebooks, each with 128 codewords of dimension 128 |
| Parameter precision | BF16 |
Stage 1 uses Dolma 3 Mix and the staged pretraining framework described in the report. The reported 5.73T-token budget describes the large-scale training run; intermediate checkpoints have consumed only the tokens preceding their saved step.
The model is optimized jointly with three objectives:
The report uses Moonlight Muon for matrix-valued parameters and AdamW for
embeddings, biases, and other non-Muon parameters. Its default learning rate
is 6e-5, with the OLMo-3-style cosine schedule.
Results below describe the report's final Stage 1 model. Scores are percentages and higher is better; deltas are absolute percentage points.
| Metric | OLMo-3-7B Stage 1 | NCP-ArchPreview Stage 1 | Delta |
|---|---|---|---|
| Overall AVG | 46.59 | 49.04 | +2.45 |
| MMLU | 62.22 | 64.80 | +2.58 |
| GSM8K | 39.27 | 45.26 | +5.99 |
| MATH-500 | 12.52 | 14.48 | +1.96 |
| HumanEval | 27.10 | 31.38 | +4.28 |
| MBPP | 34.53 | 35.91 | +1.38 |
| ARC-Challenge | 77.99 | 81.57 | +3.58 |
| PIQA | 72.25 | 80.85 | +8.60 |
| Domain average | OLMo-3-7B Stage 1 | NCP-ArchPreview Stage 1 |
|---|---|---|
| MMLU family | 54.50 | 56.73 |
| Mathematics | 20.79 | 24.54 |
| Code | 25.15 | 27.79 |
| Multiple-choice STEM | 84.47 | 86.93 |
| Multiple-choice non-STEM | 70.08 | 74.71 |
| GenQA | 54.29 | 54.76 |
Likelihood is reported separately in bits per UTF-8 byte (BPB), where lower is better.
| Likelihood metric | OLMo-3-7B Stage 1 | NCP-ArchPreview Stage 1 |
|---|---|---|
| BPB AVG | 0.824 | 0.811 |
The complete per-benchmark results are available in Table 1.
</details>Overall AVG is the unweighted mean of the 26 constituent benchmark scores in Table 1, excluding the aggregate MMLU row, domain averages, and BPB results. BPB AVG is computed separately over ten likelihood benchmarks.
Detailed evaluation settings can be found in our technical report.
The final Stage 1 model is available as
ArchSpace-Collection/NCP_ArchPreview_dolma3_8.9B_Stage1.
The report also evaluates intermediate checkpoints from 100K to 1.3M training
steps. Selected measurements from different training steps are shown below.
| Training checkpoint | MMLU | GSM8K | HumanEval | Overall AVG |
|---|---|---|---|---|
| 100K | 53.08 | 23.43 | 20.05 | 39.64 |
| 300K | 59.42 | 32.75 | 27.02 | 44.27 |
| 600K | 61.27 | 40.49 | 26.14 | 46.06 |
| 900K | 63.34 | 42.15 | 29.55 | 47.85 |
| 1.3M | 64.77 | 45.34 | 29.40 | 48.87 |
| Final | 64.80 | 45.26 | 31.38 | 49.04 |
Use the final checkpoint for the main Stage 1 comparison and intermediate checkpoints to study training dynamics. Individual benchmark scores can fluctuate between checkpoints. When using this card for an intermediate release, its own training step identifies the applicable row; the final checkpoint's scores are reference results only.
The checkpoint includes custom Transformers model code. This example uses a single prompt on one CUDA GPU with sufficient memory for the BF16 weights, activations, and cache.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "ArchSpace-Collection/NCP_ArchPreview_dolma3_8.9B_Stage1"
device = "cuda"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
torch_dtype=torch.bfloat16,
).eval().to(device)
prompt = "The role of hierarchical representations in language modeling is"
inputs = tokenizer(prompt, return_tensors="pt").to(device)
with torch.inference_mode():
output_ids = model.generate(
**inputs,
max_new_tokens=128,
do_sample=False,
use_cache=True,
pad_token_id=tokenizer.eos_token_id,
)
continuation = output_ids[0, inputs["input_ids"].shape[1]:]
print(tokenizer.decode(continuation, skip_special_tokens=True))
Use plain completion or few-shot prompts for this base model. The example is a
loading and generation example; benchmark reproduction requires the prompts,
sampling configuration, and scorers described in the report. For reproducible
runs, pin the model and tokenizer to the same Hub commit with revision.
See the inference guide for supported runtime
versions and optimized serving.
This release supports research on latent-space language modeling, evaluation of pretrained capabilities, analysis of training dynamics, continued pretraining, and concept-based adaptation.
@misc{ncpteam2026ncparchpreviewtechnicalreportmoving,
title={NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction},
author={NCP Team and Jiaqi Cao and Chiyu Chen and Shuang Cheng and Xu Cheng and Beiya Dai and Yufan Feng and Kewen Ge and Ruijun Ge and Jiayi Huang and Yang Jiao and Dahua Lin and Zhouhan Lin and Yifan Liu and Yuliang Liu and Biqing Qi and Mowen Ruan and Junzhe Shen and Yunchong Song and Hao Sun and Zhongbo Tian and Yixuan Wang and Rubin Wei and Jiaxin Xiong and Kangyu Yang and Qian Yao and Qi Zhang and Bowen Zhou},
year={2026},
eprint={2609.10715},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2609.10715},
}
The model weights are released under the Apache License 2.0.
We thank the OLMo and Dolma teams and the contributors to the open datasets, training libraries, and evaluation tools used in this work.
<!-- Editorial source: Concept_Olmo_7B.pdf, report dated 2026-09-04, SHA-256 acf619d01c0ad46e4012b9dacf1af7c2b21501a510008e6dad57b035fd1eaf9e. Architecture: Sections 2-3 and Table 10. Main evaluation: Table 1. Checkpoint trajectory: Table 14. Evaluation protocol: Appendix E / Table 13. Keep these tables separate from the frozen results in Section 5.1 / Appendix C. -->