Agents-A1-4B-NVFP4

NVFP4 (W4A4, compressed-tensors) quantization of InternScience/Agents-A1-4B — a 4B agentic model on the Qwen3.5 dense architecture (Qwen3_5ForConditionalGeneration: hybrid DeltaNet linear attention + full attention every 4 layers, Qwen3-VL vision tower, 262K context).

  • Size: 4.21 GB (from ~8 GB bf16), incl. 0.24 GB bf16 MTP draft head
  • Bonus — grafted MTP: the source checkpoint declares mtp_num_hidden_layers: 1 but ships no mtp.* weights. This repo grafts the 15-tensor bf16 MTP draft head from the base model Qwen/Qwen3.5-4B (identical text-config dims), enabling speculative decoding: 63–66% draft acceptance, ~+20% single-stream decode (measured, see below)
  • Scheme: NVFP4 W4A4, group size 16, via llm-compressor 0.11.0 / compressed-tensors 0.16.0
  • Kept bf16: lm_head, vision tower (model.visual*), DeltaNet conv1d
  • Calibration: 32 samples × 8192 seq, neuralmagic/calibration (pure-CPU calibration — no GPU used in the bake)

Serve with vLLM

vllm serve sakamakismile/Agents-A1-4B-NVFP4

Quantization is auto-detected — no --quantization flag needed. Requires a GPU with FP4 support (SM120 Blackwell) for the NVFP4 kernels.

With speculative decoding (grafted MTP draft, ~+20% single-stream):

vllm serve sakamakismile/Agents-A1-4B-NVFP4 \
  --speculative-config '{"method":"qwen3_5_mtp","num_speculative_tokens":3}'

Measured throughput — single GPU (vLLM 0.22.0, RTX PRO 2000 Blackwell SM120, KV fp8, 512-tok decode)

Config single-stream (c1) aggregate 8-way (c8) vs base
base (no MTP) 73.6 t/s 492.9 t/s
MTP n=3 (grafted) 88.6 t/s 474.5 t/s +20.4% c1 / −3.7% c8

SpecDecoding metrics (vLLM): mean acceptance length ~2.9, per-position acceptance 0.83 / 0.66 / 0.48, avg draft acceptance 63–66% — remarkable for a draft head grafted from the pre-fine-tune base model across InternScience's 3-stage agentic distillation. As usual, MTP is a latency win (single-stream / low concurrency); at saturation the extra draft forward costs slightly more than it saves.

Speculative decoding is lossless — outputs are identical to base decoding.

Recipe

Baked (pure-CPU, llm-compressor) with the same validated recipe as Qwen3.6-27B-MTP-pi-tune-NVFP4 and ThinkingCap-Qwen3.6-27B-NVFP4 (same qwen3_5 architecture family).

Downloads last month
213
Safetensors
Model size
3B params
Tensor type
F32
·
BF16
·
F8_E4M3
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sakamakismile/Agents-A1-4B-NVFP4

Quantized
(10)
this model