HuggingFaceTB/SmolLM2-135M — model analysis

Summary

  • Size: The model consists of about 135 million numbers (parameters). Its saved files total about 270MB.
  • Structure: A dense model that uses all of its parameters for every input. About 60% of the parameters are assigned to the part that holds knowledge (FFN), about 20% to the vocabulary dictionary (embedding), and about 20% to the part that reads context (attention).
  • Japanese text: Japanese is split into about 1.8 units (tokens) per character. In English, one token holds about 4.1 characters. For the same number of characters, Japanese needs about 7.6 times as many tokens as English. Of its 49,152 vocabulary entries, 89 contain kana or kanji.
  • Context length: The limit is 8,192 tokens, roughly 4,500 Japanese characters or 34,000 English characters.
  • Skew in values: Values inside the model flow in groups of 576. Position 100 was among the positions where outliers (extremely large values) appear most often, in all 30 layers. Position 100 also appeared in the tally of positions where normalization-layer weights take extreme values (39 of 61 normalization layers). This kind of skew is generally considered a cause of accuracy loss when a model is compressed (quantized).

One line: Dense · llama · 30 layers · 134.52M params · GQA (3:1) · context 8,192 · vocab 49,152

Basic information

Architecture (stage 0)

Parameters by part

Counted from the tensor shapes in the safetensors headers (measured).

Parameters by part

Part Parameters Share
FFN (dense) 79,626,240 59.20%
Embedding 28,311,552 21.05%
Attention 26,542,080 19.73%
Normalization 35,136 0.03%

Structure

Values from config.json.

Item Value
model_type / architectures llama / LlamaForCausalLM
Layers 30
Hidden size 576
FFN intermediate size 1,536
Attention type GQA (3:1)
Heads / KV heads / head dim 9 / 3 / 64
Sliding window — (use_sliding_window=—)
Layer types (layer_types) —
RoPE θ 100,000
RoPE scaling —
Partial rotary (partial_rotary_factor) —
Context length (max_position_embeddings) 8,192
Tied input/output embeddings Yes
Activation silu
Normalization ε 1.000e-05

MoE configuration

Not applicable (not an MoE model).

config vs. actual tensors

Values in config.json side by side with the actual tensor shapes (measured).

Item config Actual Match Source of actual
vocab_size 49,152 49,152 Yes model.embed_tokens.weight
hidden_size 576 576 Yes model.embed_tokens.weight
num_hidden_layers 30 30 Yes largest layer index + 1 (excluding MTP layers)
num_key_value_heads×head_dim 192 192 Yes model.layers.0.self_attn.k_proj.weight
num_attention_heads×head_dim 576 576 Yes model.layers.0.self_attn.q_proj.weight
intermediate_size 1,536 1,536 Yes model.layers.0.mlp.gate_proj.weight
tie_word_embeddings (no output-head tensor) Yes Yes Yes tensor list

Tokenizer

Results of tokenizing the fixed Japanese and English texts (see conditions) without special tokens (measured). "Tokens splitting a character" is the share of tokens that decode to a broken character (U+FFFD) on their own.

Text Characters Tokens Chars/token Bytes/token Tokens splitting a character Round-trips exactly
Japanese 449 825 0.544 1.63 86.1% Yes
English 1,201 291 4.127 4.223 0.0% Yes

Of 49,152 vocabulary entries, 89 (0.18%) contain at least one kana or kanji character and 3 contain two or more (loaded with tokenizers).

Special tokens and chat format

17 special tokens. BOS: <|endoftext|>, EOS: <|endoftext|>. Chat template: none. Thinking tags: No, tool calls: No, FIM: No, image tokens: No (found by searching special-token names and the template text).

List of special tokens

<|endoftext|>, <|im_start|>, <|im_end|>, <repo_name>, <reponame>, <file_sep>, <filename>, <gh_stars>, <issue_start>, <issue_comment>, <issue_closed>, <jupyter_start>, <jupyter_text>, <jupyter_code>, <jupyter_output>, <jupyter_script>, <empty_output>

Weights (stage 1)

Analyzed 272 of 272 tensors (compute time 1 s). Each tensor was converted to fp32 for computation.

Trends across layers

RMS is the root mean square of the elements. max|w| ÷ RMS is how many times larger the largest element is than a typical one. Excess kurtosis measures tail heaviness and is 0 for a normal distribution (all measured).

RMS

max|w| ÷ RMS

Excess kurtosis

Matrices with the largest values: max|w| ÷ RMS 52.021 (model.layers.27.mlp.gate_proj.weight), excess kurtosis 25.005 (model.layers.0.self_attn.o_proj.weight), outlier share 0.111% (model.layers.17.self_attn.k_proj.weight).

Outliers and outlier channels

An outlier is an element exceeding 6 times the standard deviation of its matrix. For each matrix that takes the hidden dimension as input, the columns with the most outliers (top 16) were collected, and the number of layers in which each column appeared was counted (measured).

Hidden-dimension column Layers it appeared in (of 30)
100 30
260 25
446 22
371 21
261 20
8 19
507 18
109 16
230 16
247 16

Normalization-layer weights

For each of the 61 normalization layers, the dimensions farthest from the median (top 16) were collected, and the number of normalization layers in which each dimension appeared was counted (measured).

Dimension Normalization layers it appeared in
258 41
100 39
507 38
162 37
247 34
490 34
127 33
397 33
446 32
8 30

Dimensions shared with the top 10 outlier channels above: 8, 100, 247, 446, 507.

Effective rank

Stable rank = (Frobenius norm)² ÷ (largest singular value)², divided by the matrix's shorter side to scale it to 0–1. The largest singular value is approximated by power iteration (30 steps) (estimated).

Effective rank

Stable rank ÷ shorter side by part:

Part Matrices Min Median Max Matrix with the minimum
Attention 120 0.0068 0.1148 0.4742 model.layers.0.self_attn.q_proj.weight
FFN (dense) 90 0.0287 0.1162 0.3189 model.layers.28.mlp.up_proj.weight
Embedding 1 0.0034 0.0034 0.0034 model.embed_tokens.weight

Embedding rows

Distribution of the length (norm) of each vocabulary entry's embedding row, and the rows with the smallest norms (bottom 0.1%) (measured).

model.embed_tokens.weight (49,152 rows)

Min Bottom 0.1% Bottom 1% Median Top 1% Max
1.7569 1.9437 2.1338 3.1309 4.9269 7.4676

Rows with small norms (50 rows in the bottom 0.1%, the 20 smallest): 7937:' Michael'(1.7569), 2322:' John'(1.7763), 3214:' CO'(1.7876), 5259:' David'(1.8027), 5622:' IN'(1.8095), 8581:' Tom'(1.8238), 6356:' Robert'(1.8255), 5455:' Paul'(1.8283), 7288:' Peter'(1.8366), 18452:' Chris'(1.8477), 4967:' William'(1.8493), 5457:' Sam'(1.8533), 2974:' Car'(1.8572), 5037:' James'(1.8633), 1166:' Ar'(1.8738), 5370:' George'(1.8768), 5325:' Alex'(1.8811), 4098:' anti'(1.8843), 788:'In'(1.8936), 1053:' Pro'(1.8979)

A small norm does not mean that the token is unimportant or untrained.

Similarity between experts

Not applicable (not an MoE model).

Measurement conditions

  • mbscope 0.1.3 (result schema v1), Python 3.14.4, torch 2.13.0+cu130, safetensors 0.8.0, huggingface_hub 0.36.2, tokenizers 0.21.4, transformers 4.49.0
  • Estimation criteria: outlier_sigma=6.0, embed_low_pct=0.1, power_iters=30, sketch_dim=32, top_channels=16, seed=20261002
  • Input texts: ja: Natsume Sōseki, I Am a Cat, opening (Aozora Bunko 789_ruby_5639; source: Natsume Sōseki Zenshū 1, Chikuma Bunko), sha256 prefix 5d6b40f79b6a / en: Jane Austen, Pride and Prejudice, opening of Chapter 1 (Project Gutenberg eBook #1342), sha256 prefix 0a714a270435
  • Run: 2026-10-02T11:31:00+09:00 – 2026-10-02T11:31:13+09:00, computed on GPU NVIDIA GeForce RTX 3090 (CUDA 13.0, TF32 off) in fp32
  • The text of this record was generated with mbscope 0.2.0 (measured with 0.1.3)

Comments

…

Comments are public. Please don't include personal information. Inappropriate comments may be removed.

← All notes