slm-125m · byte-level BPE · runs entirely in your browser

Legal & financial tokenizer

A 16,384-token vocabulary trained from scratch on 645,388 US court opinions, SEC filings and educational web text. Type below to see exactly how the model would read it.

Loading vocabulary…
Tokens
0
 
Characters
0
 
Chars / token
0.00
higher is better
Tokens / word
0.00
compression ratio
Context used
0%
of 1,024

Token stream — click any token to copy its ID · · marks a leading space

Token IDs

What this vocabulary learned

Trained with a ByteLevel pre-tokenizer and add_prefix_space=false, so every byte sequence is representable — no out-of-vocabulary token exists, and encoding is fully reversible. Legal citations, monetary amounts and statutory phrasing compress far better here than in a general-purpose vocabulary because the merges were learned from this corpus.

case law 31% SEC 44% web 25%

75% legal & financial · 1.87B training tokens

Vocabulary size16,384
Context window1,024
Training tokens1.85B
Validation tokens18.7M
Documents645,388
Merge rules16,121
Reserved tokens7
Model parameters125.8M