A 16,384-token vocabulary trained from scratch on 645,388 US court opinions, SEC filings and educational web text. Type below to see exactly how the model would read it.
Trained with a ByteLevel pre-tokenizer and
add_prefix_space=false, so every byte sequence is representable — no
out-of-vocabulary token exists, and encoding is fully reversible. Legal citations,
monetary amounts and statutory phrasing compress far better here than in a
general-purpose vocabulary because the merges were learned from this corpus.
75% legal & financial · 1.87B training tokens
| Vocabulary size | 16,384 |
| Context window | 1,024 |
| Training tokens | 1.85B |
| Validation tokens | 18.7M |
| Documents | 645,388 |
| Merge rules | 16,121 |
| Reserved tokens | 7 |
| Model parameters | 125.8M |