Track · 2:31 · Liner note

Counting Tokens Like Sheep

Tokens are the chunks of text, often word fragments, that language models read and are billed by, so their count drives cost, speed and how much fits in a context window.

Language models do not read words. They read tokens, chunks of text that are often pieces of words, and every token in and out is counted for cost, latency and context space. Counting Tokens Like Sheep is the accounting habit that follows: know how many you spend, where, and whether they earned it.

Counts vary by tokenizer, language and even formatting, so estimate with the real tokenizer, not a word count. Long system prompts and pasted tables are the usual culprits. Trim whatever does not change the answer.