The article demonstrates how to quantify LLM prompt expenses by using the tiktoken library with the cl100k_base encoder to obtain exact token counts for three example strings: a plain‑English sentence (31 tokens, 1.15 tokens/word), a technical paragraph (37 tokens, 1.76 tokens/word), and a mixed‑code snippet (43 tokens, 1.59 tokens/word). It then converts token usage into cost using a pricing table that lists Claude Sonnet at $3 per 1M input tokens and $15 per 1M output tokens, and Claude Haiku at $0.80 per 1M input and $4 per 1M output. For a 87‑token system prompt sent 1,000 times daily, the daily expense is $0.26 with Sonnet ($95 per year) and $0.07 with Haiku ($25 per year). A condensed version of the prompt reduces the token count to 47 (46 % fewer), cutting the Sonnet yearly cost to $51. To illustrate temperature’s effect, the notebook simulates sampling from a five‑word distribution; at temperature 0.2 the model selects “sunny” eight times in eight draws, whereas at temperature 1.0 the same logits yield a varied set of outputs, showing that temperature alters sampling sharpness without changing underlying knowledge. The piece wraps the logic into a reusable tokenize_and_cost function and suggests Colab experiments such as scaling query volume, adding GPT‑4o‑mini pricing, and testing extreme temperature values.
Read original
dev.to