Prompt
Why are my LLM costs jumping every month?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
Your LLM bill usually jumps for one of a few predictable reasons:
1) More usage than you realize
- More users, more requests, more retries.
- Background jobs, agents, or cron workflows calling models repeatedly.
- A feature quietly becoming popular.
2) Prompt length creeping up
- Chat history getting longer.
- More RAG context being stuffed into prompts.
- System prompts, tool specs, or templates getting expanded.
- “Small” prompt changes can add a lot of tokens at scale.
3) Output length is growing
- The model is generating longer answers than before.
- You may have loosened max tokens or changed instructions.
- Summaries, explanations, or code generation can spike output tokens fast.
4) Model mix changed
- A fallback to a more expensive model.
- Upgrading from a cheap model to a premium one.
- Accidentally routing more traffic to a high-end model.
5) Retries and failures
- Timeouts, rate limits, tool errors, or malformed outputs can trigger retries.
- If your app retries automatically, each failure multiplies cost.
6) Hidden “agent” loops
- Agents can call the LLM many times per user request.
- Tool-use loops, reflection steps, and planning can inflate costs quickly.
- One “task” may actually be 5–20 model calls.
7) Retrieval costs are growing
- More documents retrieved per query.
- Larger chunks or too many chunks per prompt.
- Re-ranking or embedding updates adding extra spend.
8) More expensive tokens than expected
- Different providers/models price input and output differently.
- Long outputs are often much more expensive than inputs.
- Some models charge more for reasoning tokens or tool calls.
9) Lack of caching
- The same prompts are being recomputed over and over.
- No response cache, embedding cache, or prompt-template cache.
10) Inefficient prompt design
- Repeating instructions in every request.
- Sending full conversation history instead of a summary.
- Including irrelevant documents or tool schemas.
What to check first
Look at these metrics month over month:
- total requests
- average input tokens
- average output tokens
- retries per request
- model distribution
- number of tool calls per request
- average retrieved context size
Fast ways to reduce cost
- Cap output tokens.
- Summarize old chat history.
- Retrieve fewer, better chunks.
- Cache repeated responses.
- Use cheaper models for routine tasks.
- Add retry limits and backoff.
- Log token usage per feature, not just overall.
If you want, I can help you diagnose the most likely cause from your setup by looking at:
- your model/provider,
- your architecture,
- a sample of recent usage logs or token counts.