Measure your LLM total cost of ownership first, then implement model routing, caching, and prompt compression in that order. Together, these three levers deliver 50 to 80% savings on most production workloads, without touching output quality, as long as you protect the change with a small evaluation set and cost dashboards.
TL;DR:
- Model routing can reduce costs by 60 to 80 percent at scale, while prompt compression and caching provide immediate savings with minimal effort.
- Building a detailed total cost of ownership model reveals that API token costs often account for less than 10 percent of total expenses at moderate volume, with engineering costs dominating.
- Proper calibration of caching strategies can eliminate up to 60 percent of model calls, but only if the hit rate is accurately measured against real prompt diversity.
- For workloads under 10,000 requests daily, optimizing prompt compression and caching yields the best ROI, whereas routing becomes critical at 100,000 requests or more daily.
- Infrastructure self-hosting only becomes cost-effective at very high, steady volumes, as managing your own serving stack incurs significant time and resource costs.
Table of Contents
- The seven levers that actually move your bill
- Measure first: build a TCO model before you touch anything
- Token-level optimizations: compression, chunking, and output limits
- Caching and batching: real hit rates, not brochure numbers
- Model routing: send easy queries to cheap models, save the rest for hard ones
- Serving and infrastructure: when self-hosting actually pays off
- Product and integration choices that change your cost baseline
- Monitoring: the dashboards and alerts that catch drift before it costs you
- A practitioner's checklist for a two-week cost sprint
- When cutting tokens costs you more than it saves
- How I help: fixed-price audits and two to four week builds
- Sources
- FAQ
The seven levers that actually move your bill
Not every optimization deserves the same attention. Some levers pay for themselves in a day, others take a sprint, and a few are only worth it once you are past a certain scale. Here is the order I work through with clients, ranked by time-to-value.
- Measurement and TCO modeling: half a day of work, no direct savings yet, but it tells you where the other six levers should point.
- Model routing: one to three days to build a first version, commonly the single largest reduction at 60 to 80% savings once complexity tiers are in place.
- Prompt compression: one to two days, typically 30 to 60% fewer input tokens with no measurable quality drop when tested against a labeled set.
- Caching (exact and semantic): two to four days to build and calibrate, can eliminate 30 to 60% of calls for workloads with repeated or similar prompts.
- Batching async workloads: one day if your provider has a batch endpoint, roughly 50% cost reduction on jobs that do not need real-time responses.
- Evaluation harness: two days once, then it runs in the background to catch regressions before they cost you in rework.
- Infrastructure and serving: only relevant once you are running your own models. Weeks of work, but decisive at high volume.
The right starting point depends on volume. At 10,000 requests a day, spend your time on prompt compression and caching, since infrastructure changes will not move the needle yet. At 100,000 a day, routing becomes the priority because the spread between cheap and expensive models compounds fast. At 1 million or more a day, you need all seven levers working together, and infrastructure choices start to dominate the conversation.
Measure first: build a TCO model before you touch anything
Every optimization sprint I have run starts the same way: instrument before you manage. Guessing which lever matters most wastes more money than the inefficiency you are trying to fix.
Build a per-task cost formula that includes everything, not just the API line item:
- API token cost: input tokens times input price, output tokens times output price, summed per request.
- Infrastructure cost: any self-hosted compute, storage, or networking tied to serving that request.
- Engineering amortization: the hours spent building and maintaining the feature, spread across its expected lifetime.
A TCO breakdown of AI integration costs found that at moderate volume, around 10,000 requests a day, API costs can be under 10% of total cost. Engineering and integration work dominates instead. That single fact changes what you optimize first: a team obsessing over token counts at that volume is often optimizing the smallest line item on the bill.
Once you have the formula, pull usage broken out by model, endpoint, and feature. Most provider dashboards let you tag requests, and if yours does not, add a request ID and log it yourself. Traffic multiplies every lever, and modeling the spike before it happens is the difference between a planned migration and an emergency one.
A statistic worth sitting with: at moderate volume, API costs can be under 10% of total LLM operational cost, with integration and engineering work making up the rest. If your cost-cutting plan only touches the API bill, you are chasing a fraction of the problem.
With the formula and usage breakdown in hand, run four to six configuration benchmarks: your current setup, a cheaper model on easy tasks, compressed prompts, and a cached version of each. Plot cost against quality for each configuration and you get a Pareto frontier: the set of configurations where you cannot improve cost without sacrificing quality, or vice versa. Everything to the right of that frontier is waste. This is the single most useful chart I build for a client in the first week, because it turns "cut costs" from a vague goal into a specific, defensible list of changes.

Token-level optimizations: compression, chunking, and output limits
Once you know where the money goes, token-level work is usually the fastest lever to pull. It touches prompts and context, not infrastructure, so you can ship changes in a day and measure the effect the same afternoon.
Compress system prompts and instructions. Long system prompts accumulate over months of patching: an edge case here, a clarifying sentence there. Most of that text is redundant once you look at it as a whole. Rewrite it for density, remove repeated instructions, and replace verbose examples with terse ones. Done systematically against a labeled validation set, prompt compression cuts input tokens by 30 to 60% with no measurable quality loss. The validation set matters here: compress blind and you risk cutting the one clause that prevents a bad output.
Get chunking right for RAG. Chunk size and overlap decide how much irrelevant context rides along with every retrieval call, and irrelevant context is pure token waste.
- Aim for chunks between 300 and 800 tokens for most text-heavy documents, and size down for structured or code-heavy content.
- Use 10 to 20% overlap between chunks so you do not sever a sentence that answers the query.
- Attach metadata (source, section, date) to each chunk so the retriever can filter before it ever touches the LLM.
- Refresh embeddings on a cadence tied to how often your source documents change, not on a fixed calendar default.
Control output length deliberately. Output tokens are the expensive half of the ledger. On most frontier models, output pricing runs 3 to 5 times higher per token than input, so a verbose response costs far more than a verbose prompt. Set max_tokens to the actual ceiling your feature needs, not a generous default. Use stop sequences to end generation the moment the answer is complete. Tighten sampling parameters (lower temperature, top-p) for tasks where consistency matters more than creative range. Fine-tuning is worth considering only once you have a stable, high-volume task where prompt-based control has plateaued: it is a bigger investment, and it only pays off at scale.
Pro Tip: Before compressing any prompt, save the original as a labeled test case. Every compression pass should be a diff you can revert, not a rewrite from scratch.
For a deeper walkthrough of filtering and structuring context for RAG and agent pipelines, I cover the mechanics in context management for solo engineers.
Caching and batching: real hit rates, not brochure numbers
Caching is the lever most likely to disappoint if you skip the calibration step. Build it in two layers.
Exact-match caching first. Hash the full prompt (or the parts that matter) and check for a hit before calling the model. This catches repeated queries verbatim: the same support question, the same report request, the same onboarding step. It costs almost nothing to build and never returns a wrong answer, since it only fires on an identical input.
Semantic caching second. Embed the incoming prompt, search for similar past prompts within a distance threshold, and return the cached response if the match is close enough. This catches paraphrased or reworded versions of the same request, which exact-match caching misses entirely.
- Calibrate your similarity threshold against a labeled set of pairs you mark as "should match" and "should not match", not a default value pulled from documentation.
- Estimate your realistic hit rate from actual prompt diversity in your logs, since a workload with highly varied prompts will cache far less than a support bot answering the same twenty questions.
- Semantic caching can eliminate 30 to 60% of calls for workloads with concentrated prompt prefixes, but that range depends entirely on your traffic shape.
- Do not assume a vendor's advertised hit rate applies to you: cache sizing depends on prompt diversity and the provider's caching window, and the only way to know your number is to measure real traffic.
Batching handles the async half of your workload. Any job that does not need a response in real time, nightly summarization, bulk classification, report generation, is a batching candidate. Group requests, submit them through a batch endpoint where your provider offers one, and let the system amortize fixed costs like the system prompt tokens across the batch. This pattern commonly reduces per-item cost by roughly 50% for async jobs.
One figure that changes the calculus: batch endpoints and prompt amortization typically deliver about 50% cost savings on jobs that can tolerate delay. If a quarter of your workload is async and you have not moved it to batch, that is the fastest remaining win on the table.
Keep batch sizes conservative at first (50 to 100 items), add retry logic for partial failures, and monitor for timeout patterns before scaling batch size up. A batch job that fails at item 400 out of 1,000 should not force a full re-run.
I go deeper on cache design and concrete token-savings examples in prompt caching for engineers.

Model routing: send easy queries to cheap models, save the rest for hard ones
Routing is usually the single biggest lever available, and it is also the one teams delay longest because it sounds harder than it is. In practice, most routing systems start simple and get more sophisticated only when the simple version stops being accurate enough.
Start with heuristic routing. Classify incoming requests by feature type, input length, or a handful of keyword rules, and send anything that fits a known-easy pattern (short factual lookups, structured extraction, simple classification) to a smaller, cheaper model. Reserve the frontier model for anything that falls outside those patterns. This gets you most of the benefit with a day or two of engineering.
Move to a learned predictor once heuristics plateau. Train a lightweight classifier on a sample of past requests, labeled by whether the small model's answer matched the frontier model's answer. This predictor decides, per request, whether the cheap model is likely to succeed. It is more work to build than heuristics, but it adapts as your traffic mix shifts.
Build complexity tiers deliberately, not by guesswork.
- Define two or three tiers (simple, moderate, complex) based on what actually predicts difficulty for your workload: input length, number of steps required, ambiguity of the request.
- Sample real production traffic to calibrate tier boundaries, rather than copying tier definitions from someone else's workload.
- Route a small percentage of "simple" traffic to the frontier model on an ongoing basis, purely to catch cases where your tiering is wrong.
Routing is the lever with the largest reported upside: complexity-based routing reduces overall cost by 60 to 80% for many production systems, with reported quality loss under 2%. That gap between cost and quality impact is why routing sits above prompt compression on most priority lists.
Consider hardware signals if you are running your own serving stack. Static routers that ignore GPU load can send traffic to an already-saturated node, causing SLO violations even though the model choice itself was correct. Incorporating hardware signals into routing decisions improves latency and utilization by avoiding that kind of queue skew, which matters once you are self-hosting rather than calling a hosted API. If you want to see how your current model mix compares across providers before committing to a routing scheme, a tool like BabyLoveGrowth's multi-LLM audit gives you a quick side-by-side.
Serving and infrastructure: when self-hosting actually pays off
Infrastructure is the lever with the highest ceiling and the highest barrier to entry. Skip this section entirely if you are calling hosted APIs and staying under a few hundred thousand requests a day: it is not yet where your money is.
If you do run your own models, KV cache memory sizing decides both your latency and your GPU bill. A larger KV cache lets you serve longer contexts and more concurrent requests per GPU, but it eats memory that could otherwise run a second model instance. GPU choice follows from that tradeoff: an H100 gives you more memory bandwidth and larger effective batch sizes, which lowers per-request cost at high throughput, but it is a poor fit for low, spiky traffic where the GPU sits idle between requests.
Autoscaling has to account for cold start time, since spinning up a new model replica can take long enough that a traffic spike causes a wave of failed requests before capacity catches up. Keep a small warm pool sized to your typical spike, and monitor per-GPU utilization directly rather than trusting request-count metrics alone, since utilization tells you whether you are paying for idle capacity.
The honest answer to "hosted API or self-hosted" is that hosted APIs win for most teams below a certain steady volume, because you are not paying for idle GPU time and you are not the one debugging a serving stack at 2 AM. Self-hosting becomes cost-effective once your volume is high and steady enough that GPU utilization stays consistently high across the day. Below that line, the engineering time spent managing your own serving infrastructure usually costs more than the API savings you are chasing.
Product and integration choices that change your cost baseline
Some of the biggest cost decisions get made before a single prompt is written, when the feature itself is being scoped.
Workload shape decides which model is cheapest. A feature that takes a long input and returns a short output (classification, extraction, tagging) has a very different cost profile than one that takes a short input and returns a long output (drafting, summarizing at length). Since output tokens run 3 to 5 times more expensive than input tokens on most frontier pricing, a feature that generates long-form output deserves more scrutiny on model choice than one that mostly classifies short inputs.
Teams that budget for the build but not the year after it are the ones surprised by a creeping bill twelve months later.
Roll out changes behind a flag, not a big-bang switch. Run a new model, prompt, or routing rule in shadow mode first: log what it would have returned, without acting on it, and compare against your current production path on real traffic. Once shadow mode looks stable, roll it out behind a feature flag to a small percentage of traffic, then expand. This is how you collect accuracy data on the new configuration without betting production quality on an untested change.
Monitoring: the dashboards and alerts that catch drift before it costs you
None of the levers above stay optimized on their own. Prompt lengths creep back up, traffic mix shifts, a new feature quietly starts calling the expensive model by default. Monitoring is what catches that before the invoice does.
Track a small, specific set of metrics rather than everything your provider exposes:
- Cost by model and by feature, so you can see exactly which part of the product is driving spend, not just a single blended number.
- Retry rate and output length, since both tend to creep upward silently and both directly inflate token cost.
- Cache hit rate, tracked over time rather than as a one-time measurement, since traffic shape drifts.
- Cost per successful task, which is the metric that actually matters. A cheaper model that fails more often on the first attempt often costs more overall once you count the reprompting and rework it triggers.
Common failure modes worth alerting on specifically include a sudden spike in retries (often a sign of a prompt regression), an unexpected shift toward the expensive model in your routing logs, and output length drifting upward on a feature that used to be short and consistent.
Alongside dashboards, keep a small labeled validation set, 50 to 200 examples is often enough, that you run before every prompt, model, or routing change ships. This catches quality regressions before they reach production traffic, which matters more than any single cost metric, since a quality regression that triggers downstream rework can cost more than the tokens it saved.
I cover dashboard and alerting setup in more detail in LLM monitoring for solo senior engineers, and the validation set methodology in evaluation metrics for production LLMs.
A practitioner's checklist for a two-week cost sprint
Here is the sequence I actually run when a client's LLM bill needs to come down without breaking what already works.
Start by instrumenting cost by model, endpoint, and feature before changing anything. Pilot the highest-ROI change on a single use case rather than rolling out five changes at once, since you need to know which one worked. When a prototype's token usage has grown unmanaged for months, the fix is often a fixed-scope refactor of the prompt and context pipeline, not another incremental patch.
A two-week sprint checklist that covers most of what matters:
- Days 1 to 2: instrument cost by model, feature, and endpoint; build the TCO formula.
- Days 3 to 5: run the Pareto benchmark across four to six configurations.
- Days 6 to 8: build a caching prototype (exact-match first, semantic second) and calibrate hit rate against real traffic.
- Days 9 to 11: build a first-pass heuristic router for the clearest complexity tiers.
- Days 12 to 14: wire up cost dashboards, alert thresholds, and the labeled validation set.
The implementation details for caching and monitoring specifically are covered in prompt caching for engineers and LLM monitoring for solo senior engineers, both written as step-by-step guides rather than overviews.
When cutting tokens costs you more than it saves
The failure mode I see most often is a team that treats token count as the only number that matters. Cut context too aggressively and you strip out the sentence that prevented ambiguity, and now the model asks a clarifying question, or worse, guesses. Removing needed context without validating against real cases increases retries and downstream human review, and that rework often costs more than the tokens you saved.
The rule of thumb: optimize for cost per successful task, not cost per token. If you have run the TCO model above and the numbers still do not make sense, that is the point to bring in outside help rather than keep iterating alone.
— Hanad Kubat
How I help: fixed-price audits and two to four week builds
Everything above works if you have the time to build the instrumentation, run the Pareto tests, and calibrate the caching layer yourself. Most founders and small teams do not, and the prototype is the spec: whatever is running in production right now already tells us where the token churn is coming from.
I offer a fixed-price Prototype Audit that includes cost instrumentation, a prioritized plan ranked by ROI, and a clear read on which of the seven levers matters most for your workload. If you want the fix built, not just diagnosed, I deliver routing, caching, or prompt changes as a fixed-scope engagement, with full handoff documentation. Details and pricing are on my landing page, where you can see current offers and get in touch.
Sources
- Hybrid ML‑LLM: token-level optimizations
- LLM cost optimization guide
- Sonar: LLM cost optimization overview
- HW-Router: hardware-aware routing for LLM serving
FAQ
What is the fastest way to reduce LLM costs this week?
Start with prompt compression and exact-match caching, since both can ship in a day or two and require no infrastructure changes. Compression can cut input tokens by 30 to 60% when tested against a labeled validation set, and caching removes repeated calls entirely.
Does model routing hurt output quality?
Routing done with proper complexity tiers keeps quality loss low. Reported cases show complexity-based routing cutting cost by 60 to 80% with quality loss under 2%, as long as the tiers are calibrated on real traffic and a small share of "simple" requests still gets checked against the frontier model.
How much of my LLM bill is actually the API cost versus engineering?
It depends heavily on volume. At moderate scale, around 10,000 requests a day, API costs can be under 10% of total cost, with engineering, integration, and QA making up the rest, so a cost-cutting plan focused only on tokens misses most of the bill.
Should I use semantic caching or exact-match caching?
Build exact-match caching first since it is simpler and never returns a wrong answer. Add semantic caching once you see repeated, paraphrased requests in your logs, and calibrate the similarity threshold against labeled pairs rather than a default value, since hit rate depends entirely on your prompt diversity.
When does self-hosting a model become cheaper than a hosted API?
Self-hosting tends to pay off only once your volume is high and steady enough to keep GPU utilization consistently high throughout the day. Below that volume, hosted APIs are usually cheaper once you account for the engineering time needed to run and monitor your own serving infrastructure.
