LLM cost per feature in 2026: four workload patterns

By LLMTest Team · Jul 1, 2026 · 6 min read costsolopreneursapivibe-coders
On this page

On this page

  1. Chat: support bots, FAQ assistants, copilots
  2. Extraction: invoices, PDFs, structured forms
  3. Summarization: transcripts, articles, support tickets
  4. Coding: PR review, code generation, bug fixing
  5. Subscription vs API
  6. The workload snapshot

Before you pick a model, count your tokens. The choice of Claude Haiku 4.5 versus Claude Sonnet 5 versus GPT-5.5 matters a lot less than understanding your workload's token profile: how many tokens go in, how many come out, and whether the inputs repeat enough for caching to bite. The math runs in five minutes on a napkin. Skipping it is the fastest way to ship an AI feature with unit economics you can't explain six months in.

Four patterns cover the majority of AI features solo builders ship: chat, extraction, summarization, and coding. Each has a different token shape and different cost behavior across model tiers.

Chat: support bots, FAQ assistants, copilots

A typical chat request sends 1,000 input tokens (a 500-token system prompt plus conversation history and the user message) and receives 300 output tokens.

Model Per request Per 1,000 requests
GPT-4o-mini ($0.15/$0.60 per MTok) $0.00033 $0.33
Gemini 2.5 Flash ($0.30/$2.50) $0.00105 $1.05
Claude Haiku 4.5 ($1/$5) $0.0025 $2.50
Claude Sonnet 5 ($3/$15) $0.0075 $7.50
GPT-5.5 ($5/$30) $0.014 $14.00

The 42x spread between GPT-4o-mini and GPT-5.5 comes mostly from output pricing, not input. Output tokens cost 4–6x more than input tokens within the same model, so a chatbot that's verbose is disproportionately expensive.

The system prompt (the static 500 tokens) can be cached. Anthropic's prompt caching charges cache reads at 10% of the base input rate after a 1.25x write cost. At a 90% cache-hit rate, Haiku 4.5 drops from $2.50 to about $2.00 per 1,000 requests. The savings are modest here because the total token count is small. Caching earns its keep on heavier workloads.

Extraction: invoices, PDFs, structured forms

Extraction is input-heavy: you send the full document and get a compact JSON payload back. A typical invoice or form runs 3,000 input tokens and 500 output tokens.

Model Per 1,000 requests With 50% batch discount
GPT-4o-mini $0.75 $0.38
Gemini 2.5 Flash $2.15 $1.08
Claude Haiku 4.5 $5.50 $2.75
Claude Sonnet 5 $16.50 $8.25
GPT-5.5 $30.00 $15.00

Extraction jobs are naturally async: a user submits a batch of invoices and checks back later. OpenAI's Batch API and Anthropic's Message Batches API both offer 50% off for async processing. At that discount, Haiku 4.5 drops from $5.50 to $2.75 per 1,000 pages, and GPT-4o-mini becomes nearly trivial at $0.38.

The tradeoff: batch mode adds minutes to hours of latency depending on queue depth. It only makes sense where real-time response isn't required. For a document extraction SaaS selling at $0.05/page (a setup explored in the AI feature pricing worked examples), Haiku 4.5 batch at $2.75/1k leaves a strong margin at that price point.

Summarization: transcripts, articles, support tickets

Summarization flips the ratio: very long input, short output. A 60-minute meeting transcript runs 5,000–8,000 tokens. Assume 6,000 input and 400 output.

Model Per 1,000 summaries With batch discount
GPT-4o-mini $1.14 $0.57
Gemini 2.5 Flash $2.80 $1.40
Claude Haiku 4.5 $8.00 $4.00
Claude Sonnet 5 $24.00 $12.00
GPT-5.5 $42.00 $21.00

Summarization is the clearest case for the Batch API. Users who upload meeting recordings already expect to wait. The 50% discount is essentially free money.

Prompt prefix caching stacks on top of the batch discount and is worth modeling separately. If your summarization feature uses the same 1,000-token instruction prefix across thousands of requests (formatting rules, output length, language), cached reads cost 10% of the base input rate. On Haiku at a 90% cache-hit rate, that prefix saves $0.81 per 1,000 requests before the batch discount is applied, a non-trivial reduction on a feature billed at $8/1k. For the break-even analysis on when caching pays off, prompt caching explained runs through the math for different hit rates and prefix lengths.

Coding: PR review, code generation, bug fixing

Code review sends the largest payloads. A typical PR review: 4,000 input tokens (code context plus diff plus instructions) and 800 output tokens (review comments and suggestions).

Model Per 1,000 reviews Notes
Claude Haiku 4.5 $8.00 Adequate for simple bug fixes
Claude Sonnet 5 $24.00 Strong on architecture feedback
Claude Opus 4.8 $40.00 Best depth, overkill for routine PRs
GPT-5.5 $44.00 Competitive coding performance

The $16 gap between Haiku and Sonnet 5 compounds fast. At 5 PR reviews per developer per week across 10 developers, that's about 2,600 reviews per quarter: a $41.60 cost difference per quarter between models. The question is whether Sonnet 5's review depth is worth $42 per year per developer. Based on our head-to-head on coding tasks, the output quality gap between Haiku and Sonnet 5 is real for multi-file refactors but smaller for isolated bug fixes.

One practical tactic: route by PR size. Small PRs (under 300 changed lines) go to Haiku. Complex or large PRs go to Sonnet 5. At a 70/30 split, the blended cost falls to roughly $13 per 1,000 reviews, about the same as Haiku-only but with Sonnet 5 covering the hard cases.

Subscription vs API

Provider API Subscription What the subscription covers
Anthropic $1–5/M tokens Claude Pro $20/mo, Max $100–200/mo Personal claude.ai + Claude Code access; does not cover your users' API calls
OpenAI $0.15–5/M tokens ChatGPT Plus $20/mo, Pro $200/mo Personal ChatGPT access; separate from your API balance
Google $0.15–$X/M Gemini Advanced (in Google One) Personal Gemini access; your API runs on separate AI Studio quota

Subscriptions don't fund your product's users. Claude Pro and ChatGPT Plus give you, the developer, personal access for your own workflow. Your product's users consume API tokens billed to your provider balance separately.

The break-even on a developer subscription: Claude Max at $100/month makes sense if your daily Claude Code sessions would exceed $100 at API rates. At Sonnet 5's $3/M input and $15/M output, a heavy coding session with 200k input tokens and 20k output tokens costs about $0.90. If you run 120 sessions per month, API-direct beats Max at $108, close enough that the session-limit-free Max plan is worth it for heavy users.

For your product, use API rates throughout. Subscriptions are a separate developer tooling line item.

The workload snapshot

Workload Token profile Haiku 4.5 /1k Sonnet 5 /1k GPT-5.5 /1k
Chat 1k in / 300 out $2.50 $7.50 $14.00
Extraction 3k in / 500 out $5.50 $16.50 $30.00
Summarization 6k in / 400 out $8.00 $24.00 $42.00
Coding 4k in / 800 out $8.00 $24.00 $44.00

Two patterns stand out. First, high input-to-output ratios (extraction, summarization) benefit from the Batch API far more than chat does. Second, GPT-5.5's output price ($30/M) is 2x Sonnet 5's ($15/M). That gap compounds on high-output workloads: for anything generating 1,500+ output tokens per request, Sonnet 5 has a structural cost advantage over GPT-5.5 at comparable quality.

The numbers above use estimated token counts. Your actual prompts may be shorter or longer. The LLMTest proxy logs cost and latency per call in production, which lets you find the tail of expensive requests that a simple average hides. A batch extraction job that averages 3,000 tokens might have 5% of requests hitting 8,000 tokens. Those tail requests determine whether your margin holds on a bad day, not just a typical one.


To test your actual prompts across these models and see per-request costs before you commit to one, start with LLMTest.

Ship LLM features without burning your budget.

LLMTest proxies your OpenAI / Anthropic calls, tracks cost per feature, and auto-rewrites prompts to be cheaper while holding quality. Free to start.

Create a free account

Related articles

Pricing an AI feature in 2026: three worked examples
Work backward from your margin target to a token budget: break-even math for AI features in 2026, with three concrete worked examples in USD.
Prompt caching explained: Anthropic, OpenAI, and Gemini in 2026
Prompt caching cuts LLM API costs up to 90%, but Anthropic, OpenAI, and Gemini implement it differently. Here's how each vendor's billing actually works.