As LLMs and semantic search power more SaaS products in 2026, vendors are still wrestling with a core pricing question: how do you charge for compute‑intensive, variable inference workloads without undermining margins or customer predictability? The dominant approaches — per‑token, per‑embed (per‑vector), and per‑call (per‑transaction) pricing — each make different tradeoffs between cost pass‑through, billing complexity and revenue stability.

Why pricing model matters now

Three shifts have sharpened the need for precise pricing design. First, inference costs are more material to gross margin than they were for traditional multi‑tenant SaaS: GPU hours, memory footprints, and retrieval costs for vector search can represent a sizable, variable outflow. Second, product behavior diversification — from chat interfaces to high‑concurrency API access and large‑scale embedding pipelines — means a one‑size pricing approach misaligns incentives. Third, buyer expectations for predictability have hardened: finance teams want predictable ARR and clear cost drivers, while engineering teams demand pricing that reflects true resource usage.

Pricing models defined

  • Per‑token pricing: chargers bill based on input+output tokens processed by the model. Common for chat/LLM inference.
  • Per‑embed (per‑vector) pricing: charges are based on the number of embeddings created or stored and sometimes on retrieval (per query retrieval count), often relevant for semantic search and RAG workflows.
  • Per‑call pricing: a flat fee per API call or per transaction regardless of token length or embedding dimension. Typical for feature‑rich endpoints where usage is counted in transactions.

Cost drivers and how they map to pricing

Understanding which model maps best to your cost base requires breaking out the main cost drivers:

  • Inference compute — scales with model size, tokens processed, and latency requirements. Per‑token pricing most directly aligns with this driver.
  • Embedding creation — scales with total embeddings generated and model used; per‑embed pricing aligns here.
  • Vector index storage & retrieval — storage costs scale with vector count and dimensionality; retrieval costs scale with query volume and index architecture (ANN vs exact). Per‑embed plus a retrieval component is common.
  • Orchestration and augmentation costs — costs for retrieval, pre/post processing, and external API calls can be significant and may be uncaptured by simple per‑token models.

Comparative analysis: margin, predictability, and customer behavior

This section evaluates each model on four axes: margin alignment (how well price reflects marginal cost), revenue predictability, customer adoption friction, and billing simplicity.

Per‑token

  • Margin alignment: High. Because inference compute often scales with tokens, per‑token billing closely tracks the largest single variable cost.
  • Predictability: Moderate to low for customers. Token usage can burst (long responses or abusive loops) and is sensitive to prompt engineering and model updates.
  • Customer behavior: Encourages customers to optimize prompts and reduce verbosity; can deter use cases that require long context or lengthy generation.
  • Complexity: Moderate. Requires accurate tokenization and transparent reporting; tokens differ by model/tokenizer which vendors must disclose.

Per‑embed

  • Margin alignment: Strong for embedding‑centric products. It directly charges for generation and often storage; however, retrieval costs can be undercharged if retrieval is heavy.
  • Predictability: High for generation; moderate overall because query volume (retrieval) can introduce variability.
  • Customer behavior: Favors architectures that batch embed generation and reuse embeddings; incentivizes customers to avoid excessive regeneration.
  • Complexity: Requires tracking counts and dimensions and communicating policies for when embeddings must be refreshed (e.g., model upgrades).

Per‑call

  • Margin alignment: Weak to moderate. Flat per‑call fees can subsidize expensive calls (long outputs or heavy retrieval) or overcharge simple ones.
  • Predictability: Very high for customers — simple to forecast a cost per transaction — and thus popular for product‑led adoption and customer finance teams.
  • Customer behavior: Encourages many small calls rather than fewer large ones unless the vendor prices tiers carefully; risk of inefficient usage.
  • Complexity: Low. Easiest to explain and implement, but requires caps/guards to prevent misuse or create premium tiers for heavy calls.

When to use each model — product archetypes

Choose pricing based on product behavior and buyer priorities, not fashion. Below are practical archetypes and recommended models.

  • Conversational AI/chat apps: Per‑token is usually best because response length and context size dominate cost. If customers highly value predictability, combine a base subscription with token allotments and clear overage pricing.
  • Semantic search / RAG & knowledge base products: Per‑embed plus retrieval‑query fees (or a small per‑call retrieval fee) aligns best. Charge per embedding generation plus a retrieval‑quota or per‑query price to capture both storage and query compute.
  • Event‑driven inference (notifications, classification APIs): Per‑call works well where each event is a predictable unit and responses are bounded. If classification requires long context, consider hybrid per‑call + per‑token overage.
  • Marketplace or platform aggregators: Use composite pricing that allocates a base per‑seat or per‑app fee, per‑embed for indexing, and per‑token for on‑demand generations. This preserves predictability for platform hosts while capturing variable costs.

Design patterns that blend tradeoffs

Most successful pricing strategies are hybrids that combine the strengths of these models:

  • Base subscription + usage credits: A fixed monthly fee with included token/embedding credits reduces churn and delivers predictable ARR while allowing cost pass‑through for heavy use.
  • Tiered per‑unit pricing: Progressive unit pricing (cheaper per‑token or per‑embed above thresholds) can align with volume economies and smooth revenue recognition.
  • Guardrails + transparency: Offer per‑unit billing with enforcement tools (rate limits, warning thresholds, throttles) and real‑time reporting so customers can manage spend.
  • Value‑based adjuncts: For enterprise buyers, supplement usage pricing with value metrics (records indexed, seats with analytics) so sellers capture willingness to pay beyond raw compute consumption.

Practical framework for choosing a model

  1. Map product actions to cost centers: identify which actions drive GPU, storage, or retrieval costs.
  2. Rank buyer sensitivity to predictability versus unit fairness: enterprise finance often prefers predictability; developers prefer usage elasticity.
  3. Estimate behavioral elasticity: will customers change behavior in response to per‑token pricing (compress prompts), or will they accept cost if it unlocks value?
  4. Prototype hybrid billing and measure MRR volatility and churn in a small cohort before full rollout.

Conclusion — align incentives, not just costs

Per‑token, per‑embed and per‑call pricing are not interchangeable; each maps to different cost structures and customer expectations. Vendors who succeed in 2026 will be those that map pricing to the dominant cost drivers of their product, combine predictability mechanisms that customers value, and instrument billing to show customers the levers to control spend. In practice that means hybrid models — a predictable base, usage credits that track primary compute drivers, and transparent tooling — rather than ideological loyalty to a single unit of charge.

For SaaS pricing teams, the actionable takeaway is simple: run experiments that pair cost telemetry (GPU hours, index IO, queries) with billing telemetry (tokens, embeddings, calls), and pick the model that minimizes margin leakage while preserving the customer behaviors that deliver product value.