Skip to main content
OpenAI combines prepaid fiat credits for API usage with flat-rate subscriptions for its consumer products. The prepaid side gives OpenAI cash up front, and developers can scale usage without a sales conversation. Many AI companies copy this model.

Why OpenAI’s Model is the Standard

Traditional SaaS billing doesn’t handle the variable cost of AI usage well. OpenAI’s model solves three problems at once:
  1. Predictable Revenue and Low Risk: Because API usage is prepaid, users can’t run up bills they can’t pay. OpenAI receives the money up front, and the user spends it as they use the service.
  2. Scalability for Developers: A $5 top-up is a low barrier to entry. As an application grows, developers can automate top-ups or buy larger packs. Starting is cheap, and usage can grow without a plan change.
  3. User Psychology: Credits denominated in US dollars, instead of abstract “tokens” or “points”, make the value clear. The balance works like a prepaid account for AI services, which makes budgeting easier for companies.

How OpenAI Bills

OpenAI runs two billing models for different users.
  1. API (Pay-as-you-go): The API uses prepaid, dollar-denominated credits. Users top up their accounts with $5, $10, $50, or more. The credits show a dollar value but can’t be used outside OpenAI. OpenAI bills per token, with different rates for input and output tokens. Purchased credits expire one year after purchase and are non-refundable (OpenAI Help Center). When a balance reaches $0, API calls fail.
  2. ChatGPT Plus, Business, and Enterprise: These are flat-rate subscriptions. ChatGPT Plus costs $20 per month, and the Business plan (formerly Team) costs $25 per user per month when billed monthly. They have soft usage caps: heavy users move to a smaller model instead of being blocked.
  3. Spend-based rate tiers: As total spend grows over time, the account unlocks higher API rate limits. Access grows with billing history.
The token prices below are the published rates at the time of writing:

What Makes It Unique

Four characteristics make OpenAI’s billing effective for AI services:
  • Fiat-denominated credits: Credits are in US dollars, so they feel like money. Developers can read the price of a request directly.
  • Long expiry: Purchased credits last a year, which reduces “use it or lose it” pressure. Users are comfortable topping up larger amounts.
  • Multi-dimensional metering: Input and output tokens are tracked separately but deduct from the same balance. OpenAI can price expensive output tokens higher than input tokens.
  • Trust tiers: Rate limits that rise with total spend reward long-term customers and encourage them to stay.

Strategic Advantages

The model reinforces itself. Low entry costs bring in developers. Prepaid credits provide immediate cash flow. Usage-based pricing means OpenAI earns more as developers succeed. Subscriptions add a steady baseline of revenue from non-developers.

Build This with Dodo Payments

You can build OpenAI’s billing model with Dodo Payments. Use Credit-Based Billing for the API side and standard subscriptions for the ChatGPT Plus side.
1

Create a Fiat Credit Entitlement

In your Dodo Payments dashboard, go to Products → Credits and click Create Credit. This credit is the central balance for each user.
  • Credit Type: Fiat Credits, with Unit Currency set to USD
  • Credit Expiry: Custom, 365 days (matches OpenAI’s one-year expiry), or Never
  • Rollover: Not needed (credits don’t reset each cycle)
  • Allow Overage: Disabled
Fiat credits use two decimal places, so one credit is one dollar and balances track cents. Dodo Payments doesn’t block usage when a balance reaches zero. To make API calls fail at $0 like OpenAI, check the balance in your application before each request (see Handle Balance Depletion below).
2

Create Top-Up Products

Create one-time payment products for different credit packs, such as $5, $10, $50, and $100. Attach your fiat credit to each product.Set the number of credits issued to the pack’s dollar value. A $50 pack issues 50 credits.
3

Create Usage Meters

Create two meters to track token usage:
  • llm.input_tokens: Sum aggregation on the tokens property.
  • llm.output_tokens: Sum aggregation on the tokens property.
On your usage-based product, toggle Bill usage in Credits on both meters and select the fiat credit. Then set Meter units per credit for each.

Calculating Meter Units per Credit

To match OpenAI’s GPT-4o pricing, work out how many tokens cost $1, which is one fiat credit:
  • Input Tokens: 1,000,000 tokens / $2.50 = 400,000 tokens per $1.
  • Output Tokens: 1,000,000 tokens / $10.00 = 100,000 tokens per $1.
In the Dodo Payments dashboard, set Meter units per credit to 400,000 for input and 100,000 for output. Dodo Payments divides each meter’s aggregated tokens by this value to get the credits to deduct.
4

Send Usage Events

After each LLM request, send the usage to Dodo Payments. One request can carry both the input and the output event. This snippet reuses the client from the previous step.
5

Handle Balance Depletion

Check the user’s balance before you process an API request. If the balance is zero or negative, reject the request, for example with a 402 status.

Handling Low Balance Webhooks

Notify users before they reach $0. Set a Low Balance Threshold when you attach the credit, then send an email or in-app notification when the credit.balance_low webhook arrives.
OpenAI offers auto recharge, which buys more credits when the balance falls below a threshold the user sets.
6

Build the ChatGPT Subscription Side (Optional)

To offer a subscription plan like ChatGPT Plus, create a separate subscription product in Dodo Payments. It doesn’t need a credit entitlement.For a Team plan, use seat-based billing: a per-seat add-on whose quantity is the number of users.

Implementing Soft Caps

To build soft caps, track subscription users’ usage with the same meters but without linking them to a credit. In your application, check the usage for the current billing period.

Accelerate with the LLM Ingestion Blueprint

The steps above build and send usage events by hand. The LLM Ingestion Blueprint instead wraps your OpenAI client and tracks tokens automatically.
The blueprint reads inputTokens, outputTokens, and totalTokens from every API response and sends them, with model, as event metadata. Set your meter’s Over Property to the token key you want to bill.
The LLM Blueprint supports OpenAI, Anthropic, Groq, Google Gemini, OpenRouter, and the Vercel AI SDK. See the full blueprint documentation for provider-specific examples and advanced configuration.

Implementing Spend-Based Rate Tiers

OpenAI’s rate tiers manage capacity by trust. To build them, track each customer’s lifetime spend.
  1. Track Lifetime Spend: Listen for payment.succeeded webhooks and add the payment amount to a total_spend field for that customer in your database. Amounts are in the smallest currency unit, so 5000 is $50.00.
  2. Define Tiers: Map spend amounts to rate limits:
    • Tier 1: $0 - $50 spend -> 3 RPM
    • Tier 2: $50 - $250 spend -> 10 RPM
    • Tier 3: $250+ spend -> 50 RPM
  3. Enforce Limits: In your API middleware, look up the customer’s tier and apply its rate limit.

Full Implementation Example: The API Proxy

In production, an API proxy usually sits between your users and the LLM provider. The proxy authenticates the request, checks credits, and reports usage. The handler below implements the proxy:

Handling Edge Cases

A billing system like OpenAI’s has several edge cases to plan for.

Race Conditions

A user with a low balance can send several requests at once and exceed the balance before any event is processed. To prevent this, keep a small buffer, or hold a distributed lock on the customer’s balance during each request.

Event Ingestion Latency

Dodo Payments deducts credits asynchronously. A background worker processes new events about once a minute, so a deduction can lag the API call. For strict real-time enforcement, keep a local cache of each user’s balance and update it as you serve requests.

Refund Handling

Refunding a credit pack purchase doesn’t remove the credits it granted. When you refund, deduct those credits yourself with a debit: use Apply Credit/Debit on the customer’s Credits tab, or the Create Ledger Entry API. Then update your application’s view of the balance, so users can’t spend credits they no longer have.

Multi-Model Support

To support several models with different prices, choose one of two options:
  1. Separate Meters: Create one set of meters per model, for example gpt-4o.input_tokens and gpt-4o-mini.input_tokens, each with its own Meter units per credit.
  2. Weighted Events: Use one meter and multiply tokens by a weight before sending the event. For example, if GPT-4o costs 10 times as much as GPT-4o-mini, send 10 times the tokens for GPT-4o requests.
OpenAI publishes a separate rate for each model, and separate meters map to that structure most directly.

Architecture Overview

The loop below shows the prepaid flow from purchase to blocked calls: The meters track tokens and deduct their dollar value from the user’s credit balance at your configured rates. Your application blocks calls when the balance reaches zero.

Conclusion

With Dodo Payments, you can combine usage-based billing with the predictability of prepaid credits, as OpenAI does. Customers pay up front, spend as they go, and top up when they need more. The same pieces work for a large LLM platform or a small AI tool: a fiat credit, top-up products, token meters, and a balance check before each request.

Key Dodo Features Used

These Dodo Payments features power the implementation:

Credit-Based Billing

Manage prepaid fiat credits and entitlements for your users.

Usage-Based Billing

Track granular usage like tokens and bill for it.

One-Time Payments

Sell credit packs and top-ups through checkout.

Event Ingestion

Send high-volume usage data to Dodo Payments.

Webhooks

Stay updated on credit balance changes and low balance alerts.

LLM Ingestion Blueprint

Automatic token tracking for OpenAI and other LLM providers.
Last modified on September 26, 2026