Why OpenAI’s Model is the Standard
Traditional SaaS billing doesn’t handle the variable cost of AI usage well. OpenAI’s model solves three problems at once:- Predictable Revenue and Low Risk: Because API usage is prepaid, users can’t run up bills they can’t pay. OpenAI receives the money up front, and the user spends it as they use the service.
- Scalability for Developers: A $5 top-up is a low barrier to entry. As an application grows, developers can automate top-ups or buy larger packs. Starting is cheap, and usage can grow without a plan change.
- User Psychology: Credits denominated in US dollars, instead of abstract “tokens” or “points”, make the value clear. The balance works like a prepaid account for AI services, which makes budgeting easier for companies.
How OpenAI Bills
OpenAI runs two billing models for different users.- API (Pay-as-you-go): The API uses prepaid, dollar-denominated credits. Users top up their accounts with $5, $10, $50, or more. The credits show a dollar value but can’t be used outside OpenAI. OpenAI bills per token, with different rates for input and output tokens. Purchased credits expire one year after purchase and are non-refundable (OpenAI Help Center). When a balance reaches $0, API calls fail.
- ChatGPT Plus, Business, and Enterprise: These are flat-rate subscriptions. ChatGPT Plus costs $20 per month, and the Business plan (formerly Team) costs $25 per user per month when billed monthly. They have soft usage caps: heavy users move to a smaller model instead of being blocked.
- Spend-based rate tiers: As total spend grows over time, the account unlocks higher API rate limits. Access grows with billing history.
What Makes It Unique
Four characteristics make OpenAI’s billing effective for AI services:- Fiat-denominated credits: Credits are in US dollars, so they feel like money. Developers can read the price of a request directly.
- Long expiry: Purchased credits last a year, which reduces “use it or lose it” pressure. Users are comfortable topping up larger amounts.
- Multi-dimensional metering: Input and output tokens are tracked separately but deduct from the same balance. OpenAI can price expensive output tokens higher than input tokens.
- Trust tiers: Rate limits that rise with total spend reward long-term customers and encourage them to stay.
Strategic Advantages
The model reinforces itself. Low entry costs bring in developers. Prepaid credits provide immediate cash flow. Usage-based pricing means OpenAI earns more as developers succeed. Subscriptions add a steady baseline of revenue from non-developers.Build This with Dodo Payments
You can build OpenAI’s billing model with Dodo Payments. Use Credit-Based Billing for the API side and standard subscriptions for the ChatGPT Plus side.1
Create a Fiat Credit Entitlement
In your Dodo Payments dashboard, go to Products → Credits and click Create Credit. This credit is the central balance for each user.
- Credit Type: Fiat Credits, with Unit Currency set to USD
- Credit Expiry: Custom, 365 days (matches OpenAI’s one-year expiry), or Never
- Rollover: Not needed (credits don’t reset each cycle)
- Allow Overage: Disabled
2
Create Top-Up Products
Create one-time payment products for different credit packs, such as $5, $10, $50, and $100. Attach your fiat credit to each product.Set the number of credits issued to the pack’s dollar value. A $50 pack issues 50 credits.
3
Create Usage Meters
Create two meters to track token usage:
llm.input_tokens: Sum aggregation on thetokensproperty.llm.output_tokens: Sum aggregation on thetokensproperty.
Calculating Meter Units per Credit
To match OpenAI’s GPT-4o pricing, work out how many tokens cost $1, which is one fiat credit:- Input Tokens: 1,000,000 tokens / $2.50 = 400,000 tokens per $1.
- Output Tokens: 1,000,000 tokens / $10.00 = 100,000 tokens per $1.
4
Send Usage Events
After each LLM request, send the usage to Dodo Payments. One request can carry both the input and the output event. This snippet reuses the
client from the previous step.5
Handle Balance Depletion
Check the user’s balance before you process an API request. If the balance is zero or negative, reject the request, for example with a
402 status.Handling Low Balance Webhooks
Notify users before they reach $0. Set a Low Balance Threshold when you attach the credit, then send an email or in-app notification when thecredit.balance_low webhook arrives.6
Build the ChatGPT Subscription Side (Optional)
To offer a subscription plan like ChatGPT Plus, create a separate subscription product in Dodo Payments. It doesn’t need a credit entitlement.For a Team plan, use seat-based billing: a per-seat add-on whose quantity is the number of users.
Implementing Soft Caps
To build soft caps, track subscription users’ usage with the same meters but without linking them to a credit. In your application, check the usage for the current billing period.Accelerate with the LLM Ingestion Blueprint
The steps above build and send usage events by hand. The LLM Ingestion Blueprint instead wraps your OpenAI client and tracks tokens automatically.inputTokens, outputTokens, and totalTokens from every API response and sends them, with model, as event metadata. Set your meter’s Over Property to the token key you want to bill.
Implementing Spend-Based Rate Tiers
OpenAI’s rate tiers manage capacity by trust. To build them, track each customer’s lifetime spend.- Track Lifetime Spend: Listen for
payment.succeededwebhooks and add the payment amount to atotal_spendfield for that customer in your database. Amounts are in the smallest currency unit, so 5000 is $50.00. - Define Tiers: Map spend amounts to rate limits:
- Tier 1: $0 - $50 spend -> 3 RPM
- Tier 2: $50 - $250 spend -> 10 RPM
- Tier 3: $250+ spend -> 50 RPM
- Enforce Limits: In your API middleware, look up the customer’s tier and apply its rate limit.
Full Implementation Example: The API Proxy
In production, an API proxy usually sits between your users and the LLM provider. The proxy authenticates the request, checks credits, and reports usage. The handler below implements the proxy:Handling Edge Cases
A billing system like OpenAI’s has several edge cases to plan for.Race Conditions
A user with a low balance can send several requests at once and exceed the balance before any event is processed. To prevent this, keep a small buffer, or hold a distributed lock on the customer’s balance during each request.Event Ingestion Latency
Dodo Payments deducts credits asynchronously. A background worker processes new events about once a minute, so a deduction can lag the API call. For strict real-time enforcement, keep a local cache of each user’s balance and update it as you serve requests.Refund Handling
Refunding a credit pack purchase doesn’t remove the credits it granted. When you refund, deduct those credits yourself with a debit: use Apply Credit/Debit on the customer’s Credits tab, or the Create Ledger Entry API. Then update your application’s view of the balance, so users can’t spend credits they no longer have.Multi-Model Support
To support several models with different prices, choose one of two options:- Separate Meters: Create one set of meters per model, for example
gpt-4o.input_tokensandgpt-4o-mini.input_tokens, each with its own Meter units per credit. - Weighted Events: Use one meter and multiply
tokensby a weight before sending the event. For example, if GPT-4o costs 10 times as much as GPT-4o-mini, send 10 times the tokens for GPT-4o requests.
Architecture Overview
The loop below shows the prepaid flow from purchase to blocked calls: The meters track tokens and deduct their dollar value from the user’s credit balance at your configured rates. Your application blocks calls when the balance reaches zero.Conclusion
With Dodo Payments, you can combine usage-based billing with the predictability of prepaid credits, as OpenAI does. Customers pay up front, spend as they go, and top up when they need more. The same pieces work for a large LLM platform or a small AI tool: a fiat credit, top-up products, token meters, and a balance check before each request.Key Dodo Features Used
These Dodo Payments features power the implementation:Credit-Based Billing
Manage prepaid fiat credits and entitlements for your users.
Usage-Based Billing
Track granular usage like tokens and bill for it.
One-Time Payments
Sell credit packs and top-ups through checkout.
Event Ingestion
Send high-volume usage data to Dodo Payments.
Webhooks
Stay updated on credit balance changes and low balance alerts.
LLM Ingestion Blueprint
Automatic token tracking for OpenAI and other LLM providers.