Skip to main content
Replicate is a platform for running open-source machine learning models in the cloud, and one of the purest examples of usage-based pricing in AI. There is no monthly subscription and no flat rate per model run. Most models are billed for the compute time they use, measured per second, at a rate that depends on the hardware. Some models are billed by input and output instead, such as per image or per token (Replicate pricing). Per-second billing suits AI workloads because run times are unpredictable. One user might run a lightweight model for a few seconds, and another a large generative model for several minutes. Because the price follows the compute used rather than the model, it stays transparent at any scale.

How Replicate Bills

Replicate’s time-based pricing doesn’t depend on the model. Whether you generate an image with SDXL or run Llama 3, the bill depends on the hardware tier and the run time. Replicate can host thousands of open-source models without a separate price for each. The implementation below also uses example meters for A40 and A100 (40GB) tiers, which Replicate’s pricing page no longer lists. Create one meter for each hardware tier you offer. Each run is metered on the meter for the hardware it used:
  1. Hardware-Specific Rates: The price per second depends on the compute resources used. Each hardware tier has its own rate.
  2. Pure Usage-Based Model: There are no monthly fees, no included quotas, and no overages. Users pay for exact compute time, for example “12.4 seconds on an A100”, not per generation.
  3. Per-Second Granularity: Providers that bill by the hour or minute charge for unused time on short tasks. Per-second billing removes that waste for small experiments and large production workloads alike.
For public models, Replicate charges only for the time a model is actively processing your request. Setup (cold boot) and idle time are free. Private models and deployments run on dedicated hardware, so you pay for all the time their instances are online: setup, idle, and active (Replicate billing docs).

What Makes It Unique

  • Hardware-specific metering: The same model costs more on faster hardware, so users trade speed against cost. A T4 GPU suits tasks that aren’t time-sensitive, and an A100 suits real-time applications.
  • Per-second granularity: Billing is calculated to the second, so users don’t pay for unused time on short tasks.
  • No subscription: Users start with no commitment, and cost scales with usage. This suits startups and developers trying different models.
  • Model-agnostic: The same billing logic applies to image generation, text processing, audio transcription, and video synthesis. The platform supports a large model catalog without complex pricing tables.

Build This with Dodo Payments

You can build this model with Dodo Payments usage-based billing. Create one meter per hardware tier and attach all of them to a single product.
1

Create Usage Meters (One Per Hardware Class)

Create a separate meter for each hardware tier. Each tier has its own price per second, so separate meters let Dodo Payments price each tier and itemize the invoice.Sum aggregation over execution_seconds gives the total compute time per hardware tier for the billing period.
2

Create a Usage-Based Product

Create a product in the Dodo Payments dashboard with these settings:
  • Pricing type: Usage Based Billing
  • Base Price: $0/month (no subscription fee)
  • Billing frequency: Monthly
Add every meter with its price per unit:Set the Free Threshold to 0 on every meter, so every second of execution is billable.
3

Send Usage Events

Send a usage event to Dodo Payments when each model run completes. Give each prediction a unique event_id, so a retried event isn’t counted twice.
4

Measure Execution Time Precisely

Time each model run with performance.now(), and round to the nearest tenth of a second for billing.
5

Create Checkout

When a user signs up, create a checkout session for the usage-based product. Dodo Payments then bills the usage each cycle and issues the invoices.

Accelerate with the Time Range Ingestion Blueprint

The Time Range Ingestion Blueprint shortens per-second compute tracking. Create one ingestion instance per hardware tier, and call trackTimeRange after each run.
The blueprint builds and sends the event, with the duration as durationMs in the metadata. Set each meter’s Over Property to durationMs, or send durationSeconds to keep per-second pricing. With one ingestion instance per hardware tier, this maps directly to Replicate’s multi-tier metering.
For long-running jobs, combine the Time Range Blueprint with interval-based heartbeat tracking, shown in Advanced: Heartbeat Metering. See the full blueprint documentation for more patterns.

Cost Estimation for Users

Usage-based bills can be hard to predict, so show users a cost estimate before they run a model. Estimates prevent surprise bills and build trust.

Example Cost Calculations

Building a Cost Calculator

This function multiplies the tier’s per-second rate by the estimated run time:

Enterprise: Reserved Capacity

For customers who need dedicated capacity and no cold boots, Replicate offers deployments: dedicated instances billed for all the time they’re online. To model reserved capacity with Dodo Payments, sell it as a subscription product:
  • Product Type: Subscription
  • Price: Fixed monthly price (for example, “Reserved A100 Instance - $500/month”)
  • Billing Cycle: Monthly
You can still send usage events for monitoring and analytics, while the subscription covers the cost. As a customer’s volume grows, reserved capacity often costs less than pay-as-you-go.

Advanced: Heartbeat Metering

For tasks that run for minutes or hours, one event at the end is risky: if the process crashes, you lose the usage data. Instead, send a usage event every 30 to 60 seconds while the task runs.

Key Dodo Features Used

Usage-Based Billing

Set up products that bill based on consumption.

Meters

Define the metrics you want to track and bill for.

Event Ingestion

Send usage data to Dodo Payments as it happens.

Subscriptions

Manage recurring billing for reserved capacity and enterprise plans.

Time Range Blueprint

Per-second compute tracking with duration helpers.
Last modified on September 26, 2026