Skip to main content
Replicate क्लाउड में open-source machine learning models चलाने का एक platform है और AI में usage-based pricing के सबसे शुद्ध उदाहरणों में से एक है। इसमें कोई monthly subscription या प्रति model run flat rate नहीं है। अधिकांश models से उनके द्वारा उपयोग किए गए compute time के लिए, प्रति second के हिसाब से billing की जाती है, जिसकी दर hardware पर निर्भर करती है। कुछ models से इसके बजाय input और output के आधार पर billing की जाती है, जैसे प्रति image या प्रति token (Replicate pricing)। AI workloads के लिए per-second billing उपयुक्त है, क्योंकि run times का अनुमान लगाना कठिन होता है। कोई user कुछ seconds के लिए lightweight model चला सकता है, जबकि दूसरा कई minutes के लिए बड़ा generative model चला सकता है। चूँकि price model के बजाय उपयोग किए गए compute के अनुसार तय होती है, इसलिए यह किसी भी scale पर पारदर्शी रहती है।

रेप्लिकेट कैसे बिल करता है

Replicate की time-based pricing model पर निर्भर नहीं करती। चाहे आप SDXL से image generate करें या Llama 3 चलाएँ, bill hardware tier और run time पर निर्भर करता है। Replicate प्रत्येक model के लिए अलग price रखे बिना हजारों open-source models host कर सकता है। नीचे दिया गया implementation A40 और A100 (40GB) tiers के लिए example meters का भी उपयोग करता है, जिन्हें Replicate की pricing page अब सूचीबद्ध नहीं करती। आपके द्वारा उपलब्ध कराए जाने वाले प्रत्येक hardware tier के लिए एक meter बनाएं। प्रत्येक run को उस hardware के meter पर मापा जाता है जिसका उसने उपयोग किया:
  1. Hardware-Specific Rates: प्रति second price उपयोग किए गए compute resources पर निर्भर करती है। प्रत्येक hardware tier की अपनी rate होती है।
  2. Pure Usage-Based Model: कोई monthly fee, included quota या overage नहीं है। Users exact compute time के लिए भुगतान करते हैं, उदाहरण के लिए “12.4 seconds on an A100”, न कि प्रति generation।
  3. Per-Second Granularity: जो providers hour या minute के आधार पर billing करते हैं, वे short tasks पर unused time का भी charge लेते हैं। Per-second billing छोटे experiments और बड़े production workloads, दोनों के लिए इस waste को समाप्त करती है।
Public models के लिए, Replicate केवल उस समय का charge करता है जब कोई model आपके request को actively process कर रहा हो। Setup (cold boot) और idle time निःशुल्क हैं। Private models और deployments dedicated hardware पर चलते हैं, इसलिए उनकी instances के online रहने के पूरे समय—setup, idle और active—का भुगतान करना पड़ता है (Replicate billing docs).

इसे विशिष्ट क्या बनाता है

  • Hardware-specific metering: तेज़ hardware पर वही model अधिक महँगा होता है, इसलिए users speed और cost के बीच trade-off चुनते हैं। ऐसे tasks के लिए जो time-sensitive नहीं हैं, T4 GPU उपयुक्त है, जबकि real-time applications के लिए A100 उपयुक्त है।
  • Per-second granularity: Billing को second तक calculate किया जाता है, इसलिए users short tasks पर unused time के लिए भुगतान नहीं करते।
  • No subscription: Users बिना किसी commitment के शुरुआत कर सकते हैं और cost usage के अनुसार बढ़ती है। यह उन startups और developers के लिए उपयुक्त है जो अलग-अलग models आज़मा रहे हैं।
  • Model-agnostic: यही billing logic image generation, text processing, audio transcription और video synthesis पर लागू होता है। यह platform complex pricing tables के बिना बड़े model catalog को support करता है।

Dodo Payments के साथ इसे बनाएँ

आप Dodo Payments usage-based billing के साथ इस model को बना सकते हैं। प्रत्येक hardware tier के लिए एक meter बनाएँ और उन सभी को एक ही product से attach करें।
1

Create Usage Meters (One Per Hardware Class)

प्रत्येक hardware tier के लिए एक अलग meter बनाएँ। प्रत्येक tier की अपनी प्रति second price होती है, इसलिए अलग meters Dodo Payments को प्रत्येक tier की price तय करने और invoice में अलग-अलग item दिखाने की सुविधा देते हैं।execution_seconds पर Sum aggregation से billing period के दौरान प्रत्येक hardware tier का कुल compute time प्राप्त होता है।
2

Create a Usage-Based Product

इन settings के साथ Dodo Payments dashboard में एक product बनाएँ:
  • Pricing type: Usage Based Billing
  • Base Price: $0/month (कोई subscription fee नहीं)
  • Billing frequency: Monthly
प्रत्येक meter को उसकी price per unit के साथ जोड़ें:प्रत्येक meter पर Free Threshold को 0 पर set करें, ताकि execution का प्रत्येक second billable हो।
3

Send Usage Events

प्रत्येक model run पूरा होने पर Dodo Payments को usage event भेजें। प्रत्येक prediction को एक unique event_id दें, ताकि retried event दो बार count न हो।
4

Measure Execution Time Precisely

प्रत्येक model run का समय performance.now() से मापें और billing के लिए second के निकटतम दसवें भाग तक round करें।
5

Create Checkout

जब कोई user sign up करे, तो usage-based product के लिए checkout session बनाएँ। इसके बाद Dodo Payments प्रत्येक cycle में usage की billing करता है और invoices जारी करता है।

Time Range Ingestion Blueprint के साथ तेज़ी लाएँ

Time Range Ingestion Blueprint per-second compute tracking को सरल बनाता है। प्रत्येक hardware tier के लिए एक ingestion instance बनाएँ और प्रत्येक run के बाद trackTimeRange call करें।
Blueprint event बनाकर भेजता है, जिसमें duration metadata में durationMs के रूप में होता है। प्रत्येक meter का Over Property durationMs पर set करें, या per-second pricing बनाए रखने के लिए durationSeconds भेजें। प्रत्येक hardware tier के लिए एक ingestion instance के साथ, यह Replicate के multi-tier metering से सीधे map होता है।
लंबे समय तक चलने वाले jobs के लिए, Time Range Blueprint को interval-based heartbeat tracking के साथ combine करें, जैसा Advanced: Heartbeat Metering में दिखाया गया है। अधिक patterns के लिए full blueprint documentation देखें।

Users के लिए Cost Estimation

Usage-based bills का अनुमान लगाना कठिन हो सकता है, इसलिए users को model run करने से पहले cost estimate दिखाएँ। Estimates unexpected bills को रोकते हैं और trust बनाते हैं।

Example Cost Calculations

Cost Calculator बनाना

यह function tier की per-second rate को estimated run time से multiply करता है:

Enterprise: Reserved Capacity

जिन customers को dedicated capacity और cold boots से मुक्ति चाहिए, उनके लिए Replicate deployments उपलब्ध कराता है: dedicated instances जिनके online रहने के पूरे समय billing होती है। Dodo Payments के साथ reserved capacity को model करने के लिए इसे subscription product के रूप में बेचें:
  • Product Type: Subscription
  • Price: Fixed monthly price (उदाहरण के लिए, “Reserved A100 Instance - $500/month”)
  • Billing Cycle: Monthly
आप monitoring और analytics के लिए usage events भेजना जारी रख सकते हैं, जबकि subscription cost को cover करती है। जैसे-जैसे customer का volume बढ़ता है, reserved capacity की cost अक्सर pay-as-you-go से कम होती है।

Advanced: Heartbeat Metering

जो tasks कई minutes या hours तक चलते हैं, उनके लिए अंत में केवल एक event भेजना जोखिमपूर्ण है: यदि process crash हो जाए, तो usage data खो जाता है। इसके बजाय, task चलने के दौरान हर 30 से 60 seconds में usage event भेजें।

उपयोग की गई प्रमुख Dodo सुविधाएँ

Usage-Based Billing

ऐसे products set up करें जिनकी billing consumption के आधार पर हो।

Meters

उन metrics को define करें जिन्हें आप track और bill करना चाहते हैं।

Event Ingestion

Usage data होने के साथ-साथ Dodo Payments को भेजें।

Subscriptions

Reserved capacity और enterprise plans के लिए recurring billing manage करें।

Time Range Blueprint

Duration helpers के साथ per-second compute tracking।
अंतिम संशोधन 26 सितंबर 2026