रेप्लिकेट कैसे बिल करता है
Replicate की time-based pricing model पर निर्भर नहीं करती। चाहे आप SDXL से image generate करें या Llama 3 चलाएँ, bill hardware tier और run time पर निर्भर करता है। Replicate प्रत्येक model के लिए अलग price रखे बिना हजारों open-source models host कर सकता है।
नीचे दिया गया implementation A40 और A100 (40GB) tiers के लिए example meters का भी उपयोग करता है, जिन्हें Replicate की pricing page अब सूचीबद्ध नहीं करती। आपके द्वारा उपलब्ध कराए जाने वाले प्रत्येक hardware tier के लिए एक meter बनाएं।
प्रत्येक run को उस hardware के meter पर मापा जाता है जिसका उसने उपयोग किया:
- Hardware-Specific Rates: प्रति second price उपयोग किए गए compute resources पर निर्भर करती है। प्रत्येक hardware tier की अपनी rate होती है।
- Pure Usage-Based Model: कोई monthly fee, included quota या overage नहीं है। Users exact compute time के लिए भुगतान करते हैं, उदाहरण के लिए “12.4 seconds on an A100”, न कि प्रति generation।
- Per-Second Granularity: जो providers hour या minute के आधार पर billing करते हैं, वे short tasks पर unused time का भी charge लेते हैं। Per-second billing छोटे experiments और बड़े production workloads, दोनों के लिए इस waste को समाप्त करती है।
Public models के लिए, Replicate केवल उस समय का charge करता है जब कोई model आपके request को actively process कर रहा हो। Setup (cold boot) और idle time निःशुल्क हैं। Private models और deployments dedicated hardware पर चलते हैं, इसलिए उनकी instances के online रहने के पूरे समय—setup, idle और active—का भुगतान करना पड़ता है (Replicate billing docs).
इसे विशिष्ट क्या बनाता है
- Hardware-specific metering: तेज़ hardware पर वही model अधिक महँगा होता है, इसलिए users speed और cost के बीच trade-off चुनते हैं। ऐसे tasks के लिए जो time-sensitive नहीं हैं, T4 GPU उपयुक्त है, जबकि real-time applications के लिए A100 उपयुक्त है।
- Per-second granularity: Billing को second तक calculate किया जाता है, इसलिए users short tasks पर unused time के लिए भुगतान नहीं करते।
- No subscription: Users बिना किसी commitment के शुरुआत कर सकते हैं और cost usage के अनुसार बढ़ती है। यह उन startups और developers के लिए उपयुक्त है जो अलग-अलग models आज़मा रहे हैं।
- Model-agnostic: यही billing logic image generation, text processing, audio transcription और video synthesis पर लागू होता है। यह platform complex pricing tables के बिना बड़े model catalog को support करता है।
Dodo Payments के साथ इसे बनाएँ
आप Dodo Payments usage-based billing के साथ इस model को बना सकते हैं। प्रत्येक hardware tier के लिए एक meter बनाएँ और उन सभी को एक ही product से attach करें।1
Create Usage Meters (One Per Hardware Class)
प्रत्येक hardware tier के लिए एक अलग meter बनाएँ। प्रत्येक tier की अपनी प्रति second price होती है, इसलिए अलग meters Dodo Payments को प्रत्येक tier की price तय करने और invoice में अलग-अलग item दिखाने की सुविधा देते हैं।
execution_seconds पर Sum aggregation से billing period के दौरान प्रत्येक hardware tier का कुल compute time प्राप्त होता है।2
Create a Usage-Based Product
इन settings के साथ Dodo Payments dashboard में एक product बनाएँ:
- Pricing type: Usage Based Billing
- Base Price: $0/month (कोई subscription fee नहीं)
- Billing frequency: Monthly
प्रत्येक meter पर Free Threshold को 0 पर set करें, ताकि execution का प्रत्येक second billable हो।
3
Send Usage Events
प्रत्येक model run पूरा होने पर Dodo Payments को usage event भेजें। प्रत्येक prediction को एक unique
event_id दें, ताकि retried event दो बार count न हो।4
Measure Execution Time Precisely
प्रत्येक model run का समय
performance.now() से मापें और billing के लिए second के निकटतम दसवें भाग तक round करें।5
Create Checkout
जब कोई user sign up करे, तो usage-based product के लिए checkout session बनाएँ। इसके बाद Dodo Payments प्रत्येक cycle में usage की billing करता है और invoices जारी करता है।
Time Range Ingestion Blueprint के साथ तेज़ी लाएँ
Time Range Ingestion Blueprint per-second compute tracking को सरल बनाता है। प्रत्येक hardware tier के लिए एक ingestion instance बनाएँ और प्रत्येक run के बादtrackTimeRange call करें।
durationMs के रूप में होता है। प्रत्येक meter का Over Property durationMs पर set करें, या per-second pricing बनाए रखने के लिए durationSeconds भेजें। प्रत्येक hardware tier के लिए एक ingestion instance के साथ, यह Replicate के multi-tier metering से सीधे map होता है।
Users के लिए Cost Estimation
Usage-based bills का अनुमान लगाना कठिन हो सकता है, इसलिए users को model run करने से पहले cost estimate दिखाएँ। Estimates unexpected bills को रोकते हैं और trust बनाते हैं।Example Cost Calculations
Cost Calculator बनाना
यह function tier की per-second rate को estimated run time से multiply करता है:Enterprise: Reserved Capacity
जिन customers को dedicated capacity और cold boots से मुक्ति चाहिए, उनके लिए Replicate deployments उपलब्ध कराता है: dedicated instances जिनके online रहने के पूरे समय billing होती है। Dodo Payments के साथ reserved capacity को model करने के लिए इसे subscription product के रूप में बेचें:- Product Type: Subscription
- Price: Fixed monthly price (उदाहरण के लिए, “Reserved A100 Instance - $500/month”)
- Billing Cycle: Monthly
Advanced: Heartbeat Metering
जो tasks कई minutes या hours तक चलते हैं, उनके लिए अंत में केवल एक event भेजना जोखिमपूर्ण है: यदि process crash हो जाए, तो usage data खो जाता है। इसके बजाय, task चलने के दौरान हर 30 से 60 seconds में usage event भेजें।उपयोग की गई प्रमुख Dodo सुविधाएँ
Usage-Based Billing
ऐसे products set up करें जिनकी billing consumption के आधार पर हो।
Meters
उन metrics को define करें जिन्हें आप track और bill करना चाहते हैं।
Event Ingestion
Usage data होने के साथ-साथ Dodo Payments को भेजें।
Subscriptions
Reserved capacity और enterprise plans के लिए recurring billing manage करें।
Time Range Blueprint
Duration helpers के साथ per-second compute tracking।