Running costs and infrastructure
Pricing AI Video Plans Against Provider Costs
Short answer
To price AI video plans, find your real cost per output minute from your provider invoices, add retries, storage, payment fees and support, then set each plan's allowance so worst-case cost stays below price with headroom. Enforce the allowance on the server before a job is queued, so no customer can spend more than the plan allows.
Key takeaways
- Measure cost per output minute from invoices across real jobs, including retries and failures; a provider's list price is only a starting point.
- Provider cost is one line of five: storage, delivery, payment fees, support and retries sit on top.
- Price in credits that carry weights for premium avatars, cloning and translation, so one allowance covers every kind of job.
- Compute each plan's worst case from its cap, not from average use, and keep margin positive even then.
- Check limits on the server before queueing; overage is then prevented, not billed afterward.
On this page 9 sections
AI video generation cost per minute is the number your whole pricing rests on. Every video your customers make consumes minutes on a provider account you pay for, and that bill grows with use while your subscription price stays fixed. Set each plan's allowance from your measured cost per output minute, add the other costs on top, and make the server refuse jobs beyond the allowance.
This is an operator worksheet. It builds the unit economics step by step with labelled invented numbers, because we do not state any provider's price. Get a current quote from each provider you use, put the date on it, and rerun the tables. If you run a white-label HeyGen clone, the plan limits described here map onto settings the platform checks before it queues a job. For how the pipeline spends those minutes, read how an AI avatar video generator works.
Know your cost per output minute
Provider pricing pages and invoices use different units. Some bill by seconds or minutes of generated video, some by credits, some by a plan with included minutes and overage. Convert everything to one unit: cost per minute of finished output video, the minute your customer actually receives. Nothing else compares across providers.
How to measure it
- Pick a test set. Twenty or more jobs that look like real customer work: different script lengths, avatars, voices and languages, with a share of premium options.
- Run them through each provider you are considering. Use production accounts, not free demos, since demo behavior and limits can differ.
- Add up the invoice. Take the total provider charge for the test period, including jobs that failed and jobs you re-ran.
- Divide by finished output minutes. Do not divide by input script length or by minutes attempted. Divide by minutes that reached a customer.
- Repeat per tier. Standard avatar, premium avatar, voice cloning and each translated language each get their own number.
- Date it. Write the date and the provider plan next to the number. Prices in this field move.
Our platform records each generation attempt with the provider and the settings used, which makes the reconciliation easier: you can match provider charges to attempts instead of guessing. The HeyGen clone platform lets you change providers from admin settings, so you can run the same test set through two providers and compare, then keep the one that gives acceptable quality at a cost you can sell.
Invented starting numbers
| Item (invented, per finished output minute) | Standard avatar | Premium avatar |
|---|---|---|
| Avatar render | 0.60 | 1.60 |
| Voice synthesis | 0.10 | 0.10 |
| Provider subtotal | 0.70 | 1.70 |
| Translated minute, per extra language | 0.75 | 1.75 |
The premium option costs about 2.4 times the standard one in this example. That ratio becomes the credit weight later.
Cost lines beyond the provider
The provider invoice is the biggest recurring cost but not the only one. Add each line to get a loaded cost.
| Cost line | What drives it | How to bound it | Invented figure |
|---|---|---|---|
| Retries and failures | Provider errors, bad inputs, re-runs a customer requests | Track failed share of jobs; cap re-runs per video | 8% extra on provider cost |
| Storage | Every finished video you keep | Retention rule per plan; purge schedule | 0.03 per output minute, monthly average |
| Delivery bandwidth | Downloads and playback | Export limits; presets | 0.02 per output minute |
| Payment fees | Every plan charge | Annual plans reduce fee count | 3% plus 0.30 per charge |
| Support | Tickets, onboarding, abuse handling | Docs, status page, plan tiers | 2.00 per small workspace per month |
| Hosting and monitoring | Servers, database, Redis, alerting | Fixed monthly cost spread over customers | Treat as overhead, not per job |
Loaded cost per standard credit in the invented example: provider 0.70 times 1.08 for retries is 0.756. Add 0.03 storage and 0.02 delivery, and the loaded cost is 0.806, which we round to 0.81. Payment fees and support are per plan, not per job, so they come in the plan table. For the lines every operator forgets, read hidden running costs of a creator platform.
Storage deserves a decision before launch. A customer who cancels but leaves a large archive still costs you storage, so decide how long files are kept after a subscription ends, and write the rule into the plan terms. Our HeyGen clone development cost page covers storage modes and retention work as running costs you carry.
Setting plan allowances
Do not sell "minutes" if your jobs have different costs. Sell credits with weights, so every job type draws from one allowance and the worst-case cost per credit stays flat.
| Job type (invented weights) | Provider subtotal | Credits per output minute | Loaded cost per credit |
|---|---|---|---|
| Standard avatar video | 0.70 | 1.0 | 0.81 |
| Premium avatar video | 1.70 | 2.5 | 0.75 |
| Translated minute, per language (standard) | 0.75 | 1.0 | 0.86 |
The translated row shows why weights need a check: 0.75 times 1.08 plus 0.05 storage and delivery is 0.86 per minute, so a translated minute at one credit costs more than a standard minute. Either weight it at 1.1 credits or accept the small difference. Use the highest loaded cost per credit, here 0.86, as your worst case for plan math, which is conservative.
Now set the plans. Margin is price minus payment fee minus support minus allowance times worst-case cost.
| Plan (invented) | Price per month | Included credits | Worst-case provider cost at 0.86 | Payment fee | Support | Margin | Margin % of price |
|---|---|---|---|---|---|---|---|
| Starter | 29.00 | 10 | 8.60 | 1.17 | 2.00 | 17.23 | 59% |
| Pro | 99.00 | 40 | 34.40 | 3.27 | 4.00 | 57.33 | 58% |
| Team | 299.00 | 150 | 129.00 | 9.27 | 10.00 | 150.73 | 50% |
Check one row by hand. Pro: 99.00 minus 34.40 minus 3.27 minus 4.00 gives 57.33. That is the margin when the customer spends every credit. Because the allowance is a hard ceiling, the margin cannot fall below that figure no matter what the customer does. Average customers use less, and unused allowance is extra margin.
The invented margins here are high because the example ignores fixed overhead, sales cost and your own time. Subtract hosting, monitoring and staff before deciding the price is comfortable. A plan is safe when margin stays positive after overhead at your expected customer count.
Price against value, not only cost
Cost sets the floor. The ceiling is what the customer would pay to get the same video another way: a freelancer, a studio, or a hosted competitor. Look at what your target buyers pay today, then place your plan between your floor and their alternative. For who the buyers are, read who buys a white-label AI video platform. For the revenue models these plans fit into, see the HeyGen clone business model.
Limits enforced before the job
An allowance only protects you if the server enforces it. If limits are checked in the browser, anyone can bypass them with a modified request. If they are checked after a job runs, you have already paid the provider.
| Control | What it caps | Where to enforce |
|---|---|---|
| Credits per period | Total spend per workspace | Service layer, before the job is queued |
| Concurrent jobs per workspace | Burst spend and worker load | Queue and service layer |
| Maximum video length | Cost of a single job | Validation on submit |
| Premium avatars, cloning, high resolution | Expensive options | Plan entitlements |
| API rate limits and key quotas | Automated abuse | API gateway and metering |
| Re-runs per video | Retry cost | Service layer |
In our platform, generation limits are checked on the server in the service layer before a job is queued, and export entitlements gate resolution and format by plan, so overage is prevented rather than invoiced afterward. API keys are scoped to a workspace and usage is recorded per endpoint per day, which supports a metered developer tier. Queue behaviour matters too: job-queue libraries such as Bull offer concurrency control, priorities, retries and rate limiting, which are the building blocks for the concurrency and re-run caps above.
Two controls beyond the base limits are worth planning. First, a warning at 80% of the allowance, by notification, so customers upgrade before they hit the wall. Second, a platform-wide spend alert on your own provider accounts, so a bug or an abuse pattern that slips through shows up in hours, not at the end of the month. We can set up per-customer budget dashboards beyond what the admin console shows.
Stress test: what if provider prices move
Provider prices and models change often, so test your plans against a rise before it happens. Take the Pro plan from the table, with a worst-case provider cost of 34.40 at full use, and raise the provider cost line.
| Provider cost change (invented) | Worst-case provider cost | Pro margin | Margin % of price |
|---|---|---|---|
| No change | 34.40 | 57.33 | 58% |
| Up 25% | 43.00 | 48.73 | 49% |
| Up 50% | 51.60 | 40.13 | 41% |
| Up 100% | 68.80 | 22.93 | 23% |
Even a doubling leaves this plan profitable before overhead, which is the sign of a plan with real headroom. If your own table goes negative at a 50% rise, cut the allowance or raise the price now. Keep a rule for what you do when cost moves: reprice new customers at once, give existing customers notice, and check your terms for what notice you must give. A second provider you have already tested is your best protection, because you can move traffic without waiting for a developer.
Annual plans and refunds
An annual plan collects cash early and cuts payment fees, but it also commits you to twelve months of provider exposure. Say Pro is sold for ten months' price, 990 up front. Release the 40 credits month by month, not all at once, so a customer cannot burn the year's allowance in a week and then ask for a refund. Payment fees fall from twelve charges of 3.27, which is 39.24, to one charge of 30.00, a saving of 9.24 in this example. Margin at full use is 990 minus 412.80 of provider cost minus 30.00 of fees minus 48.00 of support, or 499.20. Write the refund rule before you sell the plan: what happens to the unused months, and whether credits already used are refundable.
Worksheet
Fill this in with your own numbers. Do not skip a step.
- Provider quote per tier. For each provider, write the current cost per finished output minute for standard avatar, premium avatar, cloning and a translated minute. Date each.
- Retry rate. Failed or re-run jobs divided by total jobs, from your test set. Multiply provider cost by one plus that rate.
- Storage and delivery per output minute. Average file size times your storage and bandwidth rates, spread over the retention period.
- Loaded cost per credit. Provider cost times the retry factor, plus storage and delivery, divided by credits per minute. Take the maximum across job types.
- Payment fee per plan. Percentage times price plus any flat fee, from your processor's terms.
- Support allowance per plan. Expected support hours times hourly cost.
- Plan allowance. Choose credits per plan, then compute worst-case margin as price minus fee minus support minus credits times loaded cost.
- Overhead check. Subtract fixed monthly costs divided by expected customers. Margin must stay positive.
- Overage price. Set a price per extra credit at least as high as the plan's effective price per credit, plus margin.
- Limits. Set concurrency, maximum length, entitlements and API quotas to match the plan.
- Review date. Put a monthly reminder to re-measure cost and reprice.
Overage and the developer tier
When a customer uses up the allowance, they either wait for renewal, upgrade or buy extra credits. Price extra credits above plan credits, or customers will skip the upgrade. In the invented example, Pro's effective price per credit is 99.00 divided by 40, or 2.475. An overage credit at 3.00, less a percentage payment fee of 0.09 (ignoring any flat fee) and a worst-case cost of 0.86, leaves roughly 2.05 of margin per credit, or 68%.
Metered API access is the same logic with a meter. Stripe describes usage-based billing and billing credits for prepaid or promotional usage, which are the standard tools for charging by use and granting prepaid or free credits. Note that Stripe's billing credits are for spending on your own products and services, not for stored value or third-party spending, so check what each tool allows before you design around it. Our platform models plans and a Stripe-compatible flow, and moving real money requires your own Stripe account and webhook wiring.
Free tier and trial design
A free tier is a marketing cost, so price it like one. Compute it and cap it.
| Trial calculation (invented) | Value |
|---|---|
| Free credits per verified workspace, once | 3 |
| Worst-case cost per trial at 0.86 per credit | 2.58 |
| Trial sign-ups in a month | 100 |
| Total trial cost | 258.00 |
| Conversion to Starter | 5 customers |
| Trial cost per paying customer | 51.60 |
| Starter margin per month (from the table above) | 17.23 |
| Months to recover trial cost | about 3.0 |
If recovery takes longer than your typical customer stays, the trial is too generous or the conversion too low. Fix it by cutting trial credits, raising the quality of sign-ups or improving onboarding, in that order.
Protect the trial from abuse. Require email verification before the first job, optionally two-factor, and give one trial per workspace and per payment identity. Block voice cloning and premium avatars on trial, since they cost most and carry the highest consent risk. Cap the trial's video length. A scripted sign-up attack against an open trial could cost thousands in one night; a hard cap and verification make it a small, known number. Our platform supports email verification and optional two-factor at sign-up, and limits are checked server-side.
Consent and policy checks also belong on any plan that allows cloning or real-person avatars; see AI avatar consent, likeness and misuse.
What the platform price covers
Keep two budgets apart in your head. The platform price is the published, one-time price for the software, full source code, 60 days of technical support and a year of updates. It is not a running cost. The running costs are provider usage, storage, delivery, payment fees, hosting, monitoring and support, and they recur and scale with customers.
- One-time: the software package, rebranding and deployment. See our pricing for the published figure.
- Recurring and variable: provider minutes, storage, delivery, payment processing.
- Recurring and fixed: hosting, database, Redis, monitoring and staff time.
- Yours to arrange: provider accounts and approvals. We connect your providers; we do not supply render minutes or a bundled avatar and voice catalog.
We take no share of what your customers pay you. That means your margin per plan is the one you compute here, with no revenue share to subtract. It also means the cost discipline is yours.
What to do next
Run the measurement first. Pull twenty representative jobs through your candidate providers, compute the cost per finished output minute and date it. Then build the plan table, test the worst case and fix the limits. Only after the numbers hold should you publish prices.
When the table works, review the HeyGen clone features to confirm each limit you need exists, and keep a monthly review on the calendar, since models and prices change quickly. This article is general information, not financial or legal advice; take your terms of service, refund rules and tax treatment to a qualified adviser before you charge customers.
Questions and answers
Should I price per minute or per credit?
Per credit usually works better. A credit carries a weight, so a standard avatar minute costs one credit and a premium avatar minute costs more. That lets one plan allowance cover different jobs while keeping your worst-case cost per credit steady. Per-minute pricing forces you to publish a separate rate for every avatar tier, voice and language.
How do I stop overspend?
Check the plan limit on the server before a job is queued, cap concurrent jobs per workspace, rate-limit API keys, and gate expensive options such as cloning and premium avatars to higher plans. Add an alert when a workspace nears its allowance. The cheapest place to stop a bad job is before it reaches the provider.
Do customers bring their own AI accounts?
In our product the provider settings are platform-level, set by you in the admin console, so customers use your provider accounts and you pay the provider. We can set up per-customer provider keys for your build, with the scope confirmed at kickoff. A bring-your-own-key model changes your pricing, since you would charge for the platform alone.
Who pays for failed renders?
Decide the rule in advance and publish it. Most operators do not charge the customer for a failed job and absorb the provider cost if the provider billed for it. Check your provider's terms, because some bill for failed or retried calls. Our platform records every attempt with its provider and settings, which gives you evidence for provider disputes.
Can I switch providers when prices change?
Yes for providers supported by the integration screens, which you edit at runtime with Save and Test buttons and no redeploy. Adding another provider is tailored work we do for your build. Re-measure your cost per output minute after any switch and adjust allowances before margins fall.
How big should a free trial be?
Small and fixed. Give a few credits once per verified workspace, block cloning and premium options, and compute the cost of trials per paying customer. If trial cost per conversion exceeds a few months of plan margin, shrink the trial or tighten sign-up checks before you spend on marketing.
Sources
- Stripe Docs: Basic usage-based billing
- Stripe Docs: Billing credits
- Bull: Premium queue package for handling distributed jobs (GitHub)
Checked in October 2026. Rules, fees and programme terms change; confirm on the source before you rely on them.
Independence note. GetFame is an independent software company. HeyGen is a trademark of its owner and is named here only to describe a category of platform. GetFame is not affiliated with, sponsored by or endorsed by HeyGen.
Keep reading
How Does an AI Avatar Video Generator Work?
How does an AI avatar video generator work? Follow the pipeline from script to voice, avatar, lip sync, render and translation, stage by stage.
Who Buys a White-Label AI Video Platform? Buyer Types
How to start an AI video business: the buyer types that run one, what they need before launch, business models, niches and who should not start at all.
Hidden Running Costs of a Creator Platform After Launch
The monthly bills a creator platform carries after launch, as cost drivers and formulas: hosting, video delivery, payment fees, moderation, plus a worksheet.