Video and streaming infrastructure
Scaling Video Delivery: CDN, Storage and Transcoding Choices
Short answer
To scale a video streaming app, treat each stage separately: upload, transcode, store, package, deliver, play. Transcode once per upload, keep hot renditions in object storage, put a CDN in front with long cache lifetimes for segments, and watch delivery first because it follows minutes watched. Change a stage only when a measured threshold says so.
Key takeaways
- A clip passes through six stages, and each stage has its own cost driver and its own failure mode.
- Delivery is usually the largest recurring line because it follows minutes watched, not registered users.
- Transcode once at upload, keep the renditions you actually serve, and archive or delete the rest on a schedule.
- A CDN works through cache hits, so segment URLs that never change and long cache lifetimes matter more than which vendor you pick.
- Short clips, long videos and live streams need different settings, so one default does not fit all three.
- Leave the default setup when a number you track crosses a threshold you wrote down in advance, not when the bill feels high.
On this page 10 sections
To scale a video streaming app, stop thinking of "video" as one thing. A clip moves through upload, processing, storage, packaging, delivery and playback, and each stage has a different cost driver and a different way to fail. Scale the stage that is actually under pressure, and leave the others alone.
This is the architecture and decision guide. It follows a clip from upload to play, shows where cost and delay come from, and says when to change each part. If you are starting from a white-label TikTok clone, you already have a pipeline, and the question is what to tune as traffic grows. For the concept behind renditions and HLS, read the explainer on what video transcoding and adaptive bitrate are. For formulas to price each stage, use video hosting cost for a video sharing site.
The life of an uploaded clip
- Upload. The app sends the file to storage, ideally straight to an object-storage bucket using a short-lived signed address, so your application servers do not carry the bytes. Failure mode: slow or abandoned uploads on weak mobile connections. Fix: resumable uploads and clear size limits.
- Queue. The upload creates a job on a queue. Failure mode: a long queue after a burst, so creators wait. Fix: more workers, priority for short clips, a visible status.
- Transcode. A worker re-encodes the source into renditions and makes thumbnails. Failure mode: a bad source file crashing a worker. Fix: inspect before processing and reject early.
- Store. Renditions and segments go to object storage. Failure mode: storage growing faster than the catalog suggests because every rendition is kept. Fix: a ladder you can justify and a retention rule.
- Package and publish. The system writes manifests and marks the video live. Failure mode: a video marked live before every rendition exists. Fix: publish only when the set is complete, or publish a low rendition first.
- Deliver. A CDN serves segments to viewers, fetching from storage on a miss. Failure mode: low cache hit ratio, slow starts in far regions. Fix: cache settings and regional placement.
- Play. The player reads the manifest, picks a rendition and switches as the connection changes. Failure mode: stalls and slow first frame. Fix: a good low rung, short first segments and a tested player.
On our TikTok clone script, uploads are converted by FFmpeg into several resolutions with a thumbnail, written to secure bucket storage and served through Cloudflare or Akamai delivery with token-based stream links, deployed onto your own cloud account. That is the shape of the default setup this guide helps you tune.
Where the money goes
This table is relative on purpose. Real amounts depend on your quotes, so use the formulas in the cost guide to turn each row into money.
| Stage | What you pay for | Grows with | Typical share of the video bill | First lever |
|---|---|---|---|---|
| Transcoding | Processing time per uploaded minute and rung | Uploads | Small to moderate, spiky | Fewer rungs, better presets for popular videos |
| Storage | GB-months for renditions and sources | Catalog size | Moderate, steady | Retention rules, cooler tiers for sources |
| Delivery (egress) | GB sent to viewers | Minutes watched | Usually the largest | Cache hit ratio, bitrate, codec |
| Live ingest and processing | Stream hours and real-time encoding | Live events | Small, until peaks | Limits per creator, scheduled events |
| Application and database | Servers and queries | Active users | Small next to video | Caching, read replicas when needed |
The thing to remember is that delivery follows what people watch, so a successful week raises it immediately, while storage and processing rise with what creators post. Set alerts on delivery first. Our TikTok clone development cost page separates the one-time software price from these running bills, and hidden running costs of a creator platform lists the rest of the monthly lines.
When to transcode, and how much
Transcoding is a one-time cost per upload, so the decisions are about which renditions to make and when.
| Strategy | How it works | Good when | Risk |
|---|---|---|---|
| Full ladder at upload | Make every rung before publishing | Uploads are modest and quality matters | Creators wait longer, and unwatched videos still cost processing |
| Low rung first, rest later | Publish a small rendition quickly, add the others in the background | Short clips where speed to publish matters | Early viewers see lower quality |
| On demand | Make a rung the first time someone asks for it | A long tail of rarely watched videos | First viewer waits, more complexity |
| Tiered by popularity | Basic ladder for all, extra rungs and a better codec for videos that prove popular | Large catalogs with a few hits | Needs usage data and a re-encode job |
A good default for a young platform is a full small ladder at upload, kept short, with a re-encode path for hits. Add the more complex strategies once you have data showing a long tail of unwatched uploads.
Where transcoding runs
There are three common setups: your own queue and workers, a managed cloud transcoding service, and a video platform API that also handles packaging and delivery. The trade-off is engineering effort against price per minute, covered in the explainer. For the other lead products, the ready-made ReelShort-style platform offers adaptive HLS streaming, set up for your build, that runs through Mux or Cloudflare Stream on your own account, which is an example of the managed route. See the ReelShort clone for how episode delivery, mostly bandwidth, is set up. Move from your own workers to a managed service when queue delay hurts creators or uploads are too bursty to size for.
Adaptive bitrate and rendition ladders
Adaptive bitrate streaming lets the player change rendition as conditions change. IETF RFC 8216 describes HLS as a master playlist of variant streams, with clients expected to switch between them to adapt to the network. The scaling question is how many rungs and how far apart.
- Too few rungs: viewers on weak connections have nothing small enough to play, and viewers on good ones get worse pictures than they could.
- Too many rungs: storage and processing multiply, and rungs close together make little difference to viewers.
- Top rung too high: delivery cost climbs for viewers who cannot see the difference on a phone screen.
Start from your audience's devices and networks. If most viewing is on phones, a ladder that stops at 720p may be enough. Check the actual rendition mix your viewers use after a month and drop rungs nobody picks. A better codec shrinks delivered bytes, and MDN's codec guide notes that newer royalty-free codecs such as AV1 compress considerably better than H.264, though device support has gaps, so add one as a second option after H.264, not instead of it.
CDN choices and caching
A content delivery network keeps copies of your segments in many locations, so most requests are answered nearby and never reach your storage. In HTTP terms, that is a shared cache. MDN describes managed shared caches, including CDNs, as caches deployed to serve many users, and RFC 9111 defines a shared cache as one that stores responses for reuse by more than one user.
What makes a CDN effective
- Immutable segment URLs. A segment never changes after it is written, so give it a long cache lifetime. MDN's caching guide recommends versioned URLs with a long
max-ageand theimmutabledirective for content that does not change. - Short lifetimes for playlists that change. A live playlist is rewritten every few seconds, so cache it briefly or not at all. A video-on-demand manifest can be cached longer.
- A single cache key. If tokens or tracking values are part of the URL, every viewer gets a different key and the cache never hits. Put access tokens where the CDN can validate them without splitting the cache, or use signed cookies where supported.
- An origin shield or tiered cache. Many CDNs can funnel misses through one cache layer, so storage sees one request instead of one per location.
- Regional reach. A CDN is only as good as its presence where your viewers are. Check coverage in your target countries and test from a real device there.
Protecting paid or private video
Signed or tokenized links limit hotlinking and sharing of paid content. Our TikTok-style product uses secure token-based stream URLs for this. Test that your token scheme does not break caching, because that is the most common way to pay for a CDN and get no benefit from it.
One CDN or several
A single provider is simpler. A second provider adds resilience and negotiating power, at the cost of more configuration. Most platforms should start with one, own the account, and add a second only when an outage or a contract point justifies the work.
Storage tiers and retention rules
Object storage is the right home for video files and segments. The design decisions are which copies to keep and where.
Amazon's S3 documentation shows the general pattern: classes for frequent access, for infrequent access with a retrieval fee and a minimum storage period, and for archives that need a restore step before reading. Other providers offer equivalents under different names and terms. Use the pattern, then read your provider's current rules.
| Data | Access pattern | Suggested tier | Rule |
|---|---|---|---|
| Renditions of recent and popular videos | Read constantly | Hot | Keep behind the CDN |
| Renditions of old, rarely watched videos | Read occasionally | Hot or infrequent access, once the retrieval fee is understood | Review after 90 days of low views |
| Original source files | Almost never | Cooler or archive | Keep only if you plan to re-encode |
| Failed and abandoned uploads | Never | Delete | Auto-delete after a short window |
| Deleted videos | Never | Delete | Remove all renditions, thumbnails and manifests |
The rules here, such as 90 days, are examples to adapt. Write your own into your creator terms, so nobody is surprised when an old upload moves tier or is removed. On large catalogs, let a lifecycle rule do the moving, not a person.
The OnlyFans-style product uses S3-compatible object storage with CDN-ready delivery for premium media, and scales by moving media to larger object storage behind a CDN and adding queue workers. See the white-label OnlyFans clone for that fan-platform example, where media is private and paid, so access control matters more than raw reach.
Short video versus long form versus live
| Pattern | Short vertical clips | Long-form on demand | Live |
|---|---|---|---|
| Typical length | Seconds to a minute or two | Minutes to hours | Open-ended |
| Ladder | Few rungs, small, fast start | Fuller ladder, higher top | Real-time ladder, fewer rungs |
| Segment length | Short first segment | Standard | Short, to cut delay |
| Caching | Very effective, clips replay | Effective for popular titles | Weak at first, playlist changes |
| Prefetch | Yes, next clips load ahead | Rarely | No |
| Main cost | Delivery from heavy scrolling | Storage and delivery | Peaks and real-time processing |
| Main risk | Slow first frame loses the viewer | Buffering mid-video | Delay and outages |
Short video apps lean on preloading the next clip so swipes feel instant. That raises delivery, because some preloaded clips are never watched, so limit how many you prefetch and how much of each. Our guide to starting a short video app covers the product side. Live needs its own limits, alerts and fallback, because the newest segments are still being produced while viewers watch. MDN's guide to live streaming explains the segment and manifest model behind this. For how live income relates to the bill, see how virtual gifts work in live streaming.
Monitoring and thresholds
You cannot tune what you do not measure. Track these, weekly at first.
- Delivery volume in GB, and per viewer-hour. A rising per-hour figure means bitrates are climbing or the cache is missing.
- Cache hit ratio at the CDN, and origin egress.
- Time to first frame at the 50th and 95th percentile, by region and device class.
- Rebuffering rate, the share of playback time spent stalled.
- Rendition mix, which rungs viewers actually play.
- Queue depth and processing delay, from upload to published.
- Storage growth per week, and the share belonging to unwatched videos.
- Failed uploads and failed jobs as a share of total.
- Live peak concurrency against your tested limit.
For each metric, write a threshold that triggers a decision. Example thresholds, invented for illustration: processing delay above 10 minutes for a typical clip triggers more workers or a managed service; a hit ratio below 80% triggers a review of URLs and headers; the 95th percentile first frame over 3 seconds in a region triggers a regional check; storage from unwatched videos above a third of the total triggers a retention rule. Your numbers will differ. What matters is that the decision is written down before the pressure arrives.
When to leave the default setup
A shipped default is a good starting point and not a promise for every scale. Leave it when a tracked number crosses its threshold, and make the smallest change that fixes it.
| Signal | Likely cause | Change |
|---|---|---|
| Creators wait too long to publish | Queue too small or bursty uploads | More workers, low rung first, or a managed service |
| Delivery bill growing faster than watching | Cache misses, high bitrate, over-prefetch | Fix cache keys and lifetimes, trim the ladder, limit prefetch |
| Slow starts in one region | Weak CDN presence there | Test and, if needed, add a second CDN or regional setup |
| Storage cost rising | Sources and unwatched renditions kept | Retention and tier rules |
| Live streams fail at peaks | Concurrency over tested capacity | Load test, add capacity, cap audience per stream |
| Data must stay in a country | Legal or contract requirement | Regional infrastructure, a separate scoping job |
On our side, delivery runs through a global CDN setup, and dedicated regional infrastructure or very high concurrency targets are a separate scoping conversation, not something to assume. Decide where your data lives with data ownership and hosting choices in mind.
A growth path in three steps
Most platforms move through the same three stages, and it helps to know which one you are in.
- Prove the product. One cloud account, one bucket, a CDN in front, a small ladder, workers on the same provider. Spend effort on retention, not infrastructure. Your main job is to avoid the mistakes that make later change painful: tokens that break caching, no retention rule, sources kept forever.
- Hold the line as traffic grows. Add alerts, raise the cache hit ratio, trim the ladder to the rungs viewers use, and start a retention job. Move transcoding to a managed service if queues grow. This stage usually saves more money than any new technology.
- Optimize for scale. Add a better codec for hits, regional placement, a second delivery provider for resilience or price, and negotiated committed-use pricing. Each of these needs data from the earlier stages to justify it.
Skipping from the first stage to the third is the common waste: a multi-provider design built for traffic that never came. Move only when a written threshold says so, and re-run your cost worksheet with new quotes before and after each change, so you know whether the change paid for itself.
What to do next
- Draw your own six-stage pipeline and write the owner, the tool and the cost driver beside each stage.
- Set up the nine metrics above, with alerts, before launch.
- Write a threshold and a response for each metric.
- Test caching with a real device: check that segment responses carry long lifetimes and that tokens do not split the cache.
- Get quotes at your expected volume and ten times that, using the formulas in the cost guide.
- Review the setup each quarter, and re-read provider terms.
If you want to see how the delivery layer is set up on a ready-made product, read TikTok clone features or talk to us about your expected traffic.
Questions and answers
Do I need my own CDN?
You need a CDN, but not necessarily your own contract. Many setups put a managed CDN account, owned by you, in front of your storage. The decision is about who holds the account and who can see the bill. Own the account so you can change providers, set alerts and negotiate when volume grows.
Is object storage enough?
For storing video files and segments, yes. Object storage is the normal home for them. It is not enough for delivering them at scale to many viewers in many regions, because every request goes to one place. Add a CDN in front for delivery, and a queue and workers for processing.
How do I reduce egress?
Raise the cache hit ratio with long lifetimes on segments, lower the average delivered bitrate with a sensible ladder and a better codec, cap the top resolution, and avoid delivering video you do not need to, such as autoplaying many clips off screen. Measure first, because the biggest saving depends on where your traffic actually goes.
Is live harder to scale than video on demand?
Yes, in three ways. Live segments are new, so there is little cache benefit at first. Viewers arrive at the same moment, so peaks are sharper. And processing must keep up with real time. Plan live separately, with its own limits, alerts and a tested fallback if the stream fails.
When should I move transcoding to a managed service?
When uploads are bursty, your team is small, or queue delays are hurting creators. A managed service scales on demand and removes worker maintenance, at a higher price per minute. If volume is steady and someone can run the workers, your own queue is often cheaper. Compare quotes at your expected and your viral volume.
Where does the platform's video pipeline stop and my cloud bill start?
The software includes the application and the video pipeline it ships with. The cloud account, storage, delivery network and any managed video service are billed to you by their providers and grow with usage. We deploy onto your account, so you own the environment. Our development cost page explains what the one-time price covers.
Sources
- MDN: HTTP caching
- IETF RFC 9111: HTTP Caching
- IETF RFC 8216: HTTP Live Streaming
- Amazon S3: Understanding and managing storage classes
- MDN: Live streaming web audio and video
- MDN: Web video codec guide
- Apple Developer: HTTP Live Streaming
Checked in October 2026. Rules, fees and programme terms change; confirm on the source before you rely on them.
Independence note. GetFame is an independent software company. TikTok is a trademark of its owner and is named here only to describe a category of platform. GetFame is not affiliated with, sponsored by or endorsed by TikTok.
Keep reading
Video Hosting Cost for a Video Sharing Site, Item by Item
Video hosting cost for a video sharing website as formulas: storage GB-months, delivery GB, transcode minutes and live. Fill them with your own quotes.
What Is Video Transcoding and Adaptive Bitrate Streaming?
What is video transcoding? A plain explanation of codecs, renditions, HLS, DASH and adaptive bitrate, and what each choice means for quality, cost and speed.
Hidden Running Costs of a Creator Platform After Launch
The monthly bills a creator platform carries after launch, as cost drivers and formulas: hosting, video delivery, payment fees, moderation, plus a worksheet.