Nexcess.com Servers.com LiquidWeb.com
Back

Cloud repatriation comes to streaming: why media companies are pulling workloads back off hyperscale cloud

Last updated on September 03, 2026 10 min to read by Liam Waterson

IDC surveyed enterprise IT teams in 2024 and found that 80% expected to move at least some compute or storage back off public cloud within the year.

Streaming and media companies are a meaningful part of that number, and it's not hard to see why. Video infrastructure runs constantly, at a scale that's easy to predict most nights and impossible to predict on the nights that matter most. Hyperscale cloud is built to handle both. It's priced to handle both too, and that's where the problem starts.

What is cloud repatriation?

Cloud repatriation is the process of moving applications, data, or workloads from public cloud providers like AWS, Azure, or Google Cloud back to on-premises infrastructure, dedicated bare metal, or another hosting environment the company controls directly.

For a streaming platform, that usually means encoding, transcoding, origin storage, and packaging infrastructure. This is the part of the stack that runs every day, all year, regardless of what's trending or what's live. It's not the part that needs to scale to ten times its size overnight.

CDN and delivery are a different story, so let's go there briefly. CDN capacity is already running on bare metal in most cases, whether that's through a cloud-adjacent service like AWS CloudFront or a dedicated vendor like Akamai or Fastly. It was never really "on public cloud" in the way compute and storage are. Some platforms build and operate their own CDN instead of paying for delivery as a service, and it's often the more cost-effective option at scale. Either way, CDN isn't usually part of a repatriation decision, but it can be. Some platforms decide to bring delivery in-house, building and managing their own CDN on dedicated infrastructure. Either way, it's a separate build-vs-buy call that most platforms have already made, one way or another, before repatriation even comes up.

Why streaming platforms are moving workloads back

Hyperscaler prices keep climbing for workloads that never change

Most of a streaming platform's infrastructure spend goes toward capacity that runs at roughly the same level every day. Encoding pipelines don't get busier because a show trends on social media. Origin storage doesn't shrink on a slow Tuesday.

That steady, predictable load gets billed at the same elastic, on-demand rate as capacity that genuinely does need to scale on short notice, and hyperscaler prices have only gone up.

Egress is the line item that actually hurts

Video is bandwidth-heavy by nature, and every stream delivered out of the cloud gets billed for it. In our experience, egress and traffic costs on a hyperscaler can run roughly 10 times higher than the same load carried on dedicated bare metal. For a platform delivering video around the clock, that gap compounds fast.

The bill makes it hard to know what a viewer actually costs you

A cloud bill that moves with usage is fine when usage is genuinely unpredictable. When most of your usage is predictable and priced like it isn't, it gets harder to know your real cost per viewer, and harder to forecast margin with any confidence.

Colo and on-prem operators are feeling the same pressure

It's not only cloud-native platforms rethinking this. Colocation and on-premises operators are looking for managed partners too, as the capital cost of running their own hardware keeps rising. The direction of travel is the same either way: companies want infrastructure that's sized to what they actually run, not what they might.

The three ways companies actually do this

Repatriation isn't one move. In practice, companies take one of three approaches.

  1. Selective repatriation. Moving specific, high-cost workloads back while leaving the rest on public cloud. What makes sense depends on the platform, not the workload type: a 24/7 multi-channel platform has predictable encoding and origin load worth moving to bare metal, while a smaller operation streaming one event a week doesn't, even running the same components.
  2. Hybrid. Keeping steady, predictable workloads on dedicated infrastructure and using public cloud for capacity that genuinely needs to flex, like delivery during a live event or a traffic spike you can't fully forecast. This is the model most streaming platforms end up at, and it's the one we see working best.
  3. Full return. Moving everything off public cloud entirely. This is rare, and for good reason. It means giving up the instant, on-demand scale that public cloud is genuinely good at, which most streaming platforms need somewhere in their stack, even if it's not everywhere.

Where a full return off the cloud makes sense, and where it doesn't

A full return makes sense when a platform's traffic is well understood and genuinely stable. If you know what a normal night looks like, and your infrastructure needs don't change much from one month to the next, there's little reason to keep paying for elasticity you don't use.

It doesn't make sense in a few common situations.

  • Platforms still establishing a traffic pattern don't have anything reliable to size a move against yet.
  • Platforms with genuinely unplanned events and no lead time, like breaking news coverage, need the instant elasticity that only a hyperscaler can offer.
  • Platforms that still have meaningfully unpredictable spikes across the whole stack, not just around a handful of known events.

Hyperscaler-only vs. bare metal-only vs. hybrid

Hyperscaler-only Bare metal-only Hybrid (baseline on bare metal, burst to hyperscaler)
Cost predictability Low. Scales with usage, hard to forecast High. Flat, known cost High for baseline, metered for surge
Cost at steady-state load High. Paying elastic pricing for work that never changes Low Low
Cost during a spike High, but instant Poor. Can't absorb an unplanned spike without lead time Moderate. Surge capacity is reserved and activated as needed
Provisioning speed Instant Hours to days (pre-configured builds ready in under an hour; custom hardware can take 24+ hours) Baseline pre-provisioned; surge available on demand
Best fit Platforms with no established traffic pattern yet Stable, well-understood traffic with no need for sudden elasticity Known baseline with occasional, genuine surge events
Biggest risk Predictable load permanently priced like unpredictable load Underprepared for a real surprise spike Requires knowing your own traffic pattern well enough to size baseline correctly

What repatriation actually costs to get right

Repatriation isn't free, and it isn't instant. Let's be honest about what it takes.

Migrating live video infrastructure carries risk if it's rushed. Encoding, transcoding, and origin systems that have been running for years often have dependencies nobody's documented recently, and moving them without a phased plan can mean downtime you can't afford.

There's a skills gap too. Running your own dedicated infrastructure means someone on your team needs to know how to manage it, from hardware to network configuration. Teams that have spent years operating entirely on public cloud sometimes need to rebuild that muscle or a budget to seek external help.

And custom hardware takes time. Pre-configured builds can be ready in under an hour, but custom configurations can take 24 hours or more to provision. That's fine for baseline capacity you're planning months ahead of time. It's a poor fit for anything you need standing up mid-spike.

None of this is a reason to avoid repatriation. It's a reason to plan it as a phased move, not a weekend project.

Moving from all-in to hybrid

Picture two versions of the same platform's infrastructure. In the first, every workload, from encoding to delivery, runs on a hyperscaler, billed at the same elastic rate whether it's a quiet Tuesday or a live event.

In the second, the steady, predictable workloads (encoding, transcoding, origin, storage) run on dedicated bare metal at a flat, known cost, while only the capacity that needs to flex for a spike stays on the hyperscaler, activated when it's actually needed and scaled back down after.

The second version isn't cheaper because it uses less infrastructure. It's cheaper because each workload is running on infrastructure priced for what it actually does.

How to evaluate whether repatriation is right for your platform

  1. Pull your last 3 to 6 months of traffic and look for the average or the floor as your baseline, whichever makes more sense for your needs.
  2. Separate that baseline from your surge capacity. Encoding, transcoding, origin storage, and delivery infrastructure are the usual candidates for the baseline side.
  3. Price dedicated bare metal against your current on-demand spend for that specific baseline load, not your total cloud bill. The comparison only means something if it's apples to apples.
  4. Pilot with one workload first, whatever unit separates cleanly on your stack, rather than moving everything at once. This gives you a real read on the savings and the operational lift before you commit further.
  5. Reassess after a few months with real data from the pilot, and decide what moves next.

Cloud repatriation FAQs

Is cloud repatriation the same as hybrid cloud?

Not really. Repatriation is the act of moving a workload back off public cloud. Hybrid cloud is often the result: a mix of dedicated infrastructure for predictable load and public cloud for capacity that needs to flex.

Do most companies fully leave the cloud?

No. Full repatriation is rare. Most companies that repatriate do it selectively, moving specific high-cost workloads back while keeping the public cloud in place for the parts of their infrastructure that genuinely need to scale on demand.

How long does it take to repatriate a workload?

Plan for months, not hours. The infrastructure itself can be quick to stand up (a pre-configured bare metal build can be ready in under an hour, though colocation timelines run differently and depend on your provider and site), but that's a small part of the total move.

Most of the timeline goes into planning: mapping dependencies, sizing your real baseline, testing the migration path, and running workloads in parallel before you cut over. For a live video pipeline, that planning and validation work is what actually determines your timeline, not how fast the hardware can be turned on.

What usually goes back first?

Encoding, transcoding, and origin storage are the most common starting points, when they run at a steady, predictable level for that platform. Whether they do depends on the platform itself, its viewership and channel type, not the components.

If those components aren't steady for your platform, that's usually a sign it's too early to repatriate anything.

Getting started with cloud repatriation

Cloud repatriation isn't about leaving public cloud behind. It's about putting each workload on the infrastructure priced for what it actually does, so your predictable, steady-state load stops paying elastic prices it never needed.

Start by looking at how many channels you're running 24/7 and what your concurrent viewership actually looks like on a normal night, then price dedicated infrastructure against that specific baseline rather than your whole bill. That single comparison tells you more about whether repatriation makes sense for you than any industry benchmark will.

If the numbers look worth pursuing, this is exactly the kind of move we help streaming and media companies plan. servers.com sits between hyperscale cloud and traditional bare metal, built for platforms that want dedicated infrastructure for their baseline without giving up the ability to burst when they need to. Take a look at what that setup could look like for your traffic on our streaming industry page.

PS - I'll be at IBC2026 with the servers.com team, having this same conversation. If you plan to be there, come find us at booth 5.A81 for a coffee and a chat!

About the author

Liam Waterson, Sales Development Representative at Servers.com by Nexcess

Liam Waterson, Sales Development Representative

Liam Waterson is a Sales Development Representative at servers.com by Nexcess, where he works primarily with streaming businesses to balance growth, performance, and profitability. Outside of work, he is a big football fan and an average Shrek impersonator.