Accelerator Access Is the New Availability Zone

in #technology • 5 days ago

Accelerator Access Is the New Availability Zone

Rows of operational server racks inside a modern cloud data centre facility

For about a decade, comparing the major clouds meant comparing capabilities. One had a managed service the others lacked. One had better data tooling. One had the enterprise relationship.

That era ended quietly. All three now offer a credible managed database, a credible serverless runtime, a credible container platform, and core compute pricing inside a band narrower than the cost of designing your system badly.

So what is left to compare?

Supply.

The question that is not on any grid

The biggest practical difference between the three providers in 2026 is whether you can actually get accelerators — in your region, this quarter, at a commitment level you can live with.

That is not a product question. It varies by region, by instance family, by commitment size, and by what the provider's other customers happen to be doing this quarter. Which is exactly why it never appears in a pitch and is never volunteered.

Each provider's posture differs in a way worth understanding:

Amazon offers the widest range plus first-party silicon that is meaningfully cheaper per unit of throughput for workloads that fit it. That qualifier is the whole story — porting is genuine engineering work, not a flag. Teams that budget for the port get lasting savings. Teams that assume it is a checkbox find out in week three.

Google offers its own tensor hardware alongside commodity parts, with a stack that is unusually good on the frameworks it optimises for and noticeably rougher off them. The gap between the happy path and everything else is wider here than elsewhere. On the happy path, the best price-performance of the three.

Microsoft offers a very large commodity fleet plus privileged access to frontier model capacity through commercial relationships. If the plan is to consume a frontier model rather than train anything, that access — with the data-residency guarantees around it — is frequently decisive whatever the compute comparison says.

What to actually ask

Not what they offer. Ask:

Which accelerator families are available in my primary region right now, and what is the lead time for each? Does that change at one year of commitment? At three? If I need to scale fourfold in six months, is that path contractual or best-effort? What is the fallback region, and what does latency and egress look like from there?

A field engineer who answers concretely is telling you something real. One who redirects to a capability overview is telling you something too.

Two other questions that still discriminate

How flexibly can a committed-use discount be reallocated? Your workload shape will change inside eighteen months. A commitment locked to one instance family in one region becomes a liability the moment your architecture evolves. Flexibility is often easier to negotiate than a bigger discount, because it costs the provider less — and almost nobody asks for it.

Where do my prompts and responses go, and is training-exclusion contractual? The answers differ between providers and between service tiers within a single provider. The gap between "our documentation says we do not" and "our contract commits that we do not" is the entire question.

And one thing to do before you sign

Model what leaving would cost, layer by layer. Containerised compute is weeks. Object storage is trivial code and significant egress. Managed databases are moderately sticky. Identity is quarters and is usually the longest pole nobody estimates.

And the AI apparatus is sticky in a new way: the model call is portable in an afternoon, while the evaluation harness, prompt library and guardrail configuration built around one interface take quarters.

You need most of that analysis for disaster recovery anyway. That it doubles as negotiating leverage is a bonus — providers discount considerably more when the alternative is credible.

Full comparison: Amazon Web Services vs Google Cloud vs Azure

Frequently Asked Questions

Is first-party silicon worth the porting effort?
At sustained high volume on a stable workload, usually yes — the per-unit savings compound. Below that threshold the engineering time outweighs the saving. Budget it as a project, not a configuration change.

Which provider is best for AI?
Depends on training versus serving. Consuming a frontier model with enterprise guarantees favours Microsoft. Training on optimised frameworks often favours Google. Running open-weight models flexibly favours Amazon's neutrality.

Does data gravity outweigh technical merit?
Frequently. Compute should sit next to the data your AI features read, because egress and latency slow the experimentation loop, and that loop is what makes features good.

Is multi-cloud worth it for redundancy?
Usually not. It roughly doubles operational surface against an event rarer than your own deployment errors.

How long should a cloud decision take?
About a month: one page on workload shape, name the one binding constraint, two-week bake-off on the survivors, negotiate with a modelled exit cost.

What is the cheapest portability decision available?
Keeping the AI evaluation harness and routing layer provider-independent from day one. Nearly free at the start, very expensive to retrofit.

Sort:  

Lo de que Google rinde mejor en el happy path y bastante peor fuera de él es lo que más me hizo ruido, porque nadie te lo dice en el pitch. Pregunta concreta: cuando hablás de "fallback region", ¿hablamos de otra zona del mismo país o directamente cruzar el charco con la latencia que eso implica?

This hits on a very real operational bottleneck in modern cloud infrastructure. Accelerator supply and regional availability dictate architecture decisions much more than feature lists right now.