AI needs more than a cloud-first infrastructure strategy

Terry Storrar
Terry Storrar
Managing Director at Leaseweb UK

As organisations move AI from experimentation into production, Terry Storrar, Managing Director at Leaseweb UK, explains why infrastructure choice is becoming a strategic consideration for DevOps teams.

There is no denying that the explosion of AI workloads is profoundly shaping enterprise IT choices. The pressure is greater and the pace faster for development teams to test, refine and deploy multiple AI applications and large language models that support complex, data-heavy workloads. So, while software delivery was once the core focus for DevOps teams, the AI landscape is now forcing a review of infrastructure choices as a strategic priority for any organisation building AI-driven services.

A recent survey highlighted that almost half of organisations believed their DevOps teams lacked the bandwidth needed to support the very rapid growth of AI workloads. Cloud-first architectures based on more predictable workloads and flexible compute are struggling to provide DevOps teams with what is needed to fulfil the demands of modern AI workloads. Rising costs, lack of real-time visibility, challenges with scaling and resource allocation, compliance and GPU availability might all be infrastructure issues, but AI has noticeably exposed their direct impact on development teams. Quite simply, without the right infrastructure in place, development teams are not able to deliver what is expected of them.

Why Infrastructure-as-a-Service is now strategic

Many organisations assumed that their existing cloud setups would support their AI ambitions. So what are the key concerns that DevOps teams must address?

For DevOps teams already balancing uptime, deployments, security and developer productivity, the demands of AI create an even tougher operational environment, throwing new challenges into the mix. DevOps teams now need to ensure infrastructure access is speedy enough to support experimentation without compromising operational performance. GPU acceleration is a necessity, and a rethink of storage architectures is often required to support AI development. These factors are becoming as fundamental to designing infrastructure as the AI application requirements themselves.

In this environment, it is no surprise that Infrastructure-as-a-Service (IaaS) is increasingly viewed as a strategic consideration for DevOps, particularly where organisations need infrastructure options that can respond quickly to changing requirements. These choices make a difference to how quickly new workloads can be tested and deployed, and then scaled, governed and managed in longer-term operations. IaaS can also provide organisations with the agility to adapt to AI workload demands as and when needed, creating an environment in which DevOps teams can optimise and accelerate AI projects.

Balancing infrastructure choices

One of the key trade-offs that DevOps teams need to consider is how to balance using hyperscale cloud providers with other infrastructure options, such as private cloud, to fulfil all phases of AI deployments. At the simplest level, this means deciding where AI workloads are best allocated, weighing the scalability of public cloud against the cost predictability and control offered by dedicated infrastructure.

There is an increasing realisation that public cloud may not be the most suitable choice for every workload or every phase of development. In fact, recent research has found that organisations are shifting production AI workloads towards private cloud, with 56% running or planning to run production AI inferencing on private cloud, compared to 41% on public cloud. Interestingly, the research indicates that DevOps teams are making conscious infrastructure choices for different development stages. Public cloud is currently favoured for early AI testing and model training, but organisations are increasingly choosing private environments for AI inference at scale.

Hyperscale platforms offer significant benefits; they enable development teams to easily scale services across the globe, launch applications speedily and access advanced AI tooling. However, there are downsides too. Controlling costs is increasingly a concern when delivering AI workloads at scale, especially given GPU power requirements. In addition, large language model workloads can be an excessive drain on cloud budgets. DevOps teams are typically working within complex infrastructure environments, so it can be an additional challenge, beyond their core focus on delivery, to have visibility of where workloads are allocated and how costs are accumulating.

One of the newest and most significant demands that AI is placing on IT development is data protection and compliance. While most organisations are now facing tough questions about data residency, governance and infrastructure access, at the same time DevOps teams are torn between adhering to regulations and avoiding over-reliance on proprietary ecosystems that restrict the portability of the new applications they are evaluating.

Consider specialist providers

It’s in this situation that regional and specialist IaaS providers can offer an alternative to hyperscalers, often differentiating through flexibility, transparency and infrastructure tailored to particular workloads. For example, a specialist provider may be able to deliver specific GPU infrastructure and network and storage architectures designed for compute-heavy workloads. This flexibility can be useful for DevOps teams, as AI projects tend to move through development more quickly than conventional applications, meaning infrastructure needs to accommodate this pace.

Local or regional providers may also offer a more customised operating environment, for example through support models that provide direct access to engineers familiar with AI development environments. For some companies, this tailored option may be more effective than navigating standard support processes, particularly if an AI workload needs timely action.

With greater emphasis on data sovereignty and regulation across many industry sectors, Europe’s businesses are also paying closer attention to transparency around data residency and management, alongside the ability of providers to support their legal and compliance requirements.

Is hybrid ideal?

The appeal of hybrid is that the model can be adapted to an individual organisation’s needs, using multiple providers and retaining the flexibility to change as requirements evolve. Many DevOps teams now combine hyperscale cloud with regional IaaS providers to suit their workload criteria. For example, while hyperscale might support global services, specialist providers might be a more suitable choice for GPU-intensive AI training and managing highly complex datasets.

Hybrid infrastructures are likely to become even more common across development environments, as they are across entire organisations. DevOps teams are likely to be better placed within environments that support rapid experimentation and scalable deployment, and are built for operational resilience.

For DevOps teams, the challenge is not only to deploy software, but to be able to mould platforms that will reliably support AI at scale. Whether the choice is reliance on hyperscale cloud, using private infrastructure or adopting a hybrid approach, it is the DevOps teams that manage to fine-tune the balance for their needs that will be best placed to achieve the desired AI outcomes.

Related Articles

More Opinions

It takes just one minute to register for the leading twice weekly B2B newsletter for the data centre industry, and it's free.