As organisations look to move AI projects beyond the pilot stage, Colin Fernandes, Enablement Director at Acceldata, explains why bringing compute closer to where data already resides could offer a more flexible, controlled and cost-effective way forward.
AI seems to be everywhere, with companies looking for ways to improve their performance or reduce their costs. According to research from the UK Office for National Statistics, 25% of companies use AI today, rising to 44% in companies with more than 250 employees. The challenge is that many of these companies struggle to get beyond their initial pilot projects. One of the reasons for this is data.
AI relies on data to work effectively. However, while many companies think their data is ready for AI, the reality is that enterprises don’t always have the right approach in place to support what their AI projects actually need. According to the EDM Association Global Data Benchmark Report, only 31% of companies report that they have deployed an advanced data strategy capability.
Building that capability involves looking at how you will operate and use your data for AI. The challenge is that AI projects assume you have the right data infrastructure in place, so your data can be discovered, trusted and accessed securely by authorised users, wherever that data is stored. This does not require data to be centralised. A shared catalogue and common governance policies can provide authorised access to data that remains distributed across data centres, clouds and other platforms. Data from GLG found that 80% of large enterprise companies currently run hybrid environments with multiple data warehouses and data lakehouses. Three-quarters have four or more of these environments deployed.
Deploying data effectively
The promise around data warehouses is that organisations can consolidate structured data from multiple sources, retain historical information and make that data available for reporting and analysis. Ideally, this means modelling and governing the data so it can be used efficiently and new value created from it. While companies have considered consolidating around the public cloud and a single data platform, their data teams also have requirements around data sovereignty and vendor lock-in. Many enterprises are therefore prioritising open standards, interoperable architectures and, where appropriate, open source technologies to reduce dependency and preserve flexibility. The cost of consolidation can also be substantial and the process potentially time-consuming.
Alongside this, you will have to prepare that data for AI work. The traditional approach is to set up data pipelines that combine the data required with other sources, enrich it and then pass it through for analysis. Each of these pipelines will include multiple stages, and any change can lead to the process breaking down and needing to be rebuilt. This can be expensive in its own right. Copying data between platforms can also create additional storage requirements and more copies that have to be governed, secured, monitored and eventually deleted. So, staying with things as they are may not be a sustainable option either.
The first step is to look at where you host and store your data compared to where you process it. Previously, that data warehouse processing and storage would have been directly linked, but that is no longer always the case. A data lakehouse can separate compute from data storage, allowing organisations to take advantage of the structure of a data warehouse while using a more flexible deployment approach.
However, how you implement a data lakehouse involves making some choices too. Moving to the cloud is one option, where you can run a data lakehouse service from your choice of hyperscale cloud provider, or alternatively use a data lakehouse vendor on top of the public cloud. One approach is to run with a single cloud provider and lakehouse option. This can simplify implementation, but the trade-off is that organisations may become more closely tied to that provider or deployment approach.
For those with more complex requirements, a federated data approach stores data across different cloud providers. In these deployments, the data plane operates across different providers, while the compute runtime and control plane can be implemented separately. The choice then is to consolidate the compute and control plane in one cloud environment, or federate the compute side as well and run the control plane separately. In practice, this can mean running federated compute where data lives, whether that is across different clouds or on-premises deployments.
Running compute closer to where the data lives can reduce the amount of raw data that has to move. Workloads can also be placed on available capacity in another data centre or, for temporary peaks, burst to the public cloud. This can improve the use of existing infrastructure and reduce the need to provision every environment for maximum demand. However, federating compute also adds operational complexity because teams may have to manage multiple runtimes, networks, identity systems, security policies and monitoring environments. A common control plane can therefore be important for maintaining consistent governance and visibility across the estate.
To make the right decision for your organisation, you have to evaluate how much of your data can remain where it is, what has to stay in its original location and how much you can consolidate. For some organisations, the cost of moving data or the data handling regulations they are covered by will prevent migration from being practical. For others, the demand for AI may provide the investment needed to make changes around data management and consolidate on a single platform in the cloud. Many organisations are likely to follow a middle path, using data where it is through federation and adopting hybrid approaches where appropriate.
Wherever your data currently resides, making it ready for AI may involve some changes. But it also involves looking at how to get your compute to run more efficiently and support your business goals. For organisations with distributed or regulated data, bringing compute and AI closer to the data may offer a more practical long-term approach than moving large volumes of data between environments. Where data does need to move, that movement should be authorised, governed and monitored, with visibility over what is transferred, why it is moved and where the resulting copy is stored. Whatever approach you take, the priority should be maintaining control over where compute runs and where data resides, so the business can meet its requirements around sovereignty, cost and control.

