Here's how research organisations can get ahead of AI workload complexity, via a global namespace
No substantial innovation can exist in a silo. While research teams globally are looking to get the best out of AI workloads underpinned by an array of storage systems, disconnected and insufficiently organised information risks project stoppages and governance gaps.
A proprietary approach needs to be replaced with a system that focuses on standards, regardless of the storage types in place. Many traditional AI storage approaches require data to be migrated into a new environment before it can be used. Avoiding an upfront data migration can shorten the time needed to put AI workloads into operation, while making use of existing cloud and on-premises environments.
To properly orchestrate the entire data stack for AI, a global namespace that stores unifies file and object metadata across storage systems is vital. Policy-driven orchestration can then enforce data placement and protection requirements in the background, allowing researchers to focus on delivering project outcomes.
Research obstacles
Bench scientists and data stewards alike make use of an array of tools, cloud and file storage systems, which is a common cause of siloed data.
Floyd Christofferson, VP of Product Marketing at Hammerspace, said that research data is frequently separated due to grant requirements or specific project constraints, which limits accessibility.
“You might have a research project that is funded by a grant, or part of a graduate project. This often tends to be isolated,” said Christofferson. “In the worst-case scenario, the data, by the rules of the grant, has to be in a specific storage.”
With inference required for AI models to analyse incoming and evolving data, research teams find themselves facing a difficult choice of whether to duplicate infrastructure, copy data, or migrate it, which can complicate governance and data sovereignty and increase operational costs.
“You can look at this in the context of a pilot project,” said Christofferson. “A certain amount of data is curated for AI, but when you get into production, it's not a ring-fenced environment. It now has to deal with live data.
“If you're dealing with new inputs, for example genomic research or materials research, before you can even get access to your inference engines, you've got to do all that copying. Time to token is not really the question. It's getting your data AI-ready, or time to the very first token, that is needed before you can take advantage of GPUs.”
Maintaining inference across pipelines
For AI-facing projects like drug discovery and medical imaging to deliver optimal insights and drive value, a data-centric approach that removes storage barriers between stages of the research workflow is required. End users need an end-to-end plan from considering where data begins, through where inference needs to happen, and to output.
“If you can alleviate friction in an operationalised AI pipeline, you can dramatically reduce the time and cost associated with taking a project from pilot to production,” said Christofferson.
Partnering with a platform like Hammerspace can help bring all storage types into one singular-view namespace. By assimilating metadata from existing storage while leaving the underlying data in place, Hammerspace makes that data accessible to authorised users and applications, including AI workloads, across locations.
A unified, multi-protocol global namespace enables access to distributed data without requiring a custom client or disruptive data migration. This architecture allows for background orchestration even on live data that is actively in use, which significantly reduces the friction involved in making data available for AI workloads.
Data protection and governance
Research projects can involve sensitive personal data and valuable intellectual property, making appropriate access controls and data protection essential. Governance needs to be maintained for the masses of unstructured imaging data residing across the IT estate.
Metadata-driven policies, known as objectives, can define how data should be placed and protected. Custom metadata can be used to identify datasets at a file-granular level that are subject to particular restrictions and inform the policies, exclusion criteria, and workflows applied to them.
Christofferson said: “Using custom metadata, you can assign restriction tags to individual files, or whole volumes of folders. These can identify what can't be used for AI, and be applied automatically to reduce reliance on somebody to remember to tag things.”
In addition, policies can maintain multiple instances of the same file across storage locations, supporting protection and availability requirements. These can include remote instances and write once, read many (WORM) protection.
“This provides a global, bird’s-eye view of all your data, and allows you to determine the behaviour of the files using multiple file system and custom metadata variables to automate policy actions,” said Christofferson.
A unified namespace
When it comes to getting the whole data stack ready for AI projects, it is orchestration, not migration, that should be the focus. Unified access with metadata fully aligned in the background is needed to maintain any truly innovative research project.
With a platform like Hammerspace, global research teams working across multiple lab spaces and infrastructures can begin accessing existing data in place through a unified namespace without first migrating it to a net-new repository. Data can be continuously orchestrated in a singular namespace that operates across cloud, on-prem and edge environments. Meanwhile, a standards-based parallel file system delivers high-performance data access to GPU workloads across heterogeneous storage infrastructure.
“Most AI storage platforms require you to copy that data into certain storage. But with Hammerspace, the same application space can be utilised with the same file system space that you are used to,” said Christofferson.