VIDEO PODCAST
Diskover's Will Hall: Optimizing Unstructured Data for Enterprise AI
James Maguire
October 9, 2026

As enterprises push AI into production, one of their biggest challenges is gaining control over the enormous volumes of unstructured data spread across storage systems, cloud environments, applications, and business units.

In this TechVoices interview, Diskover CEO Will Hall explains why machine-generated unstructured data can dwarf traditional knowledge-worker data, how Diskover uses metadata to create a catalog and abstraction layer across heterogeneous infrastructure, and why identifying the right data before moving it into AI pipelines can reduce infrastructure costs while improving model quality.

Hall also discusses Diskover’s expanded partnership with NetApp and argues that, as agentic AI and highly automated workflows become more common, data governance and the ability to manage massive unstructured datasets will become increasingly important.

‍

Core Takeaways
Unstructured Data at Massive Scale
Enterprise unstructured data increasingly includes machine-generated files from industries such as chip design, life sciences, energy, and media, with some environments containing tens or even hundreds of billions of files.
A Catalog Above the Storage Layer
Diskover aims to provide a catalog across cloud and on-premises storage, giving enterprises a global view of their data so they can understand its business value, govern it, and identify what should be retained or removed.
Curating Data for AI
Rather than moving enormous datasets indiscriminately, Diskover uses lightweight metadata to help enterprises identify the most relevant data for AI pipelines, potentially reducing storage, data-movement, GPU, and token costs while improving the quality of model inputs.
NetApp Integration and Enterprise AI
Diskover’s expanded NetApp partnership integrates its data intelligence capabilities more directly with NetApp environments, with plans for deeper storage and AI integrations designed to provide faster insight into large-scale unstructured datasets.
KEY QUOTES

Unstructured Data Is the “Wild, Wild West"

“Unstructured data is the Wild, Wild West. No exaggeration. If you look at the data, it’s really hard to search, it’s really hard to find. These systems are, for the most part, opaque, meaning that the business user has no way to access the crown jewels he has sitting on these massive repositories of storage, whether it’s on-prem or in the cloud.”
“So therefore, it's challenging from an inventory perspective for AI and challenging in terms of organizing the right data, getting to the right data for your AI pipelines. And this really compounds in these data-intensive workloads.”

Machine-Generated Data Changes the Scale of the Problem

“If you look at data in the verticals we’re talking about — media and entertainment, oil and gas, chip design, life science, pharma, healthcare, so on and so forth — you have machine-generated data that absolutely dwarfs traditional unstructured data as you would think about it in a knowledge-worker sense.”
“The result of that is hundreds of millions, if not billions, of files. We have one customer that has 103 billion files. All that is unstructured files and objects. And it’s just two orders of magnitude bigger than the way people traditionally think about unstructured data.”

Finding the Signal in the Data

“You need this abstraction layer above the storage, above the cloud bucket, that gives you the ability — one platform — to understand, be able to act on that data, and be able to govern all that data across all these different surfaces. What Diskover does is we allow you to really get that global view of the data, and then you can understand what data is valuable and what data is just costing you money.”
“More often than not nowadays, that’s: I want to do something with this from an AI perspective. So you have to take this giant haystack of data, curate it down to what’s most relevant, and then very efficiently send that over to the AI pipeline.”

Better Data Can Mean More Efficient AI

“If you can isolate what’s the most relevant data set and you only have to move a portion of that data set — or you can very efficiently move it from remote sites or different locations or different BUs into one centralized place more efficiently — you’re preconditioned to actually take advantage of AI. You’re actually bringing the AI to the data rather than having to move the data off-prem or to another site for the AI infrastructure.”
“If you can move it more efficiently, and then you are just moving these precise data sets, not only are your tokens more efficient, your GPU usage is more efficient, but you’re actually putting better quality data into the model — less opportunity to skew the results of the model.”
ABOUT THE AUTHOR
James Maguire
Executive Director
An award-winning journalist, James has held top editorial roles in several leading technology publications, covering enterprise tech trends in cloud computing, AI, data analytics, cybersecurity and more. He regularly communicates with industry analysts and experts and has interviewed hundreds of technology executives. James is the Executive Director of TechVoices.