The Petabyte Pivot: When Unstructured Data Becomes Everyone’s Problem
- By Winston Thomas
- December 01, 2025

The infrastructure built for yesterday’s data is collapsing under the workload of tomorrow. And it’s far from being gradual.
An October 30 CDOTrends virtual roundtable, organized in partnership with HPE and held under the Chatham House Rules, invited tech leaders from financial services, gaming, energy, and stockbroking industries across Asia-Pacific to examine the issue of data infrastructure burden. A pattern emerged that should alarm any organization still treating unstructured data as a future problem: it’s here, expensive, and most enterprises are unprepared.
The structured data delusion
For decades, enterprise IT operated on a comfortable premise: data would arrive in neat rows and columns, queryable via SQL, manageable through established pipelines. A major bank’s AI/ML lead noted that while structured data has mature collection, storage, and ETL options with many machine learning applications built on top, unstructured data lacks similarly developed systems for managing collection and efficient storage.
Generative AI has now created greater demand for unstructured data in more ways than one and revealed how much valuable intelligence has been locked in legal contracts, compliance cases, trading communications, and decades of documentation that legacy systems were never designed to parse.
“I’ve been dealing with storage technologies for the last 26 years, and I can tell you that I’ve never seen anything even remotely approaching the data capacity, performance and variety explosion that we’re seeing today,” says Alex Veprinsky, chief architect for storage at HPE. The infrastructure gap isn’t measured in percentage points but in orders of magnitude. “Three-quarters of the internet ports in the world are 25 gigabit per second or less, when today’s cutting-edge technology is 400-800 gigabit per second.”
The hidden costs of standing still
Highly-regulated industries face a particularly acute version of this crisis.
One participant managing IT infrastructure for an oil and gas conglomerate described servers running decades-old applications on unsupported operating systems, kept alive solely because they contain irreplaceable vendor history and payment records. Migration is impossible without application rewrites. Deletion is unthinkable for compliance reasons.
Meanwhile, banks processing unstructured data for AI applications discovered that API costs fluctuate unpredictably, performance drifts over time even with the same models, and expanding from simple text summarization to reasoning tasks requires exponentially more expensive models. Traditional linear cost-value relationships don’t apply when your technology stack evolves weekly.
Why traditional storage economics are broken
The storage industry spent fifty years optimizing for one metric: access latency. Get data in, get it back fast. But modern AI-driven workloads demand something fundamentally different.
“It appears that for all those tools that sit on top of a storage system, the latency of getting and putting the data is less critical,” Veprinsky explains. “What’s important for them is how soon they can derive value that is relevant for your business from this data. So storage stops being where information lives. It starts being a place where information actually gets generated.”
This shifts storage from a passive repository to an active processor. The best I/O operation is the one you don’t need to make: transformations performed during storage or access operations reduce the time required to create metadata and indexes.
That philosophical shift also introduces architectural implications. “The best I/O operation is the one you don’t need to make,” Veprinsky notes. “Essentially, you have a set of transformations that are performed on data as part of it being stored or accessed, and that shortens the time to you being able to create those metadata indexes.”
It’s only going to get more complex. Data lakes are already evolving into Data Lake Houses, some using Apache Iceberg format to enable SQL-like queries over unstructured data. Agentic AI pipelines will enable specialized AI models to communicate via protocols such as MCP (Model Context Protocol), creating meshes rather than monolithic solutions.
“Making them interoperate and set up correctly is no less of a challenge than everything else,” Veprinsky warns about the complexity of orchestration.
The APAC reality check
Christophe Maisonnave, APAC sales director for hybrid cloud services at HPE, observes a troubling maturity gap: “When we look at some of the analyst reports that we have seen recently… maybe less than 30% of the companies actually in the market have end-to-end, fully efficient FinOps readiness. Everybody wants to address it, but the hard reality is that everybody is having just one portion of what they would like to have.”
The gap between aspiration and execution is widening. Data localization requirements mean GPU infrastructure might be centralized in one country while data must remain distributed across regulatory boundaries. “More than 60% of the companies in Southeast Asia are basically already seeing latency as the biggest problem that they are going to have to benefit from structured data for AI,” Maisonnave adds.
Cloud repatriation is accelerating not just from cost, but from physics: you can’t train models on data you can’t access quickly. “This whole agentic question will also have to create a form of network in a way, which is again, going to be an evolution compared to what it was a couple of years ago,” Maisonnave notes. “The edge and the data residing at the edge are becoming more and more critical and complex, and also need to be managed.”
It’s not just a technological issue. Our practices around our technology also need to evolve. Imagine: in 2013, building a 10-petaflop data center required infrastructure comparable to that of a nuclear plant; today, it requires a couple of racks. Technology cycles that once took years now complete in weeks.
Traditional capital expenditure planning, which includes locking in a three-year roadmap, execute, refresh, guarantees obsolescence before deployment. “If you go with the plan that you have today, and go like a train, actually, it’s a very big financial risk that you are going to take,” Maisonnave warns, “because you may not be even at the right point of competition into the market, because you are already into a train which is already too old, even though it’s not even starting yet.”
Organizations need procurement flexibility to adjust when new technologies emerge. “This agility, in addition to the monitoring of the spend, is going to be also a necessity for every customer in the future,” says Maisonnave.
Keeping it simple
The advice from practitioners who’ve navigated this transition? Start with data lifecycle fundamentals and don’t give life to something already dead by migrating garbage. Segment ruthlessly: what data will you monetize versus what’s just ballast?
Simplicity rules the day, Maisonnave reflects. “We all need to go back to the fundamentals. Keep things very simple and go to the fundamentals — sometimes we like to talk about a lot of complexity, but keeping things simple is probably the best chance for people to be successful.”
Storage should also stop being where information lives and start being where information gets generated, Veprinsky argues. But that transformation requires orchestration expertise that most organizations lack internally.
The petabyte pivot is a reality across all APAC organizations. The question that matters is whether you execute it strategically or let it manage you financially in a world where innovation cycles faster than budget approval processes.
Image credit: iStockphoto/tadamichi
Winston Thomas
Winston Thomas is the editor-in-chief of CDOTrends. He likes to piece together the weird and wondering tech puzzle for readers and identify groundbreaking business models led by tech while waiting for the singularity.