Can AI Tame the Data Frankenstein It Created?
- By Winston Thomas
- July 30, 2025

The AI revolution promised to make computing smarter and more efficient. But it also created algorithms with endless appetites for data.
AI systems now consume information at unprecedented scales while simultaneously generating artificial data to fuel their ever-growing needs. So, can AI solve the massive data storage crisis it created without destroying the authenticity that makes intelligence valuable?
At a recent Pure Storage panel discussion earlier in July 2025, three industry experts examined this challenge with the technical precision that keeps enterprise executives awake at night. Their conclusion? The era of simply adding more computing power and data is coming to an end, and what happens next will determine whether AI becomes humanity’s greatest tool or its most costly mistake.
The true cost of data hunger
When discussing data, we often overlook the impact on the infrastructure that supports it. The effect can be staggering: companies are seeing 5-10 times more storage requirements as they convert unstructured data into vector embeddings for retrieval-augmented generation (RAG) systems.
Vijay Morampudi, senior vice president of AI Center of Excellence at Marsh McLennan GCC, described the challenge from real-world experience: “80% of our data is unstructured data. So now with the advancement of generative AI, we are able to bring this unstructured data and convert it into a format that we can feed into these models to enhance business decision-making.” But this transformation comes at a steep price. When the audience was asked about AI adoption challenges, cost was the most frequently cited issue at 38%, followed by performance issues.
The infrastructure demands are reshaping entire data strategies. Matthew Oostveen, Pure Storage’s chief technology officer for Asia Pacific & Japan, revealed the harsh mathematics: “Research shows about 25% of the power is dedicated towards storage. We have this new generation of Nvidia systems that require a great amount of kilowatts, and we’re now reaching thresholds in customer data centers, where they’re unable to deploy as many GPUs as they want because they’re meeting these thresholds.”
To watch the full panel discussion, click here.
The shift toward efficiency
All three panelists agreed that the “bigger is always better” approach is dying. Anirban Nandi, vice president of AI and Data at Rakuten India, delivered a clear message about the limits of current scaling methods: “I think generative AI’s progress has largely been fueled by scale, more data, more compute, and more parameters. But I think we are hitting a point of diminishing returns right now, where performance gains come at disproportionate energy and environmental cost.”
The industry is moving toward what Morampudi calls “domain-specific models”. These “lean” specialized AI systems solve specific problems without the massive computational overhead of their giant predecessors. “Large vendors who are building the models, instead of adding more GPUs, are trying to optimize some of the key operations at the kernel level so that you can do that in a much faster way,” he explained, pointing to DeepSeek’s breakthrough in low-level optimization.
Rakuten’s approach shows this shift in action. Nandi revealed their strategic decision: “Two and a half years back, we made a strategic decision to invest in our own large language models. So, Rakuten AI has two large language models.” The larger model employs a mixture-of-experts (MoE) design, which activates only the relevant parts of the system, significantly reducing computational costs.
Rethinking data for vectors and scarcity
The discussion dove deep into vector embeddings, mathematical representations that allow AI systems to understand and retrieve information. These embeddings are creating a storage crisis. Morampudi shared a real example from a private equity project: “When we created the embeddings, especially for some complex use cases, we used high-dimensional vectors, which required up to 10x more storage.”
For other use cases, the solution involves sophisticated compression techniques and optimisation of indexing algorithms to balance precision and recall during retrieval. “We applied quantized techniques to reduce the size of the embeddings. With this, we are able to reduce the size by 90%. Of course, there is a challenge in terms of the quality,” Morampudi explained, acknowledging the trade-offs that come with aggressive compression.
Oostveen provided the infrastructure perspective: “Storage isn’t about space anymore. It’s around fast retrieval and how it works with semantic search. What’s the interaction with contextual indexing, and how do you do this at a massive scale?”
On the opposite end of the data explosion lies data scarcity, which can seriously hobble AI model development. Enter synthetic data: AI-generated information designed to train other AI systems. But it also comes with serious risks. Oostveen approached the topic with careful optimism: “There is absolutely an upside to the capability for synthetic data to be able to alleviate some of the scarcity, perhaps reduce some of the bias, when it’s very well controlled.”
The medical industry represents synthetic data’s best use case, where privacy concerns and data scarcity create perfect conditions for artificial generation. But Oostveen warned of potential dangers: “I’m very cautious of moving to a future where we’re ingesting data that risks eroding authenticity, whether we’re going to start amplifying, this model echo chamber phenomenon that we’re seeing.”
Nandi emphasized the importance of human oversight: “You need to ensure that, with the AI human synergy, you annotate it properly. Make sure it’s bias-free. Make sure it has a little bit of variation as well.” The risk? Synthetic data that’s too clean and similar creates models that fail when faced with real-world complexity.
Building better infrastructure
The infrastructure implications are massive, and fortunately, vendors are taking notice. Pure Storage, for example, is developing new approaches that separate data from metadata, creating what Oostveen describes as potentially “the fastest storage system on the planet for some of these larger installations.” This approach recognizes that modern AI workloads require two distinct storage patterns: massive sequential reads for training data and millions of small, random accesses for metadata and embeddings.
Whatever approach you take or pivot to, the panel’s agreement was explicit: the future belongs to intelligent, purpose-built systems rather than simply adding more computing power. As Nandi summarized: “We are no longer just talking about scaling models. We are scaling it with responsibility and trust in mind.”
Oostveen's closing advice captures what’s at stake: “The infrastructure decisions that you make today will lead to your ability to execute in the future. Choose wisely and think about what it is your organization really needs to accomplish.”
The alternative is just too expensive and can threaten the very foundations that make AI meaningful to humans.
Image credit: iStockphoto/KIT8
Winston Thomas
Winston Thomas is the editor-in-chief of CDOTrends. He likes to piece together the weird and wondering tech puzzle for readers and identify groundbreaking business models led by tech while waiting for the singularity.