The AI Hangover: Hallucinations, Waste, and the Search for a Cure
- By Lachlan Colquhoun
- September 16, 2024

Data management is the key to delivering optimum value from the new wave of Artificial Intelligence applications, and this requires a new approach to standing up infrastructure.
At Pure Storage, Matthew Oostveen says that effectively configured storage infrastructure can deliver more flexibility, drive better data preparation and avoid scalability bottlenecks while minimizing latency and maximizing throughput.
Oostveen sees a “phase of opportunity” for AI, but optimizing that opportunity requires fresh conversations on how data is stored and delivered for processing.
“We’ve been hyper-fixated on the GPU, and the compute side of the equation, and I think we forget that the GPUs need to be fed with data, and that requires specialized systems,” says Oostveen, who is Pure Storage’s chief technology officer and vice president for the Asia Pacific and Japan.
“You’re also going to need specialized systems to be able to keep up with the abundance of data that’s produced, and for that, we need a different architecture.”
Oostveen says that organizations implementing AI find their platforms “are not tuned to the task of an AI system” when they move into the inference phase.
“These existing platforms are poorly utilized and waste excessive amounts of data center space and power, which is a big issue for many organizations in the region,” he adds.
With trust — in both data and output — as the priority, many AI implementations that used large off-the-shelf language models fell short.
Organizations need to augment this process by embracing the retrieval-augmented generation (RAG) approach. This approach taps into proprietary or domain-specific information and specialized databases and complements the generational capabilities of the large language models.
“So, if a user has an input, they’ll ask a question that will go into something that is very domain-specific,” says Oostveen.
“It could be SEC filings over a number of years, or internal databases from stored customer interactions, and they are then fed into the LLM, which is able to provide a higher level of precision in the responses.”
Matthew Oostveen will be a key panelist at the October 11 Pure Leadership Series Webinar: Conquering AI’s Twin Hurdles in Cost Optimization & Scalability. Looking to tackle the two biggest challenges of AI? To register and find out more, click here.
Fewer ‘hallucinations’
Oostveen observes that organizations implementing the RAG model were getting results that outperformed those from humans at lower cost and scale.
There was also a cutting down on AI hallucinations where systems deliver false or misleading information.
Working with generic AI and large language models delivered results where around one-third of the responses were hallucinations, but this could be significantly reduced with the RAG approach. However, this requires a different storage model.
“RAG architecture leverages vector databases, which is a different way of aligning information so it can be fed into an LLM and pulled through the data pipeline more efficiently,” says Oostveen.
In terms of cloud strategies, while some organizations were doing this in the public cloud, others — particularly in the financial services sector — were doing it in private clouds on-premises or in a hybrid public-private combination.
“Utilization rates for NVIDIA GPUs are under 20%. This is staggeringly wasteful as they are scarce and costly.”
Pure Storage’s next-generation Fusion solution had a big impact on the management, provisioning and governance of data storage, mainly when the architecture is on-premises.
Integrated into the operating system and the hyperscaler eco-system, Fusion automates and relocates the provisioning of data pools to ensure performance profiles and protection policies remain intact without disrupting the end-user experience.
“When it comes to RAG, there are some really important reasons why you might want to do it on-premises,” says Oostveen.
“If you have already invested in an NVIDIA array, then leveraging its capabilities with fat storage will help with performance and latency issues, with the added benefit of more control over compliance, regulation and sovereignty.”
“It is also likely you will be using internal and confidential datasets that you don’t necessarily want to feed into an LLM model on a large public site,” he adds.
Sustainability and utilization
Improved sustainability and utilization worked together to drive better performance in AI projects.
A key to higher performance was the increased utilization of GPUs, which reduced power usage to limit energy bills.
“Utilization rates for NVIDIA GPUs is under 20%. This is staggeringly wasteful as they are scarce and costly to acquire and to run; remember, you’re paying for that power bill regardless of the utilization rate,” says Oostveen.
“With the right storage architecture, you don’t continue to invest in more systems because we can get nearer to 100% utilization.”
Image credit: iStockphoto/Urupong
Lachlan Colquhoun
Lachlan Colquhoun is the Australia and New Zealand correspondent for CDOTrends and the NextGenConnectivity editor. He remains fascinated with how businesses reinvent themselves through digital technology to solve existing issues and change their business models.