AI's Quiet Rebellion: The Rise of Practical Transformers
- By Winston Thomas
- February 23, 2025

When DeepSeek matched the performance of much larger AI models with a fraction of the resources earlier this year, it sent shockwaves through the AI community. The wake-up call was clear: bigger isn’t always better.
“People started to realize that just training and continuing to expand the existing architectures is not the end-all solution,” says Max Vermeir, senior director of AI strategy at ABBYY.
This revelation comes at a crucial time. The AI industry’s hunger for computational resources has reached gluttonous proportions, with some models demanding “hundreds of gigabytes of memory just to run a single instance,” according to Vermeir. It’s a reality check that’s forcing the industry to confront an uncomfortable truth: the current path of scaling transformer-based architectures for multimodal models isn’t just challenging; it’s potentially unsustainable.
The crossroads of ambition: AGI dreams vs. business realities
The industry is witnessing a fascinating split that could reshape its future. “There’s this kind of dual dynamic happening,” Vermeir explains. “One is the continued race for AGI, and the other is how do we take our learnings and actually make that work for people in the business environment?” This dichotomy isn’t just theoretical — it’s driving real changes in how companies approach AI implementation.
ABBYY is betting on a different horse in this race: Small language models (SLMs). “Small language models are much more efficient,” Vermeir argues. “Whenever we’re talking about document AI and process AI... it’s all about speed, accuracy, and consistency — which is the exact opposite of what large language models are.” While large language models excel at general tasks with their probabilistic approach, SLMs deliver laser-focused, consistent results in specific business contexts.
The emergence of agent-based AI has added another layer to this complexity. While there’s considerable buzz around AI agents, Vermeir warns of “agent washing” — the overselling of basic automation as sophisticated AI agents. However, he remains “a little bit more optimistic about this one in the sense that it’s going to be less hype and actually have practical applications pretty quickly.”
The technical battleground: Fusion, attention, alignment
The key to successful agent implementation, Vermeir argues, lies in the tools they’re given. “Agents need tools; the agent in itself is just the kind of thinking brain, but it needs ‘hands’ to actually go do something.” This is where document processing becomes crucial — so crucial that even NVIDIA has recently entered the intelligent document processing market. “That’s where the data is that we can still use to train models,” Vermeir points out.
It is also an area where AI engineers run into roadblocks, particularly when it comes to handling multiple types of data from multimodal input, e.g., voice, text, visual, etc.

Lately, the industry is grappling with choices between early and late fusion strategies in multimodal transformers. Early fusion offers deeper interaction but demands more computational resources and struggles with missing modalities. Late fusion processes modalities independently until decision time, offering more flexibility but potentially losing inter-modal connections. Google’s multimodal bottleneck transformer (MBT) attempts to find the sweet spot using cross-attention to mix both strategies.
Attention mechanisms allow a neural network to focus on the most relevant parts of the input data when making predictions. Imagine reading a long document: your attention shifts to the most important sentences. Similarly, in AI, attention helps models prioritize crucial information. This is particularly vital in multimodal AI, where models must process and integrate diverse data types like text, images, and audio.
Here, self-attention mechanisms are beginning to promise, particularly with DeepMind’s Perceiver IO architecture that “extends the concept of self-attention by processing arbitrary input modalities through this unified attention mechanism,” Vermeir explains. However, cross-modal alignment remains an open challenge, with researchers still exploring hierarchical attention and multi-agent approaches.
Pragmatism triumphs: The art of the right tool
When it comes to evaluating these complex systems, Vermeir advocates for a pragmatic three-step approach: first, “realize there’s no one size fits all”; second, “look at the use case you're trying to work with”; and finally, choose the right technology for the specific task. Sometimes, he notes with a hint of irony, the solution might not even require advanced AI, pointing to an experience talking to some data scientists on the best model architecture. “The best solution to the problem was a regular expression.”
The risks of overcomplicating solutions become particularly apparent in business applications. Vermeir illustrates this with accounts payable: "If you’re trying to solve an accounts payable problem and you have a model that does overfitting... Like should I actually pay this invoice sooner because I’m getting a discount? You hope that it gets it right. Otherwise, you’re going to be paying a vendor before you actually have to.” He also notes it’s one reason why many enterprises are returning to companies like ABBYY “who already have the right processes and results.”
Looking ahead, Vermeir sees adaptive transformers and multi-agent approaches as the future of transformer-based architectures. “A lot of agents will work in parallel to process multimodal inputs efficiently, each having their specializations,” he predicts. Adaptive transformers automatically adjust to task complexity. He also sees knowledge graphs as providing crucial process understanding for businesses.
However, the industry will continue to see a "duality between large cloud-based compute clusters striving for AGI" and more practical, efficient solutions. But the key to success isn’t just about choosing sides — it’s about finding the right balance between innovation and practical application.
However, Vermeir believes the industry will continue to dance between “cloud-based AGI dreams and practical, efficient solutions.” The winners will not be those who chose one side of the other; it will be those who master the art of matching the right tool to the right task, armed with a deep understanding of processes. And in the realm of process intelligence, ABBYY believes it holds a decisive advantage.
Image credit: iStockphoto/Alona Horkova
Winston Thomas
Winston Thomas is the editor-in-chief of CDOTrends. He likes to piece together the weird and wondering tech puzzle for readers and identify groundbreaking business models led by tech while waiting for the singularity.