New AI Models Push the Envelope Again
- By Paul Mah
- July 17, 2025
The AI race shows no signs of slowing, with regular releases of newer, more powerful models. Just last week, AI company xAI released Grok 4, while Chinese firm Moonshot AI launched its Kimi K2 family.
Grok 4
Grok 4 was launched by xAI last Wednesday with two models: Grok 4 and Grok 4 Heavy. The latter is touted as a multi-agent model with enhanced performance. In a livestream introducing the product, Elon Musk said Grok 4 Heavy spawns multiple agents to work on a problem simultaneously. “With respect to academic questions, Grok 4 is better than PhD level in every subject, no exceptions,” he claimed.
The company says Grok 4 delivers frontier-level performance on several benchmarks, including Humanity’s Last Exam (HLE). HLE is an extremely difficult test containing crowdsourced questions across a wide range of subjects – mathematics, physics, biology, and more – specifically designed to assess LLM capabilities.
As reported by TechCrunch, Grok 4 scored 25.4% on Humanity’s Last Exam without “tools,” outperforming Google’s Gemini 2.5 Pro, which scored 21.6%, and OpenAI’s o3 (high), which scored 21%. Grok 4 Heavy with tools was able to achieve a score of 44.4%, outperforming Gemini 2.5 Pro with tools at 26.9%.
xAI also introduced a USD 300-per-month subscription called SuperGrok Heavy. Subscribers will get early access to Grok 4 Heavy and upcoming features. This plan is similar to the ultra-premium tiers offered by other AI firms.
Kimi K2
Launched over the weekend, Moonshot’s Kimi K2 is said to excel in frontier knowledge, mathematics, coding, and general agentic tasks. It’s a cutting-edge model built on a mixture-of-experts (MoE) architecture, with a trillion parameters and 32 billion activated parameters.
As reported by CNBC, Moonshot claims that Kimi K2 surpassed Claude Opus 4 on two benchmarks and outperformed OpenAI’s coding-focused GPT-4.1 model across several industry metrics.
Two versions of the Kimi K2 large language model have been open-sourced: the Kimi K2 foundational version and Kimi-K2-Instruct. The former is optimized for researchers and developers who want full control for fine-tuning and customization, while the latter is ideal for chatbots, autonomous agents, and complex workflow orchestration without requiring extensive retraining.
Kimi K2 is accessible via API and priced at 4 yuan (56 US cents) per million input tokens and 16 yuan per million output tokens, reports the SCMP. Support for Model Context Protocol (MCP), an open standard created by Anthropic that enables AI systems to use external tools and services, will be introduced soon.
Kimi K2 is now freely available via its web and mobile applications.
Image credit: xAI
Paul Mah
Paul Mah is the editor of DSAITrends, where he report on the latest developments in data science and AI. A former system administrator, programmer, and IT lecturer, he enjoys writing both code and prose.