Outscoring Claude Code at a Twelfth of the Cost: What to Check
- By Winston Thomas
- September 28, 2026

William Chen, a co-founder of Singapore-based Sapient Intelligence, says his startup’s AI does something coding agents don’t: it works out how to solve a problem, not just how to finish a task. He has numbers to back the claim.
On OpenAI’s MLE-bench, a test built from 75 Kaggle data science competitions, Praxist, the startup’s autonomous research system, won 49 gold medals for about USD3,054 in model costs. A baseline using Anthropic’s Claude Code, running Claude Opus 4.8, won 34 golds for USD38,370, according to the company’s launch release and paper.
A more revealing number appears in the company’s technical paper: 90,423. That’s how many experiments Praxist ran on the benchmark and then set aside after an internal integrity review of its records. On nine of the 75 tasks, the company replaced its best score with a verified one. On four, that cost it a medal.
Together, the two numbers frame the real question. Praxist’s pitch is not just about speed of research; its paper promises that every result arrives with an auditable record of how it was reached, a trail a regulator could follow.
For a CDO, that promise may matter more than any medal count. So it is fair to apply the same standard to Sapient Intelligence’s own claims: Which ones come with receipts, meaning proof that shows how the result was reached?
After an hourlong interview and my own research, I found that some claims already come with receipts while others still need testing.
The pitch: a lab notebook that never forgets
Sapient Intelligence was built on a bet. In 2024, Elon Musk’s xAI tried to hire Chen and Guan Wang after their small open-source model, OpenChat, took off. The offer was worth millions, Fortune reported. They said no.
“If we were to join xAI, it’s still going to be incremental work like post-training and engineering optimizations on top of that,” Chen tells me. Post-training means tuning a model after its main training run. Chen wanted to redesign the system itself.
Praxist is the result. “Agents give the models hands,” Chen says. “Praxist gives the models a mentor.”
A coding agent takes a defined task and completes it. Praxist targets problems where nobody knows the best approach yet. It runs several AI “research peers” in parallel, scores each attempt with an outside evaluator and passes what it learned to the next round of experiments.

What sets Praxist apart is how it treats failure. Rival systems such as Weco AI’s AIDE and Sakana AI’s AI Scientist use tree search: they keep the best-scoring attempts and prune the rest. Chen explains the risk with evolution. Crabs, lobsters and scorpions all evolved claws, a good-enough design that can crowd out better ones. In tree search, an idea that starts weak, perhaps because it was not trained long enough, gets cut and never returns.
Praxist keeps failed experiments as evidence. Think of a lab notebook that records why each idea died, so no one repeats it blindly. If that notebook drives the results, the next step is to measure how much of the benchmark win it produced. That is the first test.
Test one: the design
Researchers measure a component’s value with an ablation study: remove one part of a system and see what changes. The paper tests some individual parts this way, including one component of the rocket controller described below. I asked Chen how much the research graph itself contributed to the benchmark result.
“To be exact, I can’t say how this thing contributes to the score,” he says, “because there are a lot of compound elements in the process.”

The question matters because of the company’s history. Its first big release, the Hierarchical Reasoning Model, drew attention in 2025 for its brain-inspired design. Then the ARC Prize Foundation tested it. Its analysts found that a less-documented refinement loop, not the hierarchy Sapient Intelligence had promoted, drove much of the performance.
The Praxist paper itself names the problem, noting that the field often produces “gains that are hard to trace to a cause.” Measuring the research graph’s share of Praxist’s benchmark win would answer that directly. The cost claim raises a related question.
Test two: the price
The twelvefold cost gap is the number most likely to reach a board slide. But the benchmark changed two things at once: the research system and the AI model underneath it. Praxist ran on DeepSeek’s DeepSeek-v4-pro. Claude Code ran on Opus 4.8.
Chen calls DeepSeek “significantly a smaller and a weaker model” and credits Praxist’s design for closing the gap. He may be right. A test on the same model would settle it, and Chen says Praxist also runs on OpenAI’s GPT models and on a customer’s own model. Buyers can run that test themselves.

The paper is clear about its setup. It states that Sapient Intelligence ran both systems itself, once each, and did not compare the results with the public leaderboard. The launch release describes them as “controlled internal evaluations.”
Even a complete benchmark result would cover only a controlled test. The claims that matter to buyers come from outside it.
Test three: the real world
The most striking case study is a rocket. In a partner-provided simulation, Praxist built a controller that landed safely in 12,288 of 12,288 test runs. Weco’s optimizer, starting from the same 4.03% success rate, reached 17.12%. Chen says Praxist needed “12 hours on a laptop.” An experienced rocket engineer, the company says, estimated that matching the result through conventional development would take about eight engineers a month.
Sapient Intelligence rates this result at Technology Readiness Level 3: a proof of concept, in simulation. The paper reports its other case studies in the same open way. Praxist improved the main metric in two of four.
The result closest to a real shop floor comes from robot navigation. The field relies on simultaneous localization and mapping (SLAM), in which a robot builds a map while tracking its own position on it. Chen’s example is a maker of autonomous lawn-mowing machines that sells in Europe and the U.S. He says its 10-person in-house team spent three months reducing navigation error, and Praxist then cut the remaining error roughly in half in three days.
“As long as the mission has something that’s measurable, that’s something Praxist will be really good at.” — William Chen @ Sapient Intelligence
Sapient Intelligence’s materials describe three SLAM results. The launch release says a partner’s accumulated error fell from 9.37 centimeters to 5.01 centimeters in three days. The product page cites a 22.1% cut in computing load with better accuracy. The paper reports about 72% less visual processing time with accuracy unchanged. These may be separate experiments, so buyers should ask themselves which one matches their problem.
Hardware comes next: Chen says Praxist will connect to automated drug discovery labs, some in Singapore, with partners he expects to name in October. Real labs mean real consequences, and someone has to sign off on the results.
Test four: the sign-off
Humans set the goal at the start of each research cycle and review results at the end. In between, Praxist’s PI Panel, named after a lab’s principal investigator, decides what to test next. It has three AI roles: a Builder, a Skeptic and a Portfolio manager. A fourth, the External-Validity reviewer, checks whether results can be reproduced. It switches on in what the paper calls “high-stakes mode.”

Chen says users can pause and redirect a run at any time. But the volume makes live review difficult. Praxist produces about 300 research reports overnight, he says, roughly a year’s output from 10 computer science doctoral students. “It’s impractical for a human to step in,” he says. Instead, Sapient Intelligence places human researchers at the beginning and end of each cycle, with Praxist working to extend what they can do rather than replace them.
That design means one AI model checks another inside the loop. It is a common pattern, and a question buyers should weigh. Gartner forecasts that over 40% of agentic AI projects will be canceled by the end of 2027, and it names inadequate risk controls as one of three causes.
When no human sees every step, the record is the only reviewer that does. That puts the weight of the pitch on the lineage itself.
What the records leave out
Chen shared his screen to show one night of work. The lineage graph, a map of how each finding led to the next, showed yellow hypotheses, blue insights and green results. Lines between them marked which findings supported, challenged or updated others. “It’s not a black box at all,” he says.
On this evidence, it isn’t. For a CDO in pharma or banking, that graph is Praxist’s strongest feature. A research record is not yet a compliance log, though. Answers to tamper-proofing, how long records are kept and how each deployment is identified will shape how Praxist fares in an audit. Regulators have given buyers time to ask: the EU’s Digital Omnibus delayed the AI Act’s high-risk obligations to Dec. 2, 2027, for stand-alone systems, and to Aug. 2, 2028, for AI built into products already covered by EU safety law.
The records also stop at the edge of the research. They don’t show where your data goes. When Praxist runs on a hosted model, its research agents send prompts, which can include code and results, to that model’s provider. Chen says this is why Sapient Intelligence chose an open-weight model: customers can run the whole system on their own servers, with nothing leaving the building. Praxist’s README on GitHub adds that the software does not collect experiment data, only limited operational information that users can switch off. For a regulated firm, the on-premises setup is the natural starting point.
They don’t show what the agent can reach, either. For installs managed by an AI agent, the same README suggests starting OpenAI’s Codex with a setting called --yolo. The setting turns off two safety brakes: the agent no longer asks permission before running commands, and it is no longer confined to a sandbox, the isolated space that limits what software can touch. OpenAI’s own documentation recommends it only in an environment that is already hardened, so security teams will want to plan installs accordingly.
And they don’t show the bill. Praxist is source-available: anyone can read and modify the code, and use comes with conditions. Under its Fair Source license, any organization with annual revenue of USD1 million or more must negotiate a commercial license, which covers most enterprises. Chen says a commercial version is due later this year.
Four questions before you hire a robot researcher
None of this means Praxist falls short. It means the evidence buyers need most is still being built. Chen’s own rule of thumb tells you which projects qualify: “As long as the mission has something that’s measurable, that’s something Praxist will be really good at.” If your problem passes that test, ask for the receipts.
Ask for the same-model test. Run Praxist and your current coding agent on the same AI model. A good sign: Praxist keeps most of its lead, which would show the design is doing the work.
Ask for the ablation. Which part of the system produced the gain? A good sign: the vendor names the component and shows how much the score drops without it.
Ask for the rejects. Request high-stakes mode, which adds the External-Validity reviewer. Then ask how many attempts were set aside, and why. A good sign: a count, the reasons and the results that changed because of them.
Ask where it runs and what it can reach. A good sign: an open-weight model on your own servers, research agents confined to a sandbox, and installs done by a person. Keep it away from production data until you have all three.
One receipt already clears the bar. It isn’t 49 gold medals but the 90,423 experiments Sapient Intelligence’s records helped it find and set aside, even at the cost of four medals. Ask every autonomous research vendor for its version of that number.
Image credits: iStockphoto/demaerre; Figures and graphs from Sapient Intelligence
Winston Thomas
Winston Thomas is the editor-in-chief of CDOTrends. He likes to piece together the weird and wondering tech puzzle for readers and identify groundbreaking business models led by tech while waiting for the singularity.