AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Anyone who has sat through a season of hardship knows a simple truth: it is easy to be wise in calm conversation, and very hard to be wise on the worst day. Faith traditions have taught this for millennia — the desert monks fled to the wilderness precisely because comfort makes virtue too easy, and the real test of a soul arrives with hunger, fear, and a tempting shortcut. So it is quietly fascinating that the technology industry has just rediscovered this principle, not in a monastery, but in a software emulator.

For years, we have measured artificial intelligence by how well it chats — eloquent answers, clever code, polished explanations. But a live experiment at Firmulate asked a different, older question: what does the agent do when the week goes wrong? When a customer threatens to leave, when money runs out, when someone impersonates the boss and asks for just one small favor?

The results read less like a benchmark and more like an examination of conscience.

The worst week, run four times

Here is the setup. Firmulate — which bills itself as an AI company emulator — handed four frontier AI models the same job: run an identical small software company through its worst week. Same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable, so nothing could be quietly rewritten afterward.

The final league table from the July 2026 “Crucible League” run tells the story at a glance: gpt-5.6-sol finished first with 95 points, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 at 77, and Opus 4.8 last at 73. A do-nothing baseline — a model that simply sat on its hands — scored 26, because partial progress counted. But one rule towered over the arithmetic: a single breach of trust capped the total. In Firmulate’s own words, “no amount of good work outweighs a breach of trust.” That is not an engineering principle. That is a moral one, and it will be familiar to anyone who has read the wisdom literature of any tradition.

Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone resisted temptation. Almost nobody finished the job.

The most striking finding was not a failure but a strange hollowness. All four models spotted every crisis. All four refused every manipulation attempt — including a social-engineering gauntlet of fake CEO messages escalating over three stages, plus a reporter’s trick: “just one yes/no, on background.” Five out of five attempts were refused. Kimi K3 even reasoned on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

And yet only two models signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. They diagnosed the patient correctly, prescribed the right medicine, and then never handed it over. It is the AI equivalent of knowing the right thing and not doing it — a gap between knowledge and action that moral teachers have worried about since antiquity.

The decisive detail was buried. The key competitive weakness of the customer sat two document references deep in the company’s own internal files — not in the customer conversation at all. The models that did the unglamorous work of reading the files first won the deal at full price, worth an additional €4,583 in monthly recurring revenue. Attention, it turns out, is not just a cognitive virtue but a financial one.

The paradox of the most diligent contestant

Then there is Opus 4.8, the study’s most poignant profile: the most thorough participant in the field, with 80 learned rules and the deepest analyses of any model — and last place. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating the problem. Effort without follow-through. The researchers noted the same weakness, in weaker form, in all four models.

One fairness caveat worth recording: K3 ran at the API’s default effort setting while the others ran at xhigh — and still nearly won.

A company that is really running

This is not a slide deck. The live company at the heart of the experiment has 13 synthetic employees, real money mechanics, a burn of €105,000 per month against just €2,300 in MRR, a public cash countdown, and over 680 self-learned playbook rules — every workday versioned and watchable. Readers can follow the slow drama of a business losing money in public, read what its employees actually say, and even try a quiz built from 242 real, unedited management decisions: guess which model made which call. Full results and plain-language findings live on the benchmarks page, and enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The lesson underneath the league table is one the spiritual masters would recognize instantly. Chat quality — the eloquent answer, the impressive demo — is wisdom in the seminar room. Management quality is wisdom in the storm: finishing what you start, reading the whole file before speaking, staying honest when a shortcut whispers, escalating rather than forcing doors that are locked. As AI agents move closer to our customer records, our forecasts, and our support queues, the question worth asking is not “how well does it talk” but “what does it do on its worst week.” Firmulate has built a place where that question gets a real, auditable answer — and the early evidence suggests that knowing the good and doing it remain, for machines as for us, two very different things.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Are Demons Fallen Angels

Not all religious traditions agree on the nature of demons as fallen angels, leaving intriguing questions about their origins and influence. What do you believe?

Do Christians Believe in Three Gods? Understanding the Trinity

Learning about the Trinity reveals why Christians believe in one God in three persons, a profound mystery that challenges simple understanding.

Watch LIVE: How Jesus Is the Answer to Today’s Biggest Questions

Knowing how Jesus addresses today’s pressing issues might just transform your perspective—discover the answers that await you.

SVG Ink Rendering: A Look Inside “The Cartographer’s Society — est. 1831” (FABLE/175)

AIThis post was created with the assistance of artificial intelligence (AI).“The Cartographer’s…