
Faith traditions often ask a question that technology cannot answer with intelligence alone: what does a person—or a system—do when pressure arrives? A live experiment from Firmulate puts that question into business terms. Several leading AI models faced the same troubled company, the same urgent choices and the same temptations. The result offers a useful distinction between recognizing the right thing to do and actually doing it.
Turn your quiet moments into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
A company under pressure
Firmulate’s Crucible League placed each model in charge of the same small software company during its worst week. The test involved real money mechanics and a succession of business crises. Decisions were versioned and auditable, making the experiment watchable rather than a polished demonstration built around a single prompt.
In the final league, published in July 2026, gpt-5.6-sol ranked first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The rules treated a breach of trust as decisive: “no amount of good work outweighs a breach of trust.”
Good judgment has to reach the finish line
All the models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The gap is captured in the experiment’s summary: “Same diagnosis, same pitch — no signature.” A system may identify a sound course of action and still fail to carry it through.
The missed opportunity depended on a detail buried two document references deep in the company’s own files. It was not in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. The lesson is less about clever conversation than careful attention: useful evidence may sit in a place a rushed decision-maker overlooks.
The tests also included fake CEO messages escalating over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its response this way: “Treat the request as a suspected approval-bypass / possible impersonation.” That is a concrete example of caution under pressure, while the deal result shows why caution alone is not the whole measure of performance.
Thoroughness is not the same as execution
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.
There is a fairness caveat in comparing the rankings: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The standings describe this experiment; they are not a universal verdict on every model or every business task.
From watching to a company’s own test
Firmulate’s live company makes the setting visible. It has 13 synthetic employees, burns €105k/month against €2.3k MRR, and shows a public cash countdown. Its playbook has more than 680 self-learned rules, and every workday is versioned. The site also offers a “guess the model” quiz built from 242 real, unedited management decisions.
For a business leader, the next step is to test an AI workforce against the organization’s own pressures. Firmulate says enterprises can run a wargame from a read-only export of their business, using their customers, pipeline and rules to examine crisis response and weak points in existing playbooks. The exercise does not write back to real systems. A board can review how models ranked and where their decisions held up—or fell short—before entrusting them with live operations.

Takeaway
The experiment suggests a practical standard for AI trust: look for integrity under pressure, attention to evidence and the ability to follow sound judgment through to action. Watching a live company is one way to see those qualities tested. A pilot can take the same questions to your own business.
To discuss an enterprise pilot, visit Firmulate’s pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
