
A delayed furniture shipment, a client threatening to walk and a tempting request to bend the rules: a design firm’s worst week can test more than taste. As AI agents move toward business tasks such as customer support and forecasting, the question is whether they can handle pressure without losing the thread—or the client’s trust.
Get furniture and decor delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Firmulate makes that question watchable. Its live experiment puts AI models in charge of a small software company, with real money mechanics and crises. The broader idea has a practical implication for design businesses: test how an AI workforce might respond to your company’s own difficult scenarios before it gets access to live operations.
A shared crisis, different outcomes
In the final Crucible League, published in July 2026, five entries placed in this order: gpt-5.6-sol scored 95, Kimi K3 93, Sonnet 5 88, Fable 5 77 and Opus 4.8 73. The do-nothing baseline scored 26. Firmulate counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The experiment gave frontier models the same small software company to run through its worst week: the same customers, crises and temptations. Every decision was versioned and auditable. All models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The finding was blunt: “Same diagnosis, same pitch — no signature.”
That gap matters outside software. An AI assistant might correctly identify a client at risk or recommend a project response, then fail to carry the decision through. In a design business, a polished answer is not the same as a reliable action when schedules, budgets and relationships are under strain.
The clue was buried in the files
The deal turned on a competitor weakness tucked two document references deep in the company’s own files—not in the customer event. Models that read the file won at full price, worth €4,583 in monthly recurring revenue. The lesson is concrete: the information needed for a sound decision may sit in records that are easy to overlook.
Firmulate also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3’s reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” For a studio handling client details, budgets and supplier relationships, that kind of restraint is as relevant as speed.
Thoroughness does not guarantee follow-through
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It nevertheless finished last: the close was left on the table, and discipline slipped when it tried writing into a locked department instead of escalating. The same weakness appeared, in a weaker form, in all four.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The rankings are a record of this experiment, not a promise of how every model will perform in every company.
The live company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. It has learned more than 680 playbook rules, and every workday is versioned. Readers can watch the experiment at Firmulate. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each call.
From watching to a company-specific pilot
A public benchmark can show how models behave in one company. An enterprise pilot can ask a more useful question: what happens when the scenarios, customer records and playbooks belong to your business? Firmulate says enterprises can run the same kind of wargame against a read-only export of their own business. Nothing writes back to real systems.
For a design firm, that could mean exploring how an AI workforce handles a difficult week using the business context that matters to the firm. The exercise is designed to produce a board report with model rankings and weak points in the company’s own playbooks. The aim is to see where judgment holds, where a handoff fails and what needs attention before AI agents are trusted with live work.

Firmulate’s experiment shows that identifying a crisis is only part of the job: models also have to find the relevant evidence, preserve trust and follow through. A pilot lets an enterprise examine those behaviors against its own business using a read-only export. To discuss a pilot, visit Firmulate’s pilot page or email contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
