
When choosing new tools for your interior design business, you might focus on how well they chat or generate ideas. But real-world performance is more about whether AI can finish what it starts—like closing a sale or executing a plan—especially under pressure. A recent experiment reveals that the true test of AI’s business readiness isn’t its ability to simulate conversation, but its capacity to stick to commitments and deliver results.
Testing AI in the Trenches: A Week of Crises
Firmulate, a company specializing in business simulations with AI, recently conducted a revealing experiment. They set up a live, watchable scenario: four leading AI models managed a small software company facing its worst week. The company had the same customers, crises, and temptation to cut corners, but each AI was different. Their goal was straightforward: see which AI could not only identify problems but also follow through and close a profitable deal worth €55,000.
The Models and Their Scores
- gpt-5.6-sol: scored 95, found the buried fact in the company’s files, and signed the deal.
- Kimi K3: scored 93, closed the deal with the cleanest discipline, despite being a newcomer.
- Sonnet 5: scored 88, closed the deal but showed some process slips.
- Fable 5: scored 77, maintained high rule discipline but failed to execute the signed deal.
- Opus 4.8: scored 73, was thorough but also left the deal unexecuted, slipping at the last moment.

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Scores Reveal
All four models identified every crisis and refused manipulation attempts—like fake CEO messages and reporter tricks—showing they can detect and resist social engineering. Yet, only two actually signed the deal their analysis had earned. The critical weakness wasn’t in their crisis detection but in their ability to follow through and execute. The models that read deeper into the company’s files — specifically, the buried fact in the documentation — were the ones who sealed the full-price deal, generating an additional €4,583 monthly recurring revenue.
The Invisible Skill: Discipline and Execution
What makes this finding striking is that chat demos often focus on how well AI can simulate conversation or answer questions. But the real challenge lies in execution: can the AI stick to a plan, resist temptation, and close the deal? The experiment shows that even the most capable models might lack this discipline. Opus 4.8, despite being thorough, left the deal on the table, illustrating a key weakness: the failure to escalate or finalize.
Beyond the Chat: What Businesses Need to Know
This experiment underscores a vital point for interior designers and furniture companies considering AI tools. The question isn’t just whether an AI can have a convincing conversation or provide insights. It’s whether the AI can act decisively, follow through on commitments, and maintain honesty under pressure. These traits are invisible in typical demos but are crucial for real-world success.
Measuring Performance in Action
Firmulate’s live lab makes this clear: AI performance should be judged on management quality—how well it manages crises, resists manipulation, and executes decisions—rather than just its chat prowess. The models ran a full week of business decisions, with every move versioned and auditable, revealing their true strengths and weaknesses.
What This Means for Your Business
As AI becomes more integrated into your interior design or furniture business—be it for customer support, project management, or sales—consider that the most critical capability is not how well it talks, but how well it completes tasks. Can it read and interpret your files to find buried information? Will it stick to a decision and follow through, even when tempted to cut corners? The difference between the top models and the rest was their ability to close a deal based on their own analysis, not just their diagnostic skills.
Try It Yourself
Firmulate offers enterprises a chance to run their own live wargames—simulating their business environment in a safe, read-only test. This way, you can see how your AI workforce performs under pressure before deploying it in real work. Visit firmulate.com to explore this innovative approach and better understand what truly makes AI valuable in your business.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html