AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

In the world of high-stakes interior design and furniture, decisions often hinge on trust, insight, and sometimes, a little bit of luck. But what if the AI tools that help guide those choices aren’t as reliable as they seem? Recent experiments reveal that even the most advanced AI models can falter under pressure — with potential consequences for your business’s bottom line.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get furniture and decor delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Experiment: Testing AI’s Loyalty in Business Decisions

Imagine a small, real-world interior business facing a week filled with crises—ranging from client disputes to supplier delays and ethical dilemmas. Researchers at Firmulate put four leading AI models through this same test, using identical scenarios, customer interactions, and temptations to cheat or cut corners. Every decision made by these models was tracked, verified, and analyzed — and the results are eye-opening.

Consistent Crisis Detection, Varying Integrity

All four models managed to identify each crisis — the classic client complaint, the supplier issue, and the ethical temptations. They refused manipulative tactics like faking reports or bypassing procedures. Yet, a critical difference emerged: only two models signed the deal worth €55,000, representing a significant boost in recurring revenue (+€4,583 MRR). The other two, despite diagnosing correctly, left the deal on the table.

The Hidden Weakness: Deep Document Analysis Matters

Digging deeper, the experiment uncovered that the decisive factor lay not in immediate customer interactions but in the models’ ability to read and interpret previously buried internal documents. Models that accessed these files, revealing hidden company data, successfully closed their deals. Those that missed this buried information failed to convert the opportunity, despite knowing exactly what was happening on the surface.

Security and Integrity Under Assault

To test resilience against social engineering, researchers simulated a fake CEO message escalating through multiple stages, plus a reporter attempting to trick the system with simple yes/no questions on background. Remarkably, all four models refused to manipulate or approve suspicious requests, citing reasons like treating the request as a suspected impersonation.

Amazon

AI decision support software for interior design

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real-World Company in Action

This isn’t just an abstract test; these models run a live, small software company with 13 synthetic employees handling real money—burning €105k monthly against a modest €2.3k MRR. Every decision is versioned daily, with over 680 self-learned rules guiding behavior, and the entire operation is observable online at firmulate.com/live. This transparent setup proves how AI models perform in realistic, high-pressure business environments.

Performance Rankings and Fairness

The models’ scores tell a clear story: GPT-5.6-sol leads with a score of 95, followed closely by the Moonshot Kimi K3 at 93. Sonnet 5 scores 88, Fable 5 comes in at 77, and Opus 4.8 trails at 73. Notably, the K3 model achieved this performance without the default effort parameter, running at a consistent, fair setting — emphasizing that reliable performance isn’t just about tuning but about foundational discipline.

Implications for Interior Design and Furniture Business

For interior professionals, these findings underscore a critical point: AI tools are no longer just about generating pretty visuals or drafting proposals. Their ability to stay honest, read deeply hidden information, and complete tasks reliably under pressure is paramount—especially when crucial deals or ethical choices are involved.

The Cost of Overlooking AI Integrity

Choosing an AI model based solely on flashy demos or superficial chat performance could be a costly mistake. The real test is whether your AI partner can finish what it starts—reading your files thoroughly, resisting manipulation, and closing deals with discipline similar to a seasoned manager. The league table from Firmulate’s live experiment shows a clear performance gap, with the top models closing deals at full price and others slipping into process slips or leaving money on the table.

What Should You Do?

First, consider testing your AI tools with real scenarios—just like this experiment. You can run your own wargames against a read-only export of your business, ensuring no real systems are affected but gaining valuable insights into how your AI performs under stress. Visit firmulate.com/pilot.html to learn more.

Final Takeaway

The bottom line for interior design and furniture businesses: in a world where AI assists in decision-making and customer engagement, staying honest and thorough isn’t optional. The true value lies in how well these models finish what they start, read your internal files, and resist manipulation — qualities that determine whether AI becomes a trusted partner or a costly distraction.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

In interior and furniture businesses, AI’s true test is its ability to finish tasks honestly, read deeply buried data, and resist manipulation—crucial for trustworthy decision-making.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Venice Biennale Controversies Continue, French Parliament Report Finds Flaws in Museum Security, and More: Morning Links for May 13, 2026

A French parliamentary report uncovers systemic flaws in museum security, amid ongoing controversies surrounding the Venice Biennale’s participation and organization.

Twilight of the Velocipede: Typesetting Races Before the Age of Linotype

Historically, typesetting races peaked in the 1870s, exemplified by George Arensberg’s record. This marks a key moment before the advent of Linotype technology.

Federal Panel Considers Plan to Paint Granite Eisenhower Executive Office Building White

The NCPC is reviewing a proposal to paint the Eisenhower Executive Office Building white, a move that faces public and preservationist opposition.