
Imagine if your favorite woodworking project was judged not just by how it looks but by how well you handle unexpected challenges—like a sudden crack or a missing tool. Now, picture AI models running a small, real business facing its toughest week, where every decision matters just as much. This is not a sci-fi scenario but a live experiment that reveals the true personality of AI managers, and what it means for your own work or business.
The Experiment: Putting AI to the Test in a Real Business
In a groundbreaking live setup, four frontier AI models were tasked with managing a small software company during its most turbulent week. The scenario was intense: the same customers, the same crises, and the same temptations to cut corners or manipulate outcomes. Each AI was given the same set of challenges, and every decision was logged and auditable, allowing a clear view into their management styles and integrity.
What the AI Models Were Tested On
- Handling crises with customers
- Responding to internal manipulations and deception attempts
- Deciding whether to close a lucrative deal
- Reading and analyzing critical internal documents
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Findings: Honesty, Discipline, and Decision-Making
All four models successfully identified every crisis and refused every attempt at manipulation. For example, when fake CEO messages escalated over three stages, all models refused to engage or approve unverified actions, with one model explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
However, when it came to closing a critical deal worth over €55,000, the results diverged. Only two models signed the deal, even though all had arrived at the same diagnosis and pitch. The models that closed the deal were GPT-5.6-SOL and Kimi K3. Their success hinged on a key detail buried two document references deep in the company’s internal files—something that the other models missed.
The Hidden Weaknesses
The models that failed to close the deal left it on the table, showing cracks in their discipline. For instance, Opus 4.8, the most thorough participant with over 80 learned rules and deep analyses, also slipped by leaving the close opportunity unclaimed. Notably, all models showed a similar weakness: they hesitated or failed to escalate certain internal issues, revealing a vulnerability to incomplete information.
Personality Profiles: Different Models, Different Leadership Styles
The experiment underscores that AI models exhibit measurable management personalities. While all are capable of spotting crises and resisting manipulation, their decision patterns differ significantly:
- GPT-5.6-SOL: The top performer, reading deeply into internal files and closing high-value deals without hesitation.
- Kimi K3: The newcomer with a focus on fairness and discipline, closing deals cleanly without effort parameters.
- Sonnet 5: Consistent but slightly less disciplined, closing the deal with minor slips.
- Opus 4.8: Most thorough but prone to leaving opportunities unclaimed due to over-analysis or reluctance to escalate.
Implications for Your Business: Trust and Performance Matter
This live experiment reveals a crucial point: measuring an AI’s ability to finish what it starts, read internal information thoroughly, and stay honest under pressure is key to deploying AI effectively in your organization. It’s not just about how well an AI writes or responds in chat but whether it can be trusted to make critical decisions, especially in high-stakes scenarios.
Try It Yourself: Run the Same Wargame Against Your Business
Businesses interested in understanding their own AI readiness can run the same type of simulation. Using a read-only export of your operations, you can test how your AI models handle real crises without risking actual systems or data. Discover which model aligns best with your values and needs by trying the interactive quiz and live experiment at firmulate.com/quiz.html.

AI models can behave like different types of managers—some disciplined and trustworthy, others prone to slips. Testing them in real business scenarios reveals their true personality and readiness, helping you choose the right AI partner before deployment.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html