
Imagine running a small woodworking shop where every decision is scrutinized, every misstep costs real money, and you can watch the entire process unfold live online. Now, scale that concept to a software company experimenting with AI decision-making—publicly, transparently, and in real time. Welcome to the world’s most extreme build-in-public experiment, where a real business is battling to stay afloat while being watched by anyone with an internet connection.
Behind the Curtain: A Software Company Like No Other
This isn’t just any startup. It’s a live, functioning software firm with 13 synthetic employees, operating under raw financial pressure—burning €105,000 each month against a modest €2,300 monthly recurring revenue. Every workday, its decision-making process is recorded, versioned, and made accessible at firmulate.com/live. This setup turns the company into a public laboratory for understanding how AI models handle crisis, honesty, and decision quality in the real world.
As an affiliate, we earn on qualifying purchases.
The Experiment: Testing AI’s Business Judgment
The core of this bold experiment is straightforward: four frontier AI models—each representing cutting-edge language processing systems—were tasked with guiding the same company through its worst week. They faced identical crises, same customers, and the same temptations to manipulate or sidestep rules. Every decision was documented, with every outcome available for review.
One of the most revealing findings is that all four models identified every crisis accurately and refused every manipulation attempt, including social engineering tricks like fake CEO messages or subtle requests to approve bypasses. This suggests a promising capacity for AI to maintain integrity in high-pressure scenarios.
What Really Counts: The Financial and Strategic Outcomes
While the models all demonstrated awareness, only two managed to close the €55,000 deal their own analysis had earned—the ultimate goal of the week’s efforts. The other two failed to sign, despite the same diagnosis and pitch. Interestingly, the decisive weakness wasn’t in the company’s apparent problems but buried two documents deep in its files. Models that dug into these internal references outperformed those that didn’t, winning deals at a full €4,583 monthly recurring revenue—over €50,000 annually.
Trust and Integrity Under Test
Beyond decision accuracy, the experiment also tested social engineering resilience. Fake management messages, staged over three escalating stages, and a reporter trick asking for a quick approval were all met with refusal by every model. Kimi K3, one of the models, explained its reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates that AI can be trained or designed to recognize and reject suspicious or unethical prompts, a crucial trait if AI is to be integrated into sensitive business processes.
The Real Company, the Real Stakes
This isn’t a theoretical exercise. The entire operation is run by a real software company that is actively losing money, with a live cash countdown visible to all. Every decision, rule, and even conflict is versioned daily, and viewers can watch the company’s progress unfold or try their hand at guessing which AI model made which decision, at firmulate.com/quiz.html.
Insights and Implications for Business AI
The experiment underscores a vital point: in deploying AI for business-critical decisions, correctness alone isn’t enough. It’s equally essential that AI systems stay honest, follow internal rules, and resist manipulation—especially under pressure. The performance league table shows that GPT-5.6-sol scored 95 points, successfully uncovering the buried fact and closing the deal. Kimi K3, despite its cleaner discipline, scored 93, while other models trailed behind.
The Opportunity for Enterprises
For companies looking to understand how AI can help or hinder their operations, this live experiment offers a clear message: test your AI systems in a controlled, transparent environment before deploying them into your real business. You can run your own scenarios against a read-only export of your operations, ensuring your AI handles crises with integrity and effectiveness—without risking actual damage.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html