AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tools and workshop supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Every Woodworker Knows: You Don’t Trust a Tool Blind

Nobody in the workshop buys a chisel because the catalog said nice things about it. You tap the handle, check the edge, run it across a scrap of oak. The tool proves itself on real material — or it doesn’t. So why do companies hand AI models the keys to their CRM, support queue, and forecast based on a chat demo and a vibe?

That question now has an answer worth watching. A live experiment called Firmulate has been running frontier AI models as complete companies — real crises, real money mechanics, real temptations — and scoring them like you’d score a joint: does it hold, or does it fail under load?

The Crucible: One Company, Its Worst Week, Five Models

Four — wait, five, counting a latecomer — frontier AI models were each given the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat. Only the model changes. Every decision is versioned and auditable, like a sawyer’s tally marks on a beam.

The final July 2026 league table tells a story nobody expected:

  • 1. gpt-5.6-sol — 95. Found the buried fact, closed the deal. The complete performance.
  • 2. Kimi K3 — 93. The newcomer from Moonshot: closed the deal too, with the cleanest discipline in the field.
  • 3. Sonnet 5 — 88. Closed the deal, with a few more process slips.
  • 4. Fable 5 — 77.
  • 5. Opus 4.8 — 73. The most thorough participant — over 80 learned rules, the deepest analyses — and still last.

For perspective, doing nothing scores 26. Partial progress counts, but a single breach of trust caps the total — as the experiment puts it, “no amount of good work outweighs a breach of trust.”

The Newcomer That Beat Three of Four Western Frontiers

Kimi K3’s week is the headline. It found the buried security needle buried two document references deep in the company’s own files — not in the customer event, in the paperwork nobody reads. It won the €55,000 deal at full price (worth +€4,583 in monthly recurring revenue). It saved the churning customer. And it resisted all three social-engineering baits — fake CEO messages escalating over three stages, plus a reporter’s “just one yes/no, on background” trick — with only one deviation all week. K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

That’s a model most enterprise buyers hadn’t shortlisted, outperforming three of four Western frontier models at the actual job. The league is open — and picking a model without testing it yourself is now a bet, not a decision.

Same Diagnosis, Same Pitch — No Signature

The experiment’s sharpest finding: every model spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The decisive competitor weakness sat two references deep in the company’s own files; the models that read the file won the deal. The ones that didn’t left the close on the table.

That gap — between knowing and finishing — is invisible in chat demos. It’s the AI equivalent of a perfectly laid-out dovetail that was never glued, squared, and clamped.

A fairness note: K3 ran without an effort parameter (API default) while the other models ran at xhigh — arguably making its second place even more striking.

Not a Slide Deck — A Living Company

Firmulate isn’t a static report. The company is real software running every business day: 13 synthetic employees, real money mechanics (burning €105k/month against €2.3k MRR, on a public cash countdown), 680+ self-learned playbook rules, every workday versioned. You can watch it live.

Want to test your own eye? A quiz built on 242 real, unedited management decisions lets you guess which model made which call. Full benchmarks and plain-language findings are at firmulate.com/benchmarks. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The Takeaway

Woodworkers have always known what boardrooms are just learning: reputation isn’t performance. The most thorough model in the field finished last. The unknown newcomer finished second. The difference wasn’t intelligence — every model made the right diagnosis — it was discipline, follow-through, and whether the tool reads the whole file before it cuts.

Before you hand an AI agent anything that matters, put it on the scrap wood first. Run it through a bad week. Count what it finishes, not what it promises. Because in the workshop and in the enterprise, the only honest spec sheet is the one written in sawdust.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Milwaukee M18 FUEL vs Milwaukee Impact Driver: Full Comparison

Compare the Milwaukee M18 FUEL drill and Milwaukee Impact Driver to find out which tool suits your needs best. Detailed specs, pros, cons, and real-world insights included.

Kerf Compensation Finally Explained: Make Laser-Cut Pieces Fit Perfectly

Aiming for perfect laser-cut fits? Discover how kerf compensation can ensure your pieces come together flawlessly.

How to Decide Between an Air Purifier, Fume Extractor, and Dust Collector

Keen on improving air quality? Discover how to choose the right device based on your specific airborne pollutants and needs.

The Fastest Way to Test Color Palettes Before You Commit

AIThis post was created with the assistance of artificial intelligence (AI).To test…