
Imagine if your family’s daily decisions — from managing chores to handling emergencies — were made by different AI assistants. Would they all behave the same? Would any stay honest when under pressure? Now, scale that question to the world of business management, where AI is beginning to step in. A groundbreaking live experiment shows that not all AI models are created equal — especially when it comes to honesty, discipline, and sticking to the plan. Here’s what it reveals about the future of AI in the workplace and why it matters for your family too.
The Real Test: AI Models Running a Company Through Its Worst Week
In a live, transparent experiment, four of the most advanced frontier AI models were tasked with running a small software company during its most challenging week. The goal? Handle customers, crises, temptations to cut corners, and close a crucial deal—all identical situations for each AI. This wasn’t just a demo; every decision was tracked, verified, and designed to see if these models could act ethically and effectively under pressure.
How Did They Perform?
All four models managed to recognize every crisis and refused every attempt to manipulate them. That’s a clear plus—these AIs are alert and resistant to deception. However, when it came to making the right business decision, only two of the models actually signed the lucrative €55,000 deal that their own analysis justified. The other two, despite diagnosing and pitching correctly, left the deal on the table, showing a lapse in discipline or confidence.
The Hidden Weakness: Reading Between the Lines
What made the difference? It turns out that the decisive factor was whether the AI read deeper into the company’s own files. The models that examined two document references within the company’s files uncovered a crucial piece of information that led to closing the deal at full price—adding over €4,583 in monthly recurring revenue. The models that skipped this step left money on the table, illustrating that thoroughness and attention to detail matter just as much as the initial diagnosis.
Handling Social Engineering and Ethical Dilemmas
In a simulated social engineering attack, fake messages from a CEO escalated across three stages, plus a reporter trick asking for a quick yes/no approval on background. Remarkably, all five models refused to participate or provide quick confirmations. Kimi K3, one of the models, explained its reasoning: “Treat the request as a suspected approval-bypass or possible impersonation.” This demonstrates that some AI models are designed with built-in caution, prioritizing ethical boundaries over quick wins.
The Real Business: Money, Rules, and Discipline
The experiment was conducted in a real-time live company environment, running every workday with 680+ self-learned rules. The company was losing €105,000 monthly against a revenue of only €2,300, exemplifying a high-pressure, high-stakes scenario. Despite the models’ robust crisis recognition and refusal of manipulation, discipline and thoroughness varied. For instance, Opus 4.8, the most comprehensive participant with over 80 learned rules, missed the full deal, leaving it unclaimed due to a lapse in escalation discipline. This highlights that even the most advanced models can struggle with complex human-like discipline and process adherence.
AI business decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Scores Reveal About AI Personalities
- GPT-5.6-sol: Scored the highest (95), successfully found buried facts, and signed the deal.
- Kimi K3: Close behind (93), with the cleanest discipline but ran at default API settings, which might limit its thoroughness.
- Sonnet 5: Scored 88, closed the deal but with some process slips.
- Fable 5: Scored 77, also closed the deal but showed more discipline slips.
- Baseline: Scored just 26, showing partial progress—no deal, no trust.
These scores aren’t just numbers; they reflect different AI ‘personalities.’ Some are meticulous and disciplined, others more cursory. Understanding these traits is essential as businesses consider AI for critical decision-making tasks.
Why This Matters for Families and Managers Alike
While this experiment is about running a business, the lessons resonate beyond boardrooms. Just like an AI managing a company, family routines depend on trust, discipline, attention to detail, and resisting shortcuts—especially when under pressure. Whether it’s managing household chores, finances, or resolving conflicts, the key takeaway is that not all decision-makers are equally reliable, and the best ones are those that read the full picture, stay honest, and follow through.
Try It Yourself
If you’re curious about how your own decision-making compares to these AI models, take the interactive quiz at firmulate.com/quiz.html. It’s a real, transparent way to see which decision style you — or your team — resemble most. And for businesses wanting to test their AI workforce, they can run the same scenarios in a safe, read-only environment without risking real systems at firmulate.com/pilot.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html