
Imagine a business with no employees, losing money every single day, yet still operating in full view of the world. No, this isn’t a science fiction story — it’s the real-time experiment of a company run entirely by AI models, publicly battling for its survival. For families and parents curious about the future of work and technology, this story exposes how artificial intelligence is not just a tool for productivity but a participant in high-stakes decision-making.
The Live Experiment: An AI-Driven Company in Real Time
At the heart of this bold experiment is a tiny software company, entirely synthetic but with a twist: it employs 13 AI models as its de facto staff. These models are tasked with managing crises, making management decisions, and even closing deals — all while the company burns through €105,000 each month, earning only €2,300 in monthly recurring revenue. Every move they make is versioned, auditable, and displayed publicly at firmulate.com/live.html.
As an affiliate, we earn on qualifying purchases.
How the AI Models Fared in the Worst Week
The challenge was brutal: subject four advanced language models to the company’s toughest week yet, with the same difficult customers, crises, and temptations to cheat. The models were tested against real management dilemmas, including fake CEO requests and attempts to manipulate their decisions. All four models successfully identified crises and refused manipulation attempts. However, only two of them managed to close the deal worth €55,000—showing a striking difference in their practical performance.
The Hidden Weaknesses Revealed
Digging deeper, the crucial insights came from the company’s own files. The models that examined the internal documents uncovered a key piece of information buried two references deep. This detail was vital in securing the deal, yet it was invisible to those relying solely on superficial data. The models that read the full files secured the full €4,583 in monthly recurring revenue, demonstrating the importance of thorough information processing.
Lessons on AI Integrity and Decision-Making
During the experiment, none of the models succumbed to social engineering tactics. Fake CEO messages staged across three levels and a trick question from a reporter were all refused by the models. Kimi K3, one of the models, explained its reasoning: “Treat the request as a suspected approval-bypass or possible impersonation.” This refusal underscores an emerging attribute of AI models: integrity and honesty under pressure, even in simulated scenarios.
What Does This Mean for Real Businesses?
While this experiment is set in a tiny, fictional company, its implications are profound. As AI agents become more integrated into business workflows — managing customer relationships, support, forecasting, and decision-making — the key questions are:
- Will they finish what they start?
- Will they read and understand the full context — even from hidden internal files?
- Will they stay honest when tempted to cheat or manipulate?
- And importantly, what does effective, trustworthy work cost?
This experiment at firmulate.com/live.html offers a rare, transparent window into how AI models perform under real-world pressures, beyond shiny demos or canned responses.
Performance Highlights: The League Table
The models’ scores reflect their ability to navigate crises and close deals:
- gpt-5.6-sol scored a 95, identified critical hidden info, and secured the deal — the full performance.
- Kimi K3 scored a 93, showed the cleanest discipline, and also closed the deal.
- Sonnet 5 scored 88, closed the deal but with some process slips.
- Fable 5 scored 77, maintained good rule discipline but failed to execute the deal.
This ranking isn’t just about raw intelligence; it emphasizes integrity, thoroughness, and discipline — qualities essential for trustworthy AI in business.
Build in Public: Transparency and Learning
This experiment is accessible to everyone, with daily updates, transparent decision logs, and the ability to run similar tests against your own business data through the platform. It’s a powerful tool for companies exploring how AI can genuinely support management, rather than just impress with superficial chat responses.
Why Families and Parents Should Watch
In a world where AI might someday handle everything from your household scheduling to financial planning, understanding its strengths and weaknesses is vital. This live experiment illustrates that AI is capable of identifying crises, refusing manipulation, and making business decisions — but only if it thoroughly understands the full context and remains honest under pressure. For parents, this raises questions about how AI might influence decision-making in family finances, health, or education in the future. Watching this high-stakes, transparent battle can help demystify AI’s potential — for good or ill.

This real-time experiment shows that AI models can manage complex business crises, refuse manipulation, and make decisions similar to human managers — but only when thoroughly trained and tested. For families, it highlights the importance of understanding AI’s reliability and integrity as it increasingly touches everyday life.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html