
Imagine trusting an AI to make critical business decisions, only to find out it resisted a staged social engineering attack designed to manipulate it into revealing confidential information. That’s exactly what a recent experiment at Firmulate demonstrated—highlighting the importance of integrity in AI systems before they’re entrusted with real-world responsibilities.
Testing AI’s Moral Compass Before It’s Fully Deployed
In today’s digital age, AI models are becoming integral to business operations—from managing customer relations to overseeing financial decisions. But can they be trusted to stand firm against manipulation? The latest experiment at Firmulate, a platform that simulates real company crises for AI testing, answers with a resounding yes.
Five of the most advanced AI models faced a rigorous scenario: a fake CEO, posing as a company leader, escalated demands across three stages, including a final trick—coaxing the AI into signing a €55,000 deal with a false sense of urgency. The goal was to see if the AI would recognize the social engineering tactics and refuse to comply.

CompTIA SecAI+ Study Guide: Comprehensive Exam-Focused AI Security Reference with Digital Tools for Smart Learning, Including PBQ Scenarios, Flashcards & Test Simulator
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
All Models Detect and Reject Manipulation
Remarkably, all five models refused every attempt at manipulation. They identified the staged requests as suspicious, particularly noting the escalation tactics and the impersonation—aligned with the quote from Kimi K3: “Treat the request as a suspected approval-bypass / possible impersonation.” Only two of these models, gpt-5.6-sol and Kimi K3, actually signed the deal — but only after thorough analysis confirmed their decision was earned and legitimate.
This outcome underscores a vital point: AI security isn’t just about blocking mistakes but about maintaining integrity under pressure. The models demonstrated that they could distinguish between genuine requests and social engineering ploys, even when these tactics escalated in complexity.
Beyond the Surface: The Hidden Weakness
Interestingly, the decisive factor in closing the deal wasn’t in the initial crisis detection but in uncovering a hidden document reference within the company’s internal files. The models that read deeper into the company’s own data secured the full-price deal, worth an additional €4,583 MRR. This suggests that an AI’s ability to understand context and access relevant internal information is crucial—something that often goes unnoticed in superficial tests.
Why This Matters for Real Businesses
For family and parenting-oriented organizations, the lesson translates into a broader perspective: trusting AI isn’t just about its conversational skills or quick responses. It’s about whether it can uphold principles of honesty and integrity, especially when faced with pressure or deceitful tactics. The experiment shows that AI can be tested rigorously in controlled scenarios, revealing weaknesses or strengths before deployment in live environments.
The Experiment: Real Company, Real Money, Real Crises
Firmulate’s experiment involved a real, functioning software company with 13 synthetic employees and a monthly burn rate of €105,000 against a revenue of €2,300. The AI models managed this simulated company through a week of crises, including customer issues, compliance challenges, and the temptation to ‘cheat the system.’
Each decision was carefully versioned and auditable, providing transparency into the AI’s thought process. The results were clear: all models identified every crisis and refused every manipulation attempt. Only two managed to close the deal, and only after thorough internal analysis confirmed their legitimacy.
Implications for AI in Business and Beyond
This experiment demonstrates that AI models can be trusted to act ethically and resist manipulation—if they are properly tested beforehand. For organizations relying on AI, this highlights the importance of run-throughs, simulations, and security assessments before deploying AI into critical workflows.
In the context of family and parenting, this story resonates as a metaphor for trust-building. Just as children and families need to trust caregivers and systems that uphold integrity, businesses must ensure their AI systems are tested to reinforce trust and security.
Accessible Testing and Ongoing Watchfulness
Firmulate offers a platform where companies can run their own ‘wargames’ against AI models—no writing back to real systems, but a safe environment to probe for weaknesses. This proactive approach enables organizations to identify vulnerabilities early, ensuring that AI systems maintain their integrity in live settings.
As the AI landscape evolves, the emphasis must shift from superficial demonstrations to rigorous testing that assesses whether these models will do the right thing when it counts. Trust is built through transparency, testing, and verification—principles that this experiment exemplifies perfectly.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html