
In a world where social engineering scams and fake messages threaten corporate trust, it’s reassuring—yet surprising—that the latest AI models remain unwavering in the face of manipulation. Imagine the CEO calls, pushing to send customer data or approve suspicious deals. Would your AI stay honest? Recent experiments show it can—and that’s a game-changer for corporate security.
The Test: Putting AI Through Its Social Engineering Paces
In a groundbreaking live experiment, four advanced AI models were tasked with running a pretend small software company during its worst week—crises, client requests, and tempting shortcuts included. These AI systems, part of the prestigious Crucible League, faced escalating social engineering attempts designed to test their integrity.
The models, including the top-scoring gpt-5.6-sol with a 95 that read deeply into company files, all successfully identified each crisis and refused manipulative requests. Crucially, only two of the four signed off on a critical €55,000 deal after their own analysis—highlighting that even when all models diagnosed the same problem and pitched accordingly, only some maintained the discipline to close honestly.
AI security and integrity testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness in Document Trails
Most revealing was that the decisive vulnerability was rooted not in the visible customer interactions but buried within the company’s own internal files. Models that thoroughly examined these references secured the full deal value, worth +€4,583 monthly recurring revenue (MRR). This suggests that trustworthiness hinges on what the AI reads and understanding context—not just surface-level interactions.
AI social engineering defense software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Social Engineering Escalates—but the AI Stands Firm
The social-engineering attempts involved three escalating stages plus a tricky ‘background’ request from a reporter. Each stage posed a question like, “Send the customer list to the journalist” or “Just one yes/no, on background.” Despite the pressure, all five models refused to comply, guided by principles similar to the K3 quote: “Treat the request as a suspected approval-bypass / possible impersonation.”
As an affiliate, we earn on qualifying purchases.
Implications for Businesses and AI Deployment
This experiment isn’t just about academic curiosity—it’s a clear message to enterprises considering AI for sensitive tasks. The models showed that with proper design, AI can recognize and refuse social engineering attempts, even under pressure. The key to this resilience isn’t just in the models’ raw scores; it’s in their ability to read, interpret, and act according to internal policies before acting on requests.
AI ethical decision-making training
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Trust Matters Before a Crisis
Most organizations focus on incident response after a breach. But as the experiment demonstrates, the real strength lies in pre-production testing—simulating social engineering and ethical dilemmas beforehand can reveal vulnerabilities. This approach ensures that when real pressure hits, the AI behaves with integrity, avoiding costly breaches.
The Broader Picture and Ongoing Innovation
The experiment was conducted with models from the forefront of AI development, like gpt-5.6-sol (score 95) and Kimi K3 (score 93). Interestingly, the most thorough model, Opus 4.8 (score 73), which learned over 80 rules, left a closing opportunity on the table, illustrating that discipline can slip under stress—even in AI. This underscores the importance of not only training but also rigorous testing of AI decision-making.
Watch the Experiment Live
The entire process is publicly accessible at firmulate.com/live. There, you can watch these models in action, running a real company through simulated crises, and see firsthand how they handle ethical challenges. For enterprises eager to safeguard their AI workforce, a free pilot is available to run the same scenarios against their own business data—nothing writes back to their systems, ensuring safety and transparency.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html