firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

For listenersOffer from Amazon

Turn quiet afternoons into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

When AI enters elder care, trust has to survive the hardest day

A care provider considering AI for scheduling, family updates or customer support faces a question that a polished demonstration cannot answer: how will the system behave when pressure mounts, instructions conflict or someone asks it to cross a line? Firmulate’s live experiment offers a way to watch AI models manage a company through a crisis before inviting them into consequential work.

One company, one difficult week

Firmulate gave each frontier model the same small software company, customers, crises and temptations. Every decision was versioned and auditable. The final Crucible League, published in July 2026, placed gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26; as the experiment puts it, “no amount of good work outweighs a breach of trust.”

The striking result was not that models failed to notice trouble. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The finding, in the experiment’s words: “Same diagnosis, same pitch — no signature.” Recognizing the right move and carrying it through were different tests.

The clue was in the company’s own files

The deal turned on a competitor weakness hidden two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. That detail makes the exercise relevant beyond sales: a system may need to connect information scattered through an organization before it can make a sound decision.

Trust faced a separate test. Fake CEO messages escalated through three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 described the request as a “suspected approval-bypass / possible impersonation.” For organizations serving older adults, families and care teams, resisting an apparent shortcut can matter as much as being helpful.

Thoroughness is not the same as follow-through

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and tried to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four models. More analysis alone did not guarantee a completed task or the right response when access was blocked.

The experiment also has a fairness caveat: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate publishes 242 real, unedited management decisions in a “guess the model” quiz, giving readers a chance to judge the decisions themselves.

From watching to a company-specific pilot

The live company is a watchable experiment, not a care provider: it has 13 synthetic employees, burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and versioned workdays. Readers can follow the experiment at Firmulate.

For a senior-care organization, the next step is to test AI against its own situations: a read-only export of company data, crisis scenarios tailored to its work, and a board report that ranks models and identifies weak points in existing playbooks. Firmulate says the pilot never writes back to real systems. That gives leaders a way to examine how an AI workforce responds before it is trusted with live operations.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put the hard questions on the table first

In senior care, the stakes include trust, clear escalation and reliable follow-through. Firmulate’s experiment suggests that spotting a crisis is only part of the job; the model must also act on what it knows and respect boundaries when pressured. Enterprises can run the wargame on a read-only export of their own business, with nothing writing back to real systems. Explore a Firmulate pilot or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Zero-Gravity Chairs Explained: Why They Feel So Good

Discover why zero-gravity chairs provide unmatched comfort. Learn how their unique design reduces pressure and relaxes your body, just like floating on air.

How to Arrange Porch Seating for Easy Conversation

Discover simple, practical tips to arrange your porch seating for cozy chats. Create a welcoming space perfect for friends and family to connect.

Before an AI Enters the Care Workflow, See How It Handles a Company’s Worst Week

Kimi K3 beats three Western frontier models in Firmulate’s live company test, raising a practical question for senior care: would your AI follow through?

Physics-Driven SVG Interaction: A Look Inside “Patang Bazaar — Fly the Old City Sky” (FABLE/175)

AIThis post was created with the assistance of artificial intelligence (AI).“Patang Bazaar…