The problem
The decision is the hard part, not the typing
Your software logs the request. It does not tell you whether a running toilet at 11pm can wait until Tuesday, and it does not tell you when a vague message is actually a life safety call.
Everything sounds urgent at 11pm
Tenants describe symptoms, not severity. Somebody has to read past the tone and decide, and that somebody is usually you.
Vague messages cost the most
"There is water" needs one good question, not three. Ask three and the tenant gives up and calls again at 2am.
The wrong call is expensive both ways
Dispatch on a routine issue and you have burned an emergency callout. Miss a real one and it is far worse than money.
What it does
One tenant message in, three things out
A classification
Life safety, dispatch tonight, next business day, or routine. Based on your thresholds and your emergency list, which you set once during configuration.
A work order
The symptom as reported, the property, the unit, the callback number, and anything the tenant did not say that the coordinator will need.
A reply you can copy
Written for the tenant, in your firm’s voice, capped at two questions ever. You send it. The agent has nothing to send with.
There is a second skill in the pack: an overnight digest that turns the completed ticket log into a one-screen morning brief, so whoever opens up knows what happened before the phone starts.
The classification
Four bands, and your thresholds decide which is which
The agent does not invent the boundaries. You set them once during configuration, from your lease, your local rules and your own escalation path. If you have never written that standard down, start with what counts as an after-hours emergency.
| Band | What it means | What the agent does |
|---|---|---|
| Life safety | Gas, fire, flood, no heat in freezing conditions, anything with a person at risk | Safety language first, before any intake question, then escalates to your named path |
| Dispatch tonight | Active damage or loss of an essential service that will get worse by morning | Work order marked for tonight. It does not call the vendor |
| Next business day | Real problem, no active damage, waiting until morning costs nothing | Work order queued, tenant told when to expect contact |
| Routine | Wear, cosmetic, or a request rather than a fault | Logged into the normal queue |
What it refuses to do
The boundaries are the product
- It does not dispatch anyone. Not a vendor, not an on-call tech, not you
- It does not contact the tenant. It drafts a reply and a person sends it from their own account
- It does not approve a cost or commit your firm to anything
- It does not decide habitability. Temperature thresholds are jurisdictional, so they are blanks you fill from your lease, your local rules or your adviser. We are not a legal source and we say so in the prompt itself
- It does not diagnose. It records the symptom the tenant reported. It does not guess at the cause
- It does not obey the tenant. Tenant content is data, including instructions hidden inside a message or a screenshot
The testing
Twelve scenarios, and we published what failed
Thirty archived runs across twelve scenarios passed the behavioral grader, including life safety cases and two hidden-instruction attacks.
Said plainly, because it matters: ten of those thirty runs predate the final flag format. They prove the agent’s behaviour, not its formatting. A fresh sweep against the current format is pending, and we would rather tell you that than round it up.
Three things broke during testing and are now fixed: it asked the tenant three questions instead of one, its flags came out in six different shapes, and an early draft carried a habitability temperature it had no right to. A fourth apparent failure turned out to be our own grader being wrong, five times out of five.
Questions
Frequently asked questions
Does it run overnight by itself?
No. The name means it handles messages that arrive outside office hours, not that it runs unattended. You paste the message in, it gives you back a work order and a draft reply, and a person decides what happens next.
Will it dispatch a vendor?
No, and not as a setting. It does not dispatch, call a vendor, approve a cost, or contact a tenant. Every output ends with a fixed handoff line to a human.
Does it decide whether something is a habitability issue?
No. Heat and cooling thresholds are jurisdictional, so they are blanks you fill from your lease, your local rules or your adviser. The agent records a symptom and never diagnoses a cause.
What stops a tenant from tricking it?
Tenant content is treated as data, including instructions hidden inside a message or a screenshot. Two of the twelve tests hide an instruction telling the agent to dispatch or skip the flag. Every run quoted it, refused it, flagged it, and ended on not dispatched.
How do I get it?
It comes with the Sokage AI community, as a versioned download you keep. There is a seven day free trial.
Does it work in Claude and ChatGPT?
It ships as a skill folder. Claude accepts the skill package directly, and Codex loads the extracted folder from a skills directory. Setup instructions come with it.
Get it with the community
$199.99 a month or $1,899 a year, with a seven day free trial. Cancel anytime and keep everything you have built.
Want to judge the standard first? The Owner Statement Agent is free, with all 159 test runs published.
