What to do when your client's AI breaks at 2 a.m.
A no-panic playbook for AI builders when a client system goes down in the middle of the night.
What to do when your client's AI breaks at 2 a.m.
A no-panic playbook for AI builders when a client system goes down in the middle of the night.
You built something smart for a client — an AI workflow, a chatbot, an automated pipeline. It worked great in testing, the handoff went smoothly, and then your phone lights up at 2 a.m. Something broke. The client is panicking. You're half-awake and squinting at a screen.
This is the moment that separates scrappy freelancers from operators clients actually trust. Here's how to move through it without losing the relationship — or your mind.
1. Respond fast, fix slow
The single most important thing you can do in the first five minutes is acknowledge the problem — even if you have no idea what caused it yet. Send a message. Something like: "I see it, I'm on it. I'll update you in 20 minutes."
Clients don't expect you to be magic. They do expect you to be present. A quick response buys you the time and goodwill to actually diagnose the issue without someone breathing down your neck while you work.
2. Roll back before you dig in
Before you start poking at logs or tweaking prompts, ask yourself: was anything changed in the last 24–48 hours? A new API version, a prompt edit, a third-party integration update, a model deprecation from the provider — these are the usual culprits.
If you can identify the last working state, restore it first. A boring, working system beats a clever, broken one every time. Rolling back gets the client operational again while you figure out the root cause in daylight.
3. Know your failure layers
AI systems break in predictable places. When you're diagnosing, work through these in order:
- API / provider layer — Is the underlying model or service (OpenAI, Anthropic, Google, etc.) having an outage? Check their status page before you do anything else.
- Data or input layer — Did something upstream change? A form field, a webhook payload, a database schema?
- Prompt or logic layer — Did a prompt start producing unexpected outputs? Model behavior can shift with provider updates even when nothing on your end changed.
- Integration layer — Is the connector between systems (Zapier, Make, a custom script) timing out or hitting rate limits?
Narrowing the layer quickly tells you who owns the fix and how long it's likely to take.
4. Communicate like a pro, not a developer
Once you know what's going on, tell your client in plain language — not in stack traces. "The AI writing tool is down because the company that powers it (OpenAI) had a service outage at 1:47 a.m. I've confirmed it's on their end. Their status page shows they expect it resolved by 4 a.m. I'll monitor and confirm when you're back up."
That message takes 60 seconds to write. It tells the client exactly what happened, who owns it, and what happens next. It also signals that you're watching, even if there's nothing you can personally fix. Clients remember this kind of communication more than they remember the outage itself.
5. Build the safety net before the next 2 a.m.
Once the fire is out, use the incident to justify the infrastructure your clients should have had from day one. A post-mortem doesn't have to be formal — a short voice note or a 10-minute call works fine. Cover:
- What broke and why
- What you're adding to prevent or catch it faster next time (status alerts, fallback logic, a manual override)
- Any changes to your support terms or retainer scope
Uptime monitoring tools like Better Uptime or UptimeRobot take about 15 minutes to set up and will text you before your client even notices an issue. Automated alerts mean the next outage starts with you calling them — not the other way around.
Late-night incidents are a fact of life when you build AI systems. But the operational overhead around them — client communication, status updates, monitoring, incident write-ups — doesn't have to eat your time or your sleep. That's exactly the kind of ongoing work you can delegate to a Sidekyk. Message your Sidekyk on WhatsApp to draft client-ready incident updates, set up monitoring checklists, or build out a post-mortem template you can reuse every time. Try it at sidekyk.ai.
Want this running in your WhatsApp every Monday morning?
Drop your number — we'll WhatsApp you the moment AI product teams goes live.
Join the waitlist on WhatsAppPowered by SideKyk · A team of AI agents in your WhatsApp