When Microsoft Teams stops working, the first thirty minutes determine whether an organisation contains the disruption or creates a second incident through confusion. The right response is not to move every conversation to the first app that loads. It is to verify the failure, protect high-risk work, activate an approved fallback and keep one record of what changed.
This plan is designed for Teams chat, meetings and file-dependent workflows. It also works when the visible symptom is part of a wider Microsoft 365 or Azure incident. The July 2026 Microsoft outage showed why this distinction matters: users may see a collaboration failure while the confirmed root sits in regional network routing or a shared API dependency.
The goal is controlled continuity. Some work should move, some should pause, and some should wait for evidence.
The 30-minute command table
| Time | Incident lead | IT or service owner | Business owners | All users |
|---|---|---|---|---|
| 0–5 minutes | Open incident record and name one coordinator | Test scope; check official and tenant status | Report critical blocked processes | Stop repeated retries and record exact symptoms |
| 5–10 minutes | Classify severity and approve fallback level | Identify dependency and unknown-state actions | Identify meetings, approvals and deadlines at risk | Use only the named fallback |
| 10–20 minutes | Publish one situation update | Protect integrations; preserve logs | Move only priority activity | Avoid sending sensitive files to personal tools |
| 20–30 minutes | Confirm owners and next update time | Test recovery path and duplicate-action risk | Reconcile what moved or paused | Acknowledge the update; do not create parallel rooms |
| After 30 minutes | Continue timed updates | Verify functions, not only service status | Validate business outcomes | Return to Teams only when instructed |
Print or store this table outside Teams. A continuity plan that can only be opened inside the failed service is not a continuity plan.
Minute 0–5: establish facts
Create a record in the organisation’s incident system, an approved offline document or a pre-arranged emergency channel. Include:
first observed time in UTC and local time;
affected geography and user groups;
failing function: sign-in, chat, meeting, files or automation;
exact error or behaviour;
clients and networks tested;
business process currently at risk;
person coordinating the response.
Use a neutral title such as “Teams meeting join failures — APAC — investigation”. Avoid declaring a global outage before scope is known.
IT should test one known-good account on web, desktop and mobile, using a second network where practical. The objective is not exhaustive troubleshooting. It is to distinguish local endpoint failure from tenant, regional or provider impact.
Check Microsoft’s public status sources and the tenant service-health dashboard. The Azure status history is useful when a Microsoft 365 symptom may involve a cloud dependency. Capture the tracking identifier and last update time rather than relying on a screenshot that will become stale.
Minute 5–10: classify the business impact
Use three levels:
| Level | Definition | Default action |
|---|---|---|
| Level 1: inconvenience | Some users or non-critical functions fail; work has safe alternatives | Continue with local workarounds and monitor |
| Level 2: operational disruption | Teams, files or meetings block time-sensitive business activity | Activate approved fallback for named processes |
| Level 3: material incident | Safety, regulated communication, executive decision, payment or customer commitment is at risk | Invoke formal incident management and senior ownership |
Severity is based on business consequence, not the number of social-media reports. One failed executive approval before a market deadline may matter more than thousands of delayed emoji reactions.
At this stage, identify actions in an unknown state. A user may have clicked “approve”, sent a message or triggered a workflow without receiving confirmation. Do not retry until the source and destination have been checked. Duplicate payment or instruction is often more damaging than delay.
Choose a fallback by information sensitivity
Not every fallback is suitable for every conversation.
| Work type | Acceptable fallback example | Do not use |
|---|---|---|
| Public or low-risk coordination | Approved secondary meeting service, corporate email, telephone bridge | Personal social account |
| Internal confidential discussion | Pre-approved encrypted corporate channel with managed identity | Consumer group chat with unmanaged members |
| Restricted documents | Approved document repository or wait for restoration | Personal email, public link or copied USB media |
| Executive approval | Formal approval system, recorded telephone process with dual control | Informal message without audit trail |
| Customer meeting | Approved alternate link sent through verified contact path | Last-minute link from an unknown account |
| Emergency or safety communication | Established emergency notification system and telephone tree | A collaboration app with uncertain delivery |
The fallback should already exist, have managed identities and be included in retention policy. Creating a new consumer workspace during an incident introduces unreviewed permissions, weakens evidence and may expose contact details.
Minute 10–20: send one operational update
The incident lead should publish a short message through the fallback:
Teams disruption is affecting [functions/users] from [time]. Use [approved fallback] only for [named activity]. Pause [ambiguous approvals or automations]. Do not move restricted files to personal services. Next update at [time]. Incident owner: [role].
The update should be specific enough to prevent improvisation. “Use email” is incomplete if file attachments are restricted. “Use the emergency bridge for client calls; keep documents in the approved repository” is actionable.
Business owners then decide what must move. Cancel or postpone low-value meetings. Move high-value calls. Preserve decisions in a formal record rather than allowing three parallel chat groups to produce conflicting outcomes.
IT should freeze or monitor integrations that may replay. Webhooks, meeting bots, approval connectors and AI assistants can fail in partial ways. An assistant may still receive a trigger but lose access to the file it needs, producing an incomplete answer. Mark automated output as untrusted until the dependency path is healthy.
Minute 20–30: create the recovery queue
Continuity is not complete when people can talk again. Create a queue of items that must be reconciled:
messages sent without delivery confirmation;
files edited locally or in a fallback repository;
meetings moved to another platform;
approvals attempted during the incident;
bots or workflows paused;
customer commitments made through the fallback;
users added to temporary rooms;
access granted for emergency work.
Each item needs an owner and a completion test. “Check later” is not a control.
Set the next update time even if there is no change. A known 30-minute cadence reduces duplicate questions and prevents users from treating silence as recovery.
Keep the executive channel small
A major outage attracts observers. The incident bridge should contain people with decisions to make: incident lead, service owner, security or privacy representative, business owners for affected processes and communications where customer impact exists.
Separate the working log from the executive summary. The working log records tests, timestamps and hypotheses. The summary states impact, action, risk and next decision. Mixing them produces a noisy channel where important instructions disappear.
An executive update can follow four lines:
Impact: what material business activity is blocked?
Cause: confirmed, suspected or unknown?
Control: what fallback and safety boundary is active?
Decision: what requires leadership approval before the next update?
Do not convert a preliminary provider statement into a final root cause. Say “Microsoft is investigating” or “Microsoft’s preliminary PIR attributes…” with a timestamp.
Special case: the meeting begins in five minutes
If a customer or board meeting is imminent, the meeting owner should decide among three options:
move to the pre-approved secondary platform;
use a telephone bridge;
postpone.
Verify the alternate invitation through a known contact channel to reduce phishing risk. Do not send a new link from a newly created personal account. If confidential documents cannot be shared safely, conduct the verbal portion and reschedule document review.
Assign a recorder. Decisions made outside the normal platform must return to the system of record after recovery.
Special case: files are unavailable
Teams file tabs often depend on SharePoint or OneDrive. If chat works but files do not, copying documents into the conversation may worsen the problem.
Check whether the source repository is directly available. Use offline copies only when policy permits and label them clearly. Avoid simultaneous editing across fallback locations. Nominate one version owner and record changes.
When service returns, compare hashes, versions or tracked changes before uploading. Do not assume the newest timestamp contains every authorised edit.
Special case: an approval may have completed
Unknown-state actions need a two-system check. Verify the source request and the destination effect. For a purchase approval, check both the approval record and procurement system. For a message, check the sender’s outbox and recipient’s receipt. For an automation, inspect the run log and external transaction identifier.
If the action has a financial, legal or customer consequence, require a second person to approve any retry. This is slower than clicking again and faster than unwinding a duplicate.
Recovery verification
Microsoft or the tenant administrator may mark the service healthy before every user session has recovered. Test the actual workflow:
| Function | Recovery test | Evidence to retain |
|---|---|---|
| Identity | Sign in with a normal managed account | Time, account class and result |
| Chat | Send one harmless message and confirm receipt | Sender and receiver timestamps |
| Meetings | Join from two networks and test audio/video | Meeting ID and result |
| Files | Open, edit and synchronise a test file | Version history |
| Integration | Run one non-production workflow | Exactly one source and destination record |
| External access | Test a controlled guest scenario | Guest identity and permission result |
Restore services in order of risk. Communication can return before high-impact automation. Remove temporary permissions and close fallback rooms when no longer needed.
The post-incident review
Within a few business days, answer:
Did users know where to find the plan?
Was the fallback available outside Teams?
Did identity or file dependencies undermine the fallback?
Were any sensitive files moved to unapproved locations?
Did any action duplicate or disappear?
How long did it take to send the first clear update?
Which status source was most reliable?
What control can be tested next quarter?
Do not judge the response only by uptime. Measure decision clarity, data exposure, duplicate actions and recovery completeness.
Microsoft provides broader Microsoft 365 business-continuity planning guidance. Use it to connect this rapid runbook to formal resilience design.
The verdict
When Teams goes down, spend the first five minutes establishing facts, the next five classifying impact, the following ten activating a controlled fallback and the final ten building the recovery queue. Move only the work that must move.
The strongest continuity plan protects identity, confidentiality and action lineage while communication is disrupted. It gives every user one instruction, every ambiguous action one owner and every fallback a clear end.
A collaboration outage should remain a service incident. Without discipline, it becomes a data, approval and trust incident as well.




