Skip to content

Support, Service Levels, and Incident Communication

Draft 0.1 — 2 August 2026
Service objectives are operating targets. They are not contractual service-level agreements unless incorporated into an authorized agreement with defined remedies.

A product is not operationally trustworthy merely because it can be deployed. Users need to know:

  • where to ask for help;
  • what response to expect;
  • whether a problem is known;
  • what data or money may be affected;
  • who owns recovery;
  • what remedy or appeal exists.

Support must be:

  • reachable;
  • accessible;
  • appropriately private;
  • prioritized by impact;
  • linked to product evidence;
  • honest about uncertainty;
  • capable of escalation;
  • measured without gaming closure rates;
  • integrated with incidents and product learning.

Products declare supported channels such as:

  • in-product support;
  • email;
  • web form;
  • public community forum;
  • private organization channel;
  • phone or live support;
  • security disclosure channel;
  • accessibility contact;
  • payment dispute route;
  • status page;
  • emergency contact for partners.

Each channel declares:

  • audience;
  • hours;
  • response target;
  • privacy level;
  • supported languages;
  • accessibility profile;
  • authentication requirements;
  • prohibited content.

A request contains:

  • case ID;
  • requester and contact preference;
  • product;
  • category;
  • severity;
  • affected journey;
  • environment;
  • description;
  • timestamps;
  • attachments;
  • privacy classification;
  • payment, security, accessibility, or legal routing;
  • owner;
  • status;
  • linked incidents, feedback, refunds, or evidence.
  • account and access;
  • consent and data request;
  • product defect;
  • billing, refund, or payout;
  • opportunity or contribution;
  • credential or reputation;
  • accessibility barrier;
  • abuse or harassment;
  • privacy request;
  • security vulnerability;
  • service degradation;
  • legal or compliance inquiry;
  • product feedback;
  • other.

Sensitive categories bypass public queues.

No current service failure or meaningful harm.

Workaround exists; limited impact.

Important journey is degraded for one or more users.

Critical journey unavailable, material financial or data effect, significant accessibility blocker, or broad customer impact.

Active severe security, privacy, financial, safety, legal-rights, or multi-product incident.

Severity is based on impact and urgency, not customer status alone.

  • number and type of affected users;
  • criticality of journey;
  • duration;
  • data sensitivity;
  • financial exposure;
  • accessibility exclusion;
  • security exploitability;
  • reversibility;
  • legal or contractual obligation;
  • vulnerable users;
  • cross-product propagation;
  • public misinformation.
new → acknowledged → investigating → action_required → resolved → closed
↘ linked_to_incident
↘ awaiting_requester
↘ escalated
↘ appealed

Resolved means an outcome or answer exists. Closed means the case record is complete under policy.

Products publish targets by severity and support plan.

A target should distinguish:

  • acknowledgement;
  • initial qualified response;
  • update cadence;
  • resolution or workaround target;
  • escalation threshold.

Do not publish response times that staffing cannot support.

A measurement of service behavior that matters to users.

An internal or published target for an SLI.

A contractual commitment with parties, measurement, exclusions, and remedies.

Products must not label an SLO as an SLA merely because it appears on a pricing page.

Each SLO record includes:

  • service and owner;
  • user journey;
  • SLI specification;
  • SLI implementation;
  • data source;
  • target;
  • window;
  • exclusions;
  • error-budget calculation;
  • alerting policy;
  • review date;
  • approved tradeoff policy.

Google SRE guidance emphasizes that SLOs should support decisions and error-budget tradeoffs, not exist only as reporting metrics.

  • successful consent approvals / valid approval attempts;
  • correct credential status responses / status checks;
  • payout instructions processed without reconciliation error / valid instructions;
  • critical page loads under threshold / total critical page loads;
  • support cases acknowledged within target / eligible cases;
  • successful data exports / export requests;
  • accessible critical-journey completions in test suite / tested journeys.

For a reliability SLO, the error budget represents the tolerated amount of failure over a window.

The Error Budget Policy states what happens when consumption exceeds thresholds.

Possible actions:

  • slow feature rollout;
  • stop nonessential launches;
  • assign engineering time to reliability;
  • disable risky integrations;
  • reduce traffic or capability scope;
  • require executive or product-owner exception;
  • enter incident or recovery mode.

An exhausted error budget must have operational consequences.

Support SLIs may include:

  • acknowledgement time;
  • qualified-response time;
  • resolution time;
  • update-cadence compliance;
  • reopen rate;
  • escalation rate;
  • accessibility-barrier age;
  • refund completion time;
  • requester-confirmed resolution;
  • backlog age.

Avoid rewarding fast closure that causes repeated contact.

Products may offer tiers, but must preserve minimum access for:

  • security reporting;
  • privacy and data-rights requests;
  • accessibility barriers;
  • payment errors;
  • account compromise;
  • abuse reporting;
  • exit and export.

Basic rights cannot depend on a premium support plan.

A support case becomes linked to an incident when:

  • multiple users report the same failure;
  • monitoring confirms service degradation;
  • a security or privacy event exists;
  • financial reconciliation is affected;
  • a critical dependency fails;
  • incorrect public data or claims are broadly displayed.

Linked users receive incident updates according to their communication preferences and legal obligations.

  • Incident Commander;
  • technical lead;
  • communications lead;
  • support lead;
  • security or privacy lead;
  • legal/operator authority;
  • product owner;
  • provider liaison.

One person may hold several roles in a small team, but ownership remains explicit.

State observed symptoms and scope without inventing a cause.

State the confirmed cause at an appropriate disclosure level and current mitigation.

State that a fix or containment is active and what is being watched.

State service restoration, remaining impact, data or financial follow-up, and review plans.

Publish lessons, corrective actions, and evidence when appropriate.

An incident update should include:

  • affected product and capability;
  • start time and current status;
  • known user impact;
  • data or payment impact;
  • current actions;
  • workaround if safe;
  • next update time;
  • support route;
  • uncertainty.

Avoid false precision and premature blame.

  • status pages are keyboard and screen-reader usable;
  • updates do not rely on color alone;
  • critical content is available as text;
  • contact routes offer alternatives;
  • phone-only support is not the only route for deaf users;
  • text-only support is not the only route where it excludes users;
  • authentication for support is accessible;
  • outage updates are readable at zoom and narrow widths.

Payment cases require:

  • transaction reference;
  • merchant and processor roles;
  • ledger evidence;
  • refund status;
  • dispute status;
  • payout or reserve state;
  • fraud and sanctions routing;
  • user communication;
  • escalation and appeal.

Support agents must not promise funds they lack authority to release.

Security reports and privacy requests use restricted access.

Controls include:

  • secure submission;
  • identity verification appropriate to request;
  • no unnecessary sensitive-data collection;
  • specialist routing;
  • legal or policy timelines;
  • audit trail;
  • disclosure coordination;
  • protection against retaliation.

Cases may concern:

  • unclear acceptance criteria;
  • assignment conflicts;
  • payment delay;
  • evidence rejection;
  • reviewer conflict;
  • reputation correction;
  • harassment;
  • appeal.

The support route must link to the relevant dispute and appeal protocol rather than treating the issue as ordinary customer service.

Support knowledge should be:

  • versioned;
  • searchable;
  • linked to current product versions;
  • accessible;
  • reviewed after incidents;
  • explicit about jurisdiction or plan differences;
  • capable of correction;
  • separated from internal restricted runbooks.

AI-generated support answers require source grounding and escalation when uncertain.

Agents may:

  • retrieve approved knowledge;
  • classify cases;
  • suggest severity;
  • draft responses;
  • summarize history;
  • detect possible incidents;
  • propose troubleshooting steps.

Agents must not autonomously:

  • deny a high-impact appeal;
  • disclose restricted data;
  • issue uncapped refunds;
  • change payout destinations;
  • close security reports as invalid;
  • fabricate status information;
  • impersonate a human reviewer.

Material responses identify when AI assistance was used where appropriate.

Support records preserve:

  • original request;
  • actions;
  • communications;
  • approvals;
  • tool calls;
  • refunds or corrections;
  • incident links;
  • final resolution;
  • requester feedback;
  • retention and access rules.

Do not store passwords, full secrets, or unnecessary sensitive content in support tickets.

Depending on authority and terms, remedies may include:

  • explanation;
  • workaround;
  • correction;
  • restoration;
  • credential supersession;
  • refund or credit;
  • payout correction;
  • accessibility accommodation;
  • data deletion or export;
  • appeal;
  • service termination without penalty;
  • public correction.

Support staff must understand which remedies require approval.

Operational plans address:

  • coverage;
  • escalation;
  • training;
  • access control;
  • workload;
  • psychological safety for abuse cases;
  • handover;
  • after-hours emergencies;
  • language support;
  • provider coordination.

Unsustainable heroics are not a reliability strategy.

Reviews should be:

  • blameless about human error while accountable for decisions;
  • evidence-based;
  • focused on system and process improvement;
  • linked to actions with owners and dates;
  • visible at an appropriate confidentiality level;
  • checked for recurrence.

They do not erase legal, governance, or misconduct accountability.

The product reports:

  • case volume by category and severity;
  • response and resolution distributions;
  • backlog age;
  • reopen and escalation rates;
  • SLO compliance;
  • error-budget status;
  • incident count and duration;
  • accessibility and privacy case status;
  • financial corrections;
  • repeated failure themes;
  • action completion.

Metrics are segmented carefully to avoid exposing individuals or creating worker surveillance.

  • support email that nobody owns;
  • all customers marked urgent;
  • status page updated only after social media pressure;
  • 100% uptime target with no error-budget policy;
  • measuring average response only;
  • closing cases to improve metrics;
  • hiding payment or data impact;
  • requiring premium plans for security or privacy reports;
  • using AI to block escalation;
  • publishing a postmortem with no tracked actions;
  • relying on founder memory for incident history.
  • every product publishes support channels and expectations;
  • severity S0–S4 is implemented;
  • sensitive categories route privately;
  • service SLOs have owners and error-budget policy;
  • support SLOs distinguish acknowledgement and resolution;
  • incident and support records link;
  • status updates use defined stages;
  • critical communications are accessible;
  • payment, privacy, security, and contributor cases have specialist routes;
  • AI cannot make final high-impact support decisions;
  • remedies and approval authority are documented;
  • post-incident actions are tracked.
  • Sources opened and checked: 2 August 2026
  • SLO versus SLA distinction: Required
  • Error-budget operating policy: Required
  • Premium plan required for rights and safety intake: Prohibited