Skip to main content

Disaster recovery

This page describes how Percus recovers its platform from failures and data loss: which scenarios are covered, how each one is resolved, how long recovery takes, and what has actually been tested. It states the current state as of October 2026, including what is not covered yet.

πŸ“„ Download this plan as a PDF (Spanish)

Percus also maintains an internal operational runbook with the step-by-step procedure for each scenario below. It is not published because it contains infrastructure details.

Scope​

  • Viewer service: the video configuration and player delivery that a recipient's browser calls to load and play a personalised video.
  • Backoffice: the web application and APIs that Percus clients use to manage projects, templates, distribution channels and users.
  • Data: the relational database (organisations, users, projects, templates, channels), per-viewer personalisation data, template assets, and analytics.

Recovery objectives​

Contractual commitmentInternal targetMeasured
RTO β€” time to restore service4 hours30 minutes for a database restoreDatabase failover: ~15–17 s. Database restore from a snapshot: ~12 min until the database was ready (see Testing)
RPO β€” maximum data loss30 days (backup retention window)~5 minutes (point-in-time recovery granularity)β€”

The contractual values are the ones published in Service levels. The internal targets are what Percus is working towards; they are goals, not commitments. The 30-minute target is what a single recovery event must fit within for the platform to sustain 99.9% monthly availability.

How the platform is built for recovery​

  • One AWS region, two Availability Zones. Production runs in AWS us-east-1. Compute is serverless and spread across Availability Zones by AWS; the network has one NAT gateway per Availability Zone.
  • Relational database with a standby. A writer and a standby reader in two Availability Zones, with automatic failover. Encrypted at rest with a customer-managed key.
  • Continuous backups. The relational database keeps point-in-time recovery for 30 days. Per-viewer personalisation data keeps point-in-time recovery for 35 days.
  • Versioned file storage. Template and asset storage keeps previous versions of every file, so a deleted or overwritten file can be recovered.
  • Infrastructure as code. Infrastructure is defined in code and deployed through automated pipelines, so components can be rebuilt from source control. A few settings (DNS zones, and the backoffice hosting's repository connection and secrets) are configured manually.
  • Protected credentials, keys and data. Database credentials, the encryption keys and the secrets that stored data depends on are retained even if the infrastructure that created them is removed. The relational database has deletion protection enabled; the per-viewer data table and its key are also retained if that infrastructure is removed.
Backup history

The current database cluster replaced the previous one on 6 October 2026 (UTC). Its point-in-time history starts on that date and reaches the full 30-day window on 5 November 2026.

Scenarios​

ScenarioRecoveryExpected timeStatus
A database instance or an Availability Zone failsAutomatic failover to the standby in the other Availability Zone. No data is lostSecondsTested (October 2026)
Relational data is deleted or damaged by mistake (an operation, a faulty release, corruption)Point-in-time restore to a new database just before the problem, the platform is switched to it, and the result is verifiedTarget 30 minRestore measured; switch-over procedure being validated
The whole database cluster is lostRestore from the latest point in time into a new cluster, then switch the platform to itTarget 30 minRestore measured; switch-over procedure being validated
Per-viewer personalisation data is lost or damagedPoint-in-time restore of that data to a new table, or the client re-uploads its data fileTo be defined in the drillDocumented, drill pending
A template or asset file is deleted or overwrittenRestore the previous version from versioned storageMinutesDocumented
The player or embed SDK delivery is lost or brokenRedeploy from source controlMinutesDocumented
A release breaks the platformRedeploy the previous version. Redeploying does not revert database schema changes; if a release damaged data, the relational-data scenario appliesUnder an hourDocumented; automatic rollback planned
The whole AWS region becomes unavailableNot covered by an automatic or tested procedure today: the platform and its backups are in a single regionNot definedKnown limitation

Analytics data is recoverable from its raw event store when the compacted data is lost; raw events not yet compacted may be lost. Raw events are kept for 12 months. Analytics does not affect video delivery.

Incident response​

Detection. AWS CloudWatch alarms on background processing queues and email delivery alert the engineering team by email. Alerting on API availability and errors, and external synthetic monitoring that checks the platform from outside AWS the way a client would, are being implemented.

Roles.

RoleResponsibility
Incident lead (CTO)Declares the incident, decides the recovery path, and approves switching the platform to restored data
Technical operatorRuns the recovery procedure from the internal runbook and verifies the result
Client communication (Customer Success)Keeps affected clients informed until the incident is closed

Communication. Affected clients are informed through their usual Percus contact. For incidents that affect personal data, the notification commitments in the Data Protection Addendum apply, including notification within 24 hours.

After the incident. Every incident that requires a recovery procedure is reviewed, and the runbook and this page are updated with what was learned.

Testing​

Recovery is tested with drills on the staging environment, which has the same topology as production (a writer and a standby in two Availability Zones).

DateTestResult
5–6 October 2026Database restore from a snapshot into a new cluster, during the planned move to the encrypted database (staging and production)Database ready in ~13 min (staging) and ~12 min (production)
6 October 2026Forced database failover to the other Availability Zone, and back (staging)Promotion in ~17 s and ~15 s; the API returned degraded responses for ~7–9 s
PlannedFull recovery drill: point-in-time restore, switching the platform over and verification, timed end to endβ€”

A full recovery drill will run at least once a year and after any major infrastructure change. Each result is added to the table above.

Known limitations​

  • Single region. There is no tested recovery from the loss of the entire AWS region.
  • Full recovery drill pending. The restore has been measured, but the complete procedure (restore, switch-over and verification) has not yet been timed end to end.
  • No formal on-call rotation yet. A formal on-call rotation is being put in place.
  • Detection is partial. Alarms cover background processing and email delivery; alerting on API availability and errors, and external monitoring, are in progress.