Backup é seguro de vida dos dados. Ransomware, bugs, hardware failure. Recovery rápido = business salvo. Teste restores regularmente.
Conceitos Principais
Full Backup
Cópia completa do database. Baseline. Slow, large. Weekly ou monthly.
Incremental Backup
Apenas mudanças desde último backup. Fast, small. Daily ou hourly.
PITR
Point-in-Time Recovery. Restore para timestamp específico. WAL replay. PostgreSQL, MySQL.
RTO/RPO
RTO: Recovery Time Objective (quanto tempo). RPO: Recovery Point Objective (quanto data loss).
Passo a Passo
- Full Backup Postgres: pg_dump mydb > backup.sql. Ou pg_basebackup para binary. Compress: gzip backup.sql. Upload S3.
- WAL Archiving: postgresql.conf: wal_level = replica, archive_mode = on, archive_command = 'cp %p /archive/%f'. Archive para S3.
- Restore Full: createdb mydb. psql mydb < backup.sql. Ou pg_restore para binary format.
- PITR Restore: Restore base backup. Create recovery.conf: restore_command, recovery_target_time. Start Postgres, replays WAL até target.
- Automate Backups: Cron job: daily full, hourly WAL archive. Retention: 30 days. Test restore monthly. Alert on failure.
Boas Praticas
Recomendacoes
• Backup diário automático
• WAL archiving para PITR
• Off-site backup (S3, GCS)
• Encryption em trânsito e at rest
• Test restore mensalmente
• Document recovery procedures
• Monitor backup success
Erros Comuns
Evite estes erros
• Não testar restore (backup inútil)
• Single backup location
• Sem encryption
• Retention policy indefinida
• Manual backups
• Não monitorar backup jobs
Checklist
- Full backup automático
- WAL archiving ativo
- Backup em S3/cloud
- PITR testado
- Restore procedure documentada
- Monthly restore test