# The Backup That Silently Died

I had backups configured for my homelab.

The jobs existed. The destination was configured. Copies had been landing in Cloudflare R2. I had already done the work people usually mean when they say, “we have backups.”

Then I checked them after a host migration.

Nothing had run for roughly four weeks.

No failed-job alert. No obvious error. No warning telling me the protection I thought I had no longer existed.

The backups had simply stopped.

## What happened

My homelab runs on Proxmox. Its scheduled guest backups depend on `pvescheduler`.

After a reboot during the migration, `pvescheduler` ended up dead and disabled on one of the hosts.

The backup definitions were still there. The storage configuration was still there. The dashboard still looked like a system where backups had been configured.

But the thing responsible for starting the jobs was not running.

No scheduler meant no backup attempts. No attempts meant no failed jobs. No failed jobs meant nothing triggered an alert.

That is a particularly ugly failure mode: the component that would produce the evidence of failure is itself the component that failed.

## Configuration is not evidence

I re-enabled the service:

```Bash
systemctl enable --now pvescheduler
```
But a green service status was not enough.

The real verification was a new backup completing and appearing in R2. Until the artifact existed at the destination, the backup system was not fixed.

This sounds obvious when written down. In practice, infrastructure gets verified one layer too early all the time.

The service is running, so the job must work.

The job says success, so the upload must exist.

The file exists, so it must be restorable.

Each statement assumes the next boundary without proving it.

## Successful silence is dangerous

Most monitoring is designed around explicit failure.

A job runs and exits non-zero. An API returns an error. A disk fills up. A process crashes.

But some of the worst infrastructure failures are missing events:

- A backup did not run.
- A report did not arrive.
- A certificate did not renew.
- A scheduled sync stopped producing output.
- A heartbeat disappeared.
There may be no error to collect because the system responsible for producing the error never woke up.

Monitoring the scheduler would have helped. Monitoring failed backup jobs would have helped in other cases.

The control that actually covers this failure is simpler: alert when the last successful backup becomes too old.

Do not ask only whether the backup process is healthy.

Ask when the last valid artifact reached the destination.

## A backup is a claim

Saying “we have backups” is really several claims:

1. The scheduler is running.
2. The job is executing.
3. The source data is being captured.
4. The artifact reaches independent storage.
5. Retention keeps enough history.
6. The artifact can be restored.
If you verify only the first one, you do not have a backup system. You have a running service.

If you verify only that a file exists, you have a stored object. You do not yet know whether it can recover anything.

My immediate failure was the scheduler, but the lesson is broader.

Safety systems need proof at the outcome, not confidence at the configuration.

The backup that fails loudly is annoying.

The backup that never runs is worse, because it lets you keep believing you are protected.