I am trying to configure borg backups on my system, and am running into a really weird failure mode. Here is what keeps happening to me:
The job runs for around 80 seconds
The message “systemd[1]: Stopping BorgBackup job home…” appears in the systemd logs
Nothing happens for 90s
systemd times out trying to sigterm, and does a sigkill instead
That also times out, and systemd gives up, marking the job as failed, claiming that “Unit process $PID (.borg-wrapped) remains running after unit stopped.”, but when I search for that PID that process is no longer running.
Viewing the logs it is actively backup up files right until it gets stopped, so it’s not like it is timing out or crashing, as far as I can tell.
It seems to work the first time, for the first full backup, then fails when doing future incremental backups. I find this extra weird because the full backup takes several minutes, but the incremental one reliably fails after ~80s.
The repo is a mounted Samba (cifs) network share, of the form /mnt/repo, which is defined using fileSystems."/mnt/repo" with fsType = "cifs".
So according to the isLocalPath helper in borgbackup.nix, yes it’s a local path. However, while that has the potential to cause a problem when running immediately after a boot, I am confident this is not causing this problem for two reasons:
I am (currently) running the backup manually, and it has more than enough time to mount before I start the service
The backup does not fail with a drive unreachable fault. If the repo was not mounted I expect it would fail immediately, not get partway through a backup then hang.
Just tried this, it didn’t add any more useful logs. I also tried adding it to the job environment just in case, and again nothing new. Still just borg output then suddenly “Stopping job”