Borg backups keep failing for no apparent reason

I am trying to configure borg backups on my system, and am running into a really weird failure mode. Here is what keeps happening to me:

  1. The job runs for around 80 seconds
  2. The message “systemd[1]: Stopping BorgBackup job home…” appears in the systemd logs
  3. Nothing happens for 90s
  4. systemd times out trying to sigterm, and does a sigkill instead
  5. That also times out, and systemd gives up, marking the job as failed, claiming that “Unit process $PID (.borg-wrapped) remains running after unit stopped.”, but when I search for that PID that process is no longer running.

Viewing the logs it is actively backup up files right until it gets stopped, so it’s not like it is timing out or crashing, as far as I can tell.

Here is my (lightly redacted) config:

services.borgbackup.jobs.home = {
  repo = "/path/to/backups/home";
  encryption = {
    mode = "repokey";
    passCommand = "cat /etc/nixos/borg-pass";
  };
  compression = "lz4";
  paths = "/home";
  exclude = [
    "**/baloo/index"
    "**/.cache"
    "**/steamapps/common"
  ];
  extraCreateArgs = [
    "--stats"
    "--progress"
    "--verbose"
    "--checkpoint-interval" "900"
  ];
  failOnWarnings = false;
  startAt = "daily";
  prune.keep = {
    within = "1d";
    daily = 7;
    weekly = 4;
    monthly = 6;
  };
};

Any idea what could be causing this?

You cab try adding more verbosity flags to borg if possible.

Add this and see if there are more logs

systemd.services.borgbackup.environment.SYSTEMD_LOG_LEVEL = "debug";

Has it worked before with this configuration?

Is the repo a local path like in the obfucated configuration?

It seems to work the first time, for the first full backup, then fails when doing future incremental backups. I find this extra weird because the full backup takes several minutes, but the incremental one reliably fails after ~80s.

The repo is a mounted Samba (cifs) network share, of the form /mnt/repo, which is defined using fileSystems."/mnt/repo" with fsType = "cifs".

So according to the isLocalPath helper in borgbackup.nix, yes it’s a local path. However, while that has the potential to cause a problem when running immediately after a boot, I am confident this is not causing this problem for two reasons:

  1. I am (currently) running the backup manually, and it has more than enough time to mount before I start the service
  2. The backup does not fail with a drive unreachable fault. If the repo was not mounted I expect it would fail immediately, not get partway through a backup then hang.

Just tried this, it didn’t add any more useful logs. I also tried adding it to the job environment just in case, and again nothing new. Still just borg output then suddenly “Stopping job”

I seem to have been mistaken. The last three times I checked just now it ran for almost exactly 80s. Not sure if that is important.

EDIT: I have updated my earlier posts to reflect this.

EDIT 2: And of course now that I have done so it has run for over two minutes then exactly two minutes, so 80s is also not a hard limit.