Hi.
Since a few weeks, we are repeatedly experiencing an error preventing our self-hosted runners from running Bitbucket Cloud pipelines.
These runners may remain in an unhealthy state after executing a pipeline, leading to systematic failure of all subsequent pipeline ran on the faulty runner. The issue can only be resolved by restarting the VM or manually removing the offending file.
Here is the output of the pipeline:
Runner matching labels:
- linux
- fast
- self.hosted
Runner name: bitbucket-runner-fast-1
Runner UUID: {056b36f2-6db0-5784-b31e-543bb76093ca}
Runner labels: self.hosted, linux, fast
Runner version:
current: 3.1.0
latest: 3.1.0
mkfifo: /var/lib/bitbucket-pipelines-runner/056b36f2-6db0-5784-b31e-543bb76093ca/tmp/clone_result: File exists
Skipping cache upload for failed step
Searching for test report files in directories named [test-reports, TestResults, test-results, surefire-reports, failsafe-reports] down to a depth of 4
Finished scanning for test reports. Found 0 test report files.
Merged test suites, total number tests is 0, with 0 failures and 0 errors.I've been trying to figure out how to reproduce the problem, but it's still not very clear. I have feeling it happens either:
- When a running pipeline is stopped manually from the UI.
- When a pipeline is failed prematurely due to "The clone failed due to a merge conflict with the destination branch. Fix conflicts and then commit the result."
Somehow, the "clone_result" file created by mkfifo isn't cleaned up, which causes subsequent runs to fail.
We checked that our configuration and VMs were correct. These have always worked well until then. We can't figure out what's causing the problem, and we're beginning to think it might be a Bitbucket problem.
Any ideas please?