What happened?
On 3007.13 two minions got stuck in an infinite requeue loop, logging ~180 lines/s of
[salt.minion :1960][WARNING ] Failed to acquire job_queue lock for jid 2026..., queuing anyway.
for a week, until /var/log/salt/minion reached 9 GB and filled /var. 59 stale jids cycled every 0.3 s. No job_queue.lock file existed, fds were 45/8192, plenty of RAM. The loop started 9 s after a Timeout encountered while sending ... _return request.
Root cause, two defects that combine:
-
salt.utils.files.await_lock hides the real exception. The yield is inside the outer try, whose except Exception re-raises everything as FileLockError("Error encountered obtaining file lock ..."). Any exception raised by the async with await_lock(...) body is reported as a lock failure and its traceback is lost. In 3007.13 the body included _invoke_execution() and the proc-file write, so a job-start failure became "Failed to acquire job_queue lock". 3007.14 moved _invoke_execution out of the lock, but await_lock itself is unchanged in 3008.x, and the minion's except FileLockError handler still logs without the exception.
-
The disk job queue has no exit. _process_process_queue_async_impl resubmits every queued_*.p file every 0.3 s (0.2 s in 3008.x) with no retry counter, max age or backoff. A job that fails deterministically at start is requeued forever.
Workaround used: stop minion, move cachedir/job_queue aside, truncate log, start.
Expected
await_lock should not convert exceptions from the with body into FileLockError (re-raise them unchanged, keep the wrapping for the lock-acquire path only).
_handle_decoded_payload's except FileLockError should log the exception (exc_info=True).
- Queued jobs need a retry limit and/or max age; after that the job is dropped with an error and a return to the master.
Type of salt install
onedir
Major version
3007 (3007.13); defect 1 and 2 still present in 3008.2 / master
What supported OS are you seeing the problem on?
Ubuntu
salt --versions-report output
Salt Version:
Salt: 3007.13
Python: /opt/saltstack/salt/bin/python3.10 (onedir)
OS: Ubuntu, cPanel hosts
What happened?
On 3007.13 two minions got stuck in an infinite requeue loop, logging ~180 lines/s of
for a week, until
/var/log/salt/minionreached 9 GB and filled/var. 59 stale jids cycled every 0.3 s. Nojob_queue.lockfile existed, fds were 45/8192, plenty of RAM. The loop started 9 s after aTimeout encountered while sending ... _return request.Root cause, two defects that combine:
salt.utils.files.await_lockhides the real exception. Theyieldis inside the outertry, whoseexcept Exceptionre-raises everything asFileLockError("Error encountered obtaining file lock ..."). Any exception raised by theasync with await_lock(...)body is reported as a lock failure and its traceback is lost. In 3007.13 the body included_invoke_execution()and the proc-file write, so a job-start failure became "Failed to acquire job_queue lock". 3007.14 moved_invoke_executionout of the lock, butawait_lockitself is unchanged in 3008.x, and the minion'sexcept FileLockErrorhandler still logs without the exception.The disk job queue has no exit.
_process_process_queue_async_implresubmits everyqueued_*.pfile every 0.3 s (0.2 s in 3008.x) with no retry counter, max age or backoff. A job that fails deterministically at start is requeued forever.Workaround used: stop minion, move
cachedir/job_queueaside, truncate log, start.Expected
await_lockshould not convert exceptions from thewithbody intoFileLockError(re-raise them unchanged, keep the wrapping for the lock-acquire path only)._handle_decoded_payload'sexcept FileLockErrorshould log the exception (exc_info=True).Type of salt install
onedir
Major version
3007 (3007.13); defect 1 and 2 still present in 3008.2 / master
What supported OS are you seeing the problem on?
Ubuntu
salt --versions-report output