Skip to content

[FIX] queue_job: detect a dropped job runner connection - #1000

Open
HoloborodkoBohdan wants to merge 1 commit into
OCA:18.0from
HoloborodkoBohdan:18.0-fix-queue_job-jobrunner_keepalive
Open

HoloborodkoBohdan wants to merge 1 commit into
OCA:18.0from
HoloborodkoBohdan:18.0-fix-queue_job-jobrunner_keepalive

Conversation

@HoloborodkoBohdan

Copy link
Copy Markdown

The job runner holds its lock on a connection that stays open. When the network drops it without closing it (firewall, NAT timeout, database failover), neither side notices: the job runner blocks in its next query until the kernel gives up retransmitting (about 15 minutes on Linux), and the database keeps the session, and so the lock, until its own TCP keepalive fires (2 hours by default). No other job runner can take over meanwhile.

Enable TCP keepalives and a TCP user timeout on the connection, on the client side through the libpq parameters and on the server side through the session settings, configurable like the other jobrunner_db_* keys.

Fixes #984

The job runner holds its lock on a connection that stays open. When the
network drops it without closing it (firewall, NAT timeout, database
failover), neither side notices: the job runner blocks in its next query
until the kernel gives up retransmitting (about 15 minutes on Linux), and
the database keeps the session, and so the lock, until its own TCP
keepalive fires (2 hours by default). No other job runner can take over
meanwhile.

Enable TCP keepalives and a TCP user timeout on the connection, on the
client side through the libpq parameters and on the server side through
the session settings, configurable like the other jobrunner_db_* keys.

Fixes OCA#984
@OCA-git-bot

Copy link
Copy Markdown
Contributor

Hi @guewen, @sbidoul,
some modules you are maintaining are being modified, check this out!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[19.0] queue_job: Jobs stay pending after the jobrunner's database connection is silently dropped

2 participants