Skip to content

fix: never make a page wait on WebDecoy Cloud when ingest is slow or down - #104

Merged
cport1 merged 1 commit into
mainfrom
fix/ingest-outage-page-stall
Oct 2, 2026
Merged

cport1 merged 1 commit into
mainfrom
fix/ingest-outage-page-stall

Conversation

@cport1

@cport1 cport1 commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

For WebDecoy/app#1245, item 2.

Problem: for any visitor scoring 40+, the plugin sent the detection inline during init, before the block decision. That send had a 10s timeout, and installs without a stored org id did a 10s validate-key lookup first. During an ingest outage, flagged page views stalled up to ~20s.

Fix

  • Detections are queued and sent at shutdown, after fastcgi_finish_request(). Blocking and local logging never wait on the cloud.
  • The detection client uses a 3s timeout and the stored org id, so there's no per-request key lookup. When no org id is stored, the looked-up one is cached for a day.
  • Any refusal (429, 5xx, or no answer) pauses cloud calls for 60s across all requests via a transient.
  • IP enrichment (needed in the page path by ip.* filter rules) skips the call while paused and starts the pause on an unavailable answer. That removes the 2s-per-new-IP cost during an outage.

Tests: php tests/run.php gives 189 passed (7 new).

  • Includes a source-reading guard: the only ->submitDetection( must sit inside WebDecoy_Detection_Sender::defer.
  • Mutation-checked: an inline send fails the guard, and dropping the enrichment skip fails its test.

Ships in the next plugin release.

…down

Detections were sent inline during init, before the block decision, with a
10 second timeout, and installs without a stored organization id made a
second 10 second key lookup first. When ingest was slow or down, every
flagged page view stalled for up to 20 seconds (WebDecoy/app#1245).

- Detections are queued during the request and sent from a shutdown
  handler after fastcgi_finish_request(), with a 3 second timeout and the
  stored organization id, so one send is one request.
- A refusal (429, 5xx, no answer) pauses cloud calls for 60 seconds across
  all requests, so an outage costs one attempt per minute, not one per page.
- IP enrichment, which filter rules need in the page path, skips the call
  while paused and starts the pause on an unavailable answer.
- Tests pin the deferral, the backoff and, by reading the plugin source,
  that the only detection send is inside the deferred sender.
@cport1
cport1 merged commit c5d86c1 into main Oct 2, 2026
4 checks passed
@cport1
cport1 deleted the fix/ingest-outage-page-stall branch October 2, 2026 16:58
@cport1 cport1 mentioned this pull request Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant