Two refinements to the librarian health/delivery story, matching how it
actually behaves under load:
1. Busy-aware ping (case b - broken return path). A ping arriving while
the worker is grinding a search no longer queues behind it (which made
a healthy-but-busy librarian time out and look dead). The librarian
tracks worker_busy and, when set, pongs back IMMEDIATELY without
touching the queue. Being busy is fine - you can keep piling searches
on. The ping still travels the librarian->bot return path, so it keeps
catching the one thing it must: a disrupted/incompatible return path
where queries vanish. Idle pings still go through the internal queue.
2. Per-query watchdog (case a - finished but result lost). The librarian
now tracks every search uuid's lifecycle (queued -> processing ->
gone) in active_queries, exposed via a new POST /query_status. After
dispatching a search the bot records it in self.pending; watch_pending
polls /query_status for each. While the librarian still knows the uuid
the search is progressing - left alone. The moment a uuid VANISHES
there while still pending on the bot, its result was computed but never
delivered: after a grace window (to rule out an in-flight result) the
bot posts a notice to the channel - but ONLY then. A normally delivered
result is popped from self.pending by check_data_q and never flagged.
Hardening: the worker's search body is now wrapped in try/except/finally
so a crashing search can't kill the worker thread (which would freeze the
queue), and worker_busy / active_queries are always cleared. The grace
logic lives in a dependency-free librarian_watchdog.pending_verdict so it
is unit-testable without discord/pdf libs. /ping and /query_status are
plain (sync) views so they run without flask[async].
Tests: unit test_librarian_watchdog (verdict transitions); integration
test_librarian_query_lifecycle (query_status known/unknown + auth,
idle-ping-queues, busy-ping-pongs-directly). Suite: 28 integration + 48
unit green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The librarian health check was a plain GET to '/', which only proved
Flask was listening - not that the service could actually take a query,
run it through its internal queue+worker, and answer back. So the cog
could load against a librarian whose worker was wedged or that couldn't
reach the bot on the return leg.
Replace it with a ping that travels the SAME path a real search does, on
both sides:
bot: QueryControl -> OUT_COMM_Q -> scan_queue -> awaiting_q
librarian: POST /ping -> librarian_queue -> worker pulls it off
(no Crossref/DOI search) -> pongs back with the same uuid
bot: /conjurer -> incoming_q -> scan_incoming matches uuid, wakes waiter
The cog enables only when that whole loop closes within 3s. This also
proves the librarian->bot return path, which a GET never did.
Safety: uuid is random per ping; the wait and POST are both bounded so
startup can't stall; a pong that finds no waiter is dropped (never
orphaned into IN_COMM_Q, which would make the cog post a bogus 'no
results' message); and a ping whose pong never returns is swept out of
awaiting_q after PING_TTL_SECONDS so nothing leaks. All awaiting_q writes
stay within scan_queue (append) and scan_incoming (remove) - no locks,
no cross-thread mutation.
Integration tests cover: OK round-trip, timeout when accepted-but-no-pong,
unreachable, non-200, orphan-pong-dropped, and that real results still
reach IN_COMM_Q. Suite: 24 integration + 41 unit green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>