So one librarian can serve several bots (test + deploy) instead of firing
every result/pong at a single static CONJURER_MAIN_BOT.
* The bot includes its own callback address (CONJURER_SELF_CALLBACK) in
every /query and /ping.
* The librarian stores that callback with the query (persisted with the
request, so a replay after restart still answers the right bot) and, for
results, in the OUTBOX entry ({target, payload}) so the resender delivers
to the origin bot even across a librarian restart.
* Pongs go back to the pinging bot too - otherwise a second bot's health
check would be ponged to the first and always time out, so it could
never enable its librarian cog.
* Empty callback falls back to MAIN_BOT_ADDRESS, and a legacy OUTBOX entry
(raw payload, pre-callback) is still delivered to the default bot, so the
upgrade is seamless.
Tests: per-origin result delivery + legacy-shape fallback (outbox),
busy/idle pong routed to the callback bot vs default (lifecycle). Suite:
58 unit + 52 integration green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two refinements to the librarian health/delivery story, matching how it
actually behaves under load:
1. Busy-aware ping (case b - broken return path). A ping arriving while
the worker is grinding a search no longer queues behind it (which made
a healthy-but-busy librarian time out and look dead). The librarian
tracks worker_busy and, when set, pongs back IMMEDIATELY without
touching the queue. Being busy is fine - you can keep piling searches
on. The ping still travels the librarian->bot return path, so it keeps
catching the one thing it must: a disrupted/incompatible return path
where queries vanish. Idle pings still go through the internal queue.
2. Per-query watchdog (case a - finished but result lost). The librarian
now tracks every search uuid's lifecycle (queued -> processing ->
gone) in active_queries, exposed via a new POST /query_status. After
dispatching a search the bot records it in self.pending; watch_pending
polls /query_status for each. While the librarian still knows the uuid
the search is progressing - left alone. The moment a uuid VANISHES
there while still pending on the bot, its result was computed but never
delivered: after a grace window (to rule out an in-flight result) the
bot posts a notice to the channel - but ONLY then. A normally delivered
result is popped from self.pending by check_data_q and never flagged.
Hardening: the worker's search body is now wrapped in try/except/finally
so a crashing search can't kill the worker thread (which would freeze the
queue), and worker_busy / active_queries are always cleared. The grace
logic lives in a dependency-free librarian_watchdog.pending_verdict so it
is unit-testable without discord/pdf libs. /ping and /query_status are
plain (sync) views so they run without flask[async].
Tests: unit test_librarian_watchdog (verdict transitions); integration
test_librarian_query_lifecycle (query_status known/unknown + auth,
idle-ping-queues, busy-ping-pongs-directly). Suite: 28 integration + 48
unit green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>