test_search_fills_progress_with_live_positions_and_total asserted 100%
coverage after searching for a DOI that EXISTS. Once every queried DOI is
found the consumer signals TERM and the producers stop mid-file, so the
recorded offsets reach an arbitrary point - the assertion was racing the
scan and failed roughly one full-suite run in two.
Split into the two things that are actually deterministic: coverage is now
measured with an ABSENT DOI (nothing can stop the scan early, so 100% is
guaranteed), and the found-target case asserts what holds regardless of
where the producers stopped - the total is known, progress is bounded and
sane, and the hit is reported.
Verified: 5 consecutive runs of the file and 3 consecutive full integration
runs, all green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
FILE_SHARING.md explains how the feature works and assumes you already know
the node story; there was nothing describing what a fresh box has to provide.
SHARE_NODE_SETUP.md fills that: the four host directories, the uid-33
readability requirement on the media library, deploying the musician+share
stack (Portainer or CLI), forwarding /share from the reverse proxy, and the
musician-side CONJURER_SHARE_* wiring - each with its verification command.
Calls out the three couplings that actually break it in practice: the media
must be mounted at the SAME container path on both containers (index paths and
symlinks are absolute), the link/index dirs must be shared, and
CONJURER_SHARE_BASE_URL must match the real public URL - it is not set in the
bundled stack file and silently falls back to czernobog.pl.
Also records what the node does NOT need: no host Apache, no cron, no Python,
no hand-placed scripts - Apache, the vhost, scan_shares.py and
revoke_shares.py are all baked into the image.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Field report: the radio died and kept restarting. The log chain is
unambiguous - "Daemon startup failed" (pulse), then Connection refused on
input.pulseaudio_0 / buffer.consumer_0 / pulse_out, then "Shutdown
started!", then round again.
Three fixes:
* Clear stale pulse runtime state before starting the daemon. /run is part
of the container's writable layer, so "docker restart" - and the loop that
restart:unless-stopped produces - preserves /run/pulse/pid from the killed
daemon; the next start then refuses with "Daemon startup failed", which is
what makes the loop self-sustaining. We only remove it when no pulseaudio
process is actually alive.
* When pulse still won't start, say so loudly and explain the consequence
and the way out (Icecast needs no sound device; comment the pulse tor out),
instead of a one-line WARNING that gets lost above the traceback.
* output.pulseaudio(fallible=true): the local monitor output can now fail
without failing its clock and tearing down the radio. The mic input stays
a hard dependency - documented inline, since removing it changes audio
behaviour and cannot be verified without a liquidsoap runtime.
Also: RADIO_FORCE_SCRIPT=1 re-seeds radio_conjurer.liq from the image
(keeping a .bak). The live script is deliberately never overwritten so hand
edits win - but that also means image fixes never reached volumes seeded
long ago, which is why a running radio can still hold a stale script.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Field report: one httpx ReadTimeout inside habanero surfaced as
'Search <uuid> crashed', and the worker's crash handler then FORGOT the
request - so an expensive search vanished and the user was told it was
eaten, because a public API blinked once.
Two defences:
* Every habanero call goes through _crossref_call, which retries with
linear backoff (CONJURER_CROSSREF_ATTEMPTS, default 4; backoff
CONJURER_CROSSREF_BACKOFF, 5s). habanero wraps httpx errors in a plain
RuntimeError so we can't filter narrowly - retries are simply bounded
and the last error is re-raised. They now also run via asyncio.to_thread,
so a slow Crossref no longer blocks the worker's event loop.
* A crashed search is no longer dropped on the first failure: the attempt
count is persisted with the request and the search is requeued (keeping
any checkpoint, so a crashed DB scan resumes rather than restarts) until
CONJURER_SEARCH_MAX_ATTEMPTS (default 3). It stays 'queued' for the
bot's watchdog while retrying, and only after the cap is it forgotten.
Also: scrape_bot's 'Got blocked' is routine sci-hub behaviour (it backs off
an hour and carries on) - log it as WARNING, not ERROR, so it stops looking
like a fault when scanning for real problems.
Tests: retry-then-succeed, bounded re-raise, no retry on success, the
persisted attempt counter, and forget-on-give-up. Suite: 58 unit + 70
integration green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A deep scan runs for hours with nothing in the log between start and
finish, so it's impossible to tell a working search from a wedged one.
Every CONJURER_LIBRARIAN_HEARTBEAT_SECONDS (default 1200 = 20 min) a
running search now logs that it is still going, with its uuid, the search
phrase, hits so far, elapsed minutes, and a rough how-far-along.
The estimate is deliberately cheap: the producers ALREADY record a byte
offset per chunk file (the resume watermarks), and the total size is
stat()'d once per search when the chunk list is discovered. A reading is
then just a sum over ~40 ints - nothing extra happens per line, and no
cycles are spent estimating how many cycles are left.
search_for_doi takes an optional progress dict it fills with the live
positions dict + total_bytes; the librarian publishes the running search
(uuid/query/progress/live hits) while the scan runs and clears it in
finally. Nothing running => the heartbeat stays quiet.
Tests: percentage maths incl. unknown-total and >100% clamping, the
register/clear round-trip, and an end-to-end check that a real scan fills
progress so the offsets cover the chunk files on disk. Suite: 58 unit +
65 integration green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The production bot tracks a SEPARATE image, conjurer-bot-deploy, promoted
only when a commit message contains [deploy]. On such a commit the already-
built conjurer-bot:<sha> is re-tagged (same bytes, no rebuild) and pushed as
conjurer-bot-deploy:<sha>; the deploy repo's image-updater tracks it. The
test bot + librarian keep updating on every build.
Merging THIS commit is deliberately the FIRST [deploy]: it lands the
promotion step AND carries the trigger, so it bootstraps conjurer-bot-deploy
to this build - a clean baseline with every image at the same, current version.
[deploy]
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A repeat of the same query (whitespace/case-normalised, scoped by
deep-vs-shallow) returns the stored hits and skips the whole Crossref call
and DB scan. Disk-backed (survives restart), TTL'd
(CONJURER_LIBRARIAN_CACHE_TTL, default 7d; 0 disables) and size-bounded
(CONJURER_LIBRARIAN_CACHE_MAX, default 500). Reuses DiskQueue, so it's a
handful of lines. Nothing fancy - exact (normalised) match, not fuzzy.
Checked before Crossref only on a fresh search (a resume from checkpoint
still continues its scan), and stored after a completed search.
Tests: hit/miss, normalisation, deep/shallow separation, expiry, disable,
prune. Suite: 58 unit + 59 integration green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
So one librarian can serve several bots (test + deploy) instead of firing
every result/pong at a single static CONJURER_MAIN_BOT.
* The bot includes its own callback address (CONJURER_SELF_CALLBACK) in
every /query and /ping.
* The librarian stores that callback with the query (persisted with the
request, so a replay after restart still answers the right bot) and, for
results, in the OUTBOX entry ({target, payload}) so the resender delivers
to the origin bot even across a librarian restart.
* Pongs go back to the pinging bot too - otherwise a second bot's health
check would be ponged to the first and always time out, so it could
never enable its librarian cog.
* Empty callback falls back to MAIN_BOT_ADDRESS, and a legacy OUTBOX entry
(raw payload, pre-callback) is still delivered to the default bot, so the
upgrade is seamless.
Tests: per-origin result delivery + legacy-shape fallback (outbox),
busy/idle pong routed to the callback bot vs default (lifecycle). Suite:
58 unit + 52 integration green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A restart of the librarian used to throw away an in-flight search (and any
searches still queued). Now search state survives a restart:
* Resumable DB scan (search_bot): each producer records a tell()-cookie
watermark per chunk file as it goes (safe because search_for_doi drains
the work queue before returning), and can seek back to it. search_for_doi
now takes stop_event + resume and returns (result_list, positions,
interrupted).
* Persisted requests: /query writes the accepted request to a disk queue
before enqueuing; replay_requests re-enqueues unfinished ones on startup.
So even a search still waiting in the queue survives a restart.
* Checkpoints: when a graceful shutdown interrupts a scan, the librarian
writes {dois, found-so-far, per-file offsets}. On restart answer_query
loads it, skips the (already done) Crossref+refine, and continues the
scan from the saved offsets with the found DOIs pre-marked - no line is
read twice and none is missed. A finished or crashed search forgets its
request+checkpoint (no poison-pill replay).
* Graceful shutdown: SIGTERM/SIGINT set a shutdown event; the running scan
checkpoints and the worker stops. The main thread then exits within a
BOUNDED window (CONJURER_LIBRARIAN_GRACEFUL_TIMEOUT, default 45s) so the
pod can never become an un-killable zombie. Needs terminationGracePeriod
>= that in the deploy (separate PR).
Tests: search_bot resume correctness (seek past scanned, don't miss/re-scan;
stop_event -> interrupted) and librarian state mechanics (request replay,
forget, checkpoint round-trip, poison-pill drop). Suite: 58 unit + 49
integration green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two hygiene fixes on top of the work-queue OOM bound:
Result dumps: cr_results / rr_results / s_results.json were write-only
(nothing reads them) yet accumulated EVERY search forever and json.load'd
the whole growing file on each write - unbounded RAM and PVC growth, and
for a deep search the raw cr_results dump is hundreds of MB. They are now
off by default (CONJURER_LIBRARIAN_DEBUG_DUMPS) and, when enabled, are
overwritten with just the latest search - never loaded or accumulated.
not_in_db.json is untouched: it's a real queue the scraper drains.
Search logging: search_bot logged via print(), including a per-line
carriage-return progress line that flooded stdout / the log file with
millions of entries - fine for a desktop app, unreadable and bloating in
a container. All of it is now proper logging at DEBUG (with coarse
per-500k-line progress), so a normal run is quiet. The librarian log
level is configurable (CONJURER_LIBRARIAN_LOG_LEVEL, default INFO) and a
stdout handler is added so stays useful now that the
search no longer prints straight to stdout. Set DEBUG for full verbosity.
Also: make test_result_delivery_contract hermetic (point the durable
spool at a temp dir so it can't pollute or be poisoned by the real
result_inbox/ between runs) and gitignore the runtime spool dirs.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The pod restarted spontaneously mid-search (no liveness probe is set, so
it was the kernel OOM-killer against the 1Gi limit). Cause: search_bot
built its work queue with maxsize 35_500_000. The producers stream the
WHOLE DOI database (tens of millions of lines across chunks) into it while
a few consumers drain, so the queue could buffer gigabytes of lines -
blowing the 1Gi container and taking the whole in-flight search with it.
Bound the queue (default 100k lines, env CONJURER_LIBRARIAN_WORKQ_SIZE),
so producers backpressure to consumers and RAM stays in the low MB.
Because a bounded queue means a producer can now block on a FULL queue,
make the producer's put timeout-poll the TERM sentinel, so a full queue
whose consumers have already finished (all DOIs found) can never deadlock
it. New test pins that: tiny queue + target on line 1 + thousands of
trailing decoys still terminates and finds the target.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Librarian.__init__ reads the Crossref contact from CONJURER_CROSSREF_MAILTO,
then tries to override it from a 'crossref' netrc entry. When no netrc is
mounted (the normal container setup - default /root/.netrc) the read raises
FileNotFoundError and it logged 'Crossref credentials missing in netrc ...'
on EVERY search, even though the env var was set and used. Pure noise.
Only warn when there is genuinely no contact from either source (env unset
AND netrc unreadable) - which is also the case that then raises. When the
env var is set, a missing netrc is expected and logged at debug.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
check_self starts in setup() (during on_ready), and its first tick can
fire before the gateway has populated the channel cache - get_channel()
then returns None and channel.history() raises AttributeError. Because an
unhandled exception stops a tasks.loop, this also silently killed the log
rollover and spontaneous messages the loop is responsible for.
Add a before_loop that awaits wait_until_ready() (fixes the startup race)
and a None guard that skips the tick with a warning (keeps the loop alive
even if the channel is genuinely unreachable: wrong id / not in the guild
/ missing permission).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
An 8h search result must survive a transient bot outage, an api/address
misroute, or a restart of either side. Make the librarian->bot result
path durably at-least-once with idempotent rendering:
Shared: durable_queue.DiskQueue - a dependency-free, atomically-written,
one-file-per-key disk queue (unit-tested), shared by both images
(added to Dockerfile.librarian; the bot already COPYs *.py).
Librarian (sender): finished results go to a persistent OUTBOX before
sending; delivery retries with backoff; an entry is removed only on a
positive ACK; a resender thread keeps flushing the OUTBOX, so a result
survives a bot outage AND a librarian restart (OUTBOX is on the state
volume) - it simply keeps trying until acked.
Bot (receiver): /conjurer is now idempotent and durable - each result is
persisted to an INBOX before acking and only queued if its uuid was not
already delivered (dropped as a duplicate) or already pending. Once the
cog actually renders it, mark_delivered() records the uuid and clears the
inbox, so the librarian's resends become no-ops. On startup the bot
replays any accepted-but-unrendered result from the INBOX, so a bot crash
mid-flight doesn't lose it. Pongs stay ephemeral.
Together: the librarian keeps a result until the bot confirms it; the bot
keeps it until it is on screen; duplicates never double-render. Combined
with the deploy return-path fix, an expensive result no longer vanishes.
Tests: unit test_durable_queue; integration test_librarian_outbox
(retry/backoff, resend survives outage) and test_result_durable_delivery
(persist, dedup pending, dedup delivered, replay, pong not persisted).
Suite: 55 unit + 39 integration green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Diagnostic coverage for the 'search vanished' report. Proves the bot side
of result delivery is correct end to end (right shape reaches IN_COMM_Q;
empty result still delivered; wrong api-key -> 401 vanish; uuid mismatch
-> orphaned away from the querent), which isolates a SYSTEMATIC vanish to
transport: the librarian being unable to reach the bot's /conjurer at all
(CONJURER_MAIN_BOT). The bot Service is NodePort with no pinned nodePort
while the librarian hardcodes :32442 - and being in-cluster it should use
the Service DNS http://bot:5000 instead.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two refinements to the librarian health/delivery story, matching how it
actually behaves under load:
1. Busy-aware ping (case b - broken return path). A ping arriving while
the worker is grinding a search no longer queues behind it (which made
a healthy-but-busy librarian time out and look dead). The librarian
tracks worker_busy and, when set, pongs back IMMEDIATELY without
touching the queue. Being busy is fine - you can keep piling searches
on. The ping still travels the librarian->bot return path, so it keeps
catching the one thing it must: a disrupted/incompatible return path
where queries vanish. Idle pings still go through the internal queue.
2. Per-query watchdog (case a - finished but result lost). The librarian
now tracks every search uuid's lifecycle (queued -> processing ->
gone) in active_queries, exposed via a new POST /query_status. After
dispatching a search the bot records it in self.pending; watch_pending
polls /query_status for each. While the librarian still knows the uuid
the search is progressing - left alone. The moment a uuid VANISHES
there while still pending on the bot, its result was computed but never
delivered: after a grace window (to rule out an in-flight result) the
bot posts a notice to the channel - but ONLY then. A normally delivered
result is popped from self.pending by check_data_q and never flagged.
Hardening: the worker's search body is now wrapped in try/except/finally
so a crashing search can't kill the worker thread (which would freeze the
queue), and worker_busy / active_queries are always cleared. The grace
logic lives in a dependency-free librarian_watchdog.pending_verdict so it
is unit-testable without discord/pdf libs. /ping and /query_status are
plain (sync) views so they run without flask[async].
Tests: unit test_librarian_watchdog (verdict transitions); integration
test_librarian_query_lifecycle (query_status known/unknown + auth,
idle-ping-queues, busy-ping-pongs-directly). Suite: 28 integration + 48
unit green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The librarian health check was a plain GET to '/', which only proved
Flask was listening - not that the service could actually take a query,
run it through its internal queue+worker, and answer back. So the cog
could load against a librarian whose worker was wedged or that couldn't
reach the bot on the return leg.
Replace it with a ping that travels the SAME path a real search does, on
both sides:
bot: QueryControl -> OUT_COMM_Q -> scan_queue -> awaiting_q
librarian: POST /ping -> librarian_queue -> worker pulls it off
(no Crossref/DOI search) -> pongs back with the same uuid
bot: /conjurer -> incoming_q -> scan_incoming matches uuid, wakes waiter
The cog enables only when that whole loop closes within 3s. This also
proves the librarian->bot return path, which a GET never did.
Safety: uuid is random per ping; the wait and POST are both bounded so
startup can't stall; a pong that finds no waiter is dropped (never
orphaned into IN_COMM_Q, which would make the cog post a bogus 'no
results' message); and a ping whose pong never returns is swept out of
awaiting_q after PING_TTL_SECONDS so nothing leaks. All awaiting_q writes
stay within scan_queue (append) and scan_incoming (remove) - no locks,
no cross-thread mutation.
Integration tests cover: OK round-trip, timeout when accepted-but-no-pong,
unreachable, non-200, orphan-pong-dropped, and that real results still
reach IN_COMM_Q. Suite: 24 integration + 41 unit green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
When a service group's cog stays disabled the log now says WHY:
_service_health returns the actual connection error (ConnectionError
= refused/down, ConnectTimeout = firewall/slow, gaierror = DNS/wrong
host) instead of a bare 'unreachable', and on_ready logs the resolved
FILE/LIBRARIAN/RADIO addresses so an env-var that never reached the
process (address falls back to the 192.168.1.15:5000 default) is
obvious at a glance.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Seven confirmed defects from the code audit, each small and low-risk.
* ai_functions.get_random_cyclic_message: random.randint(0, len(CYCLIC_WORDS))
is inclusive -> could return len -> IndexError. Now randrange(len) + guard on
an empty CYCLIC_WORDS.
* librarian_commands.get_image_sadox: random.randrange(0, len(res)-1) never
picked the last comic and raised ValueError('empty range') on a single file.
Now randrange(len) + an empty-dir guard.
* ai_commands image generation: every DALL-E error branch replied but did not
return, so control fell through to `if response:` with response unbound ->
UnboundLocalError right after the friendly message. Each branch now returns;
response is pre-initialised; and PermissionDeniedError no longer passes a
(message, text) tuple as a single arg.
* search_bot DOI match: `item["DOI"] in data` was a substring test, so a DOI
that is a prefix of a longer one (10.1/1 vs 10.1/12) produced a false 'exists'
hit. Now matches the line's first whitespace token exactly, via an O(1) dict
index built once per consumer (also removes the O(queried-DOIs) per-line scan
- a real win for large databases).
* communication_subroutine.scan_incoming: matched records were never removed
from awaiting_q, so it grew unbounded over uptime and a reused UUID could
re-match a stale record. Matched records are now dropped after dispatch.
* communication_subroutine.id3: (resp.headers.get("icy-name") or "").title()
guards against a stream that omits headers (was AttributeError on None,
500-ing the /prepped_tracks "next" handler).
* betoniarka.scan_tracks: waits for the radio logs to exist instead of dying
with FileNotFoundError on a fresh deploy (which silently killed the
now-playing forwarder until a restart).
Verified: tests/unit/test_search_bot.py gains exact-match and trailing-metadata
cases; full unit job 43 passed. Remaining observations (image-gen stale
/home/pi fallback paths + dead FileNotFoundError-after-OSError branch; tailer
still vulnerable to mid-run log rotation; DOI-first-token assumption) noted for
follow-up - none are crashes on the normal path.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Fits the mythology pillar of the persona (Slavic/Norse/Celtic, Old Norse
phrases). New always-loaded cog oracle_commands:
* $runy [pytanie] draws three Elder Futhark runes (past/present/future, with
upright/reversed orientation - the 8 symmetric runes are never reversed) and
asks the ACTIVE AI backend to read the spread in Conjurer's voice. If the AI
is down it still shows the drawn runes with their own meanings, so the command
always answers.
* $runa_dnia gives one rune, deterministic per user per day (sha256 seed), so
it's stable if asked repeatedly - no AI call, no state file.
The full 24-rune Futhark, the draw logic and the reversal rules are pure and
unit-tested (distinct draw, non-invertible never reversed, per-day stability,
meaning fallback).
Also fixes a pre-existing unit-job breakage: test_bar_commands and
test_lore_commands each stubbed `discord` with different completeness and
shared sys.modules, so once both landed on main the one lacking `discord.ext.tasks`
shadowed the one needing it and collection failed order-dependently. A new
tests/unit/conftest.py stubs discord once, completely, before any test module -
the per-file stubs then skip. Full unit job: 41 passed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The conversation memory file grows forever (every chat appends a user+assistant
pair), so startup load gets slower and the disk fills. New always-loaded cog
lore_commands turns that growth into content: a background task summarises the
oldest slice into one in-character "legend" via the ACTIVE AI backend, replaces
those old messages with the summary (bounding the file, keeping continuity for
the next startup's context), archives it to legendy.json, and announces it on
Safety: the compaction transforms (build_transcript, apply_compaction) are pure
and unit-tested. The file rewrite is re-read -> back up -> atomic write with no
await in between, so a handle_response append that lands while the summary is
being generated can neither be lost (it's in the preserved tail) nor corrupt
the file (single-threaded, no interleave). A .bak is kept. Scope note: this
bounds the on-disk file (startup/disk); the in-RAM MESSAGE_TABLE is a separate
concern left untouched to avoid yanking context from a live conversation.
Commands: $zapisz_legende (Vykidailo) forces a compaction now; $legendy recalls
a random past legend. All thresholds env-overridable (CONJURER_MEMORY_COMPACT_*,
CONJURER_LEGENDS_CHANNEL). constants gains LEGENDS_FILE + config + seed; bot.py
registers the cog.
Verified: tests/unit/test_lore_commands.py covers prefix-replace/tail-keep,
preservation of appends made during summarisation, and transcript formatting +
head/tail truncation. Unit job 32 passed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The most in-character capability the bot has: the persona is literally a 200kg
bartender who mixes strong drinks with intriguing names. New always-loaded cog
bar_commands:
* $nalej [motyw] asks the ACTIVE AI backend (whatever $gadaj_teraz selects) to
invent one themed cocktail in Conjurer's voice - persona reused from
GPT_SETTINGS[0] as a system message, instructions as the user turn, via
handle_response request_type NONE so it never pollutes the bar's conversation
memory. Empty motyw = a surprise; "radio"/"pod muzykę" themes the drink on the
track currently playing (PREPPED_TRACKS["now_playing"]).
* every drink is appended to menu.json (new seeded state file, CONJURER_MENU_FILE
overridable) with name/theme/author/timestamp/full text - emergent bar lore.
* $menu lists the invented drinks and pours one at random from the archive.
Text-only for now; a DALL-E drink image is an easy follow-up (the render path
already exists in ai_commands, OpenAI-only).
constants gains MENU_FILE (next to pamiec.json by default) + its seed; bot.py
registers bar_commands as a core cog. Verified: tests/unit/test_bar_commands.py
covers name extraction (markers/markdown/fallback) and the menu round-trip
incl. corrupt-file tolerance; unit job 32 passed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two field-reported robustness gaps on top of the hang fix.
1) A stray non-UTF-8 byte in a chunk (0x96 in the report) raised
UnicodeDecodeError from readline() - which is a ValueError, so the earlier
`except OSError` did NOT catch it. The finally-sentinel meant no hang, but the
producer died mid-file with a loud traceback and every DOI after the bad byte
went unsearched. Now chunks are opened with errors="replace" (bad bytes become
U+FFFD; DOIs are ASCII so a match is never affected) so the read runs to EOF,
and the producer's except is broadened from OSError to Exception so no per-file
error can ever crash the thread - it's logged and the sentinel still fires.
2) The result-send back to the bot (BackgroundTaskSearch._run) now logs exactly
what goes out - target URL, uuid, DOI count and the DOI list - so the librarian
log plainly shows a result was sent and what was in it. And a failed POST is no
longer fatal: a RequestException used to propagate out of the worker loop and
kill the thread, stalling every future query until restart; it's now caught and
logged, and a non-200 from the bot is logged as a warning.
Verified: tests/unit/test_search_bot.py gains a case writing a chunk with a 0x96
byte before a valid DOI and asserting that DOI is still found (file read to
completion, not aborted). All 5 search_bot unit tests pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
test_musician_auth.py still probed /clear_pr_pls on the musician, but that
endpoint moved to betoniarka during the split - the musician now 404s it, so
all three tests failed 404 != 401/200. This was pre-existing debt, unrelated to
the AI/share/bridge work; it just kept the integration job red.
Split the coverage to match the current architecture:
* test_musician_auth.py exercises the same auth contract (no key -> 401, key ->
200, key unset -> open) against /get_share_list, an authenticated endpoint the
musician still serves, with a valid body so the permitted case is a clean 200
rather than a 400;
* new test_betoniarka_auth.py covers /clear_pr_pls where it now lives, pointing
PRIORITY_PLAYLIST_PATH at a tmp file so the authorised case can truncate it,
and checks /ping stays open;
* conftest.py adds conjurer_betoniarka to sys.path so the service imports.
Verified in a clean venv (pytest flask waitress requests), matching the CI
integration job: 13 passed, up from 3 failed / 6 passed. Unit suite unaffected
(23 passed).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two connected features.
1) AI query interface (via the comm layer). communication_subroutine gains an
AI_QUERY_Q, a submit_ai_query() in-process entry point, and an authed
POST /ai_query endpoint ({prompt, channel_id, request_type?, username?}). The
prompt is queued and answered asynchronously by a new tasks.loop worker in the
always-loaded AI cog (Events), which calls handle_response - so it runs on
whichever backend $gadaj_teraz currently selects (GPT or Claude) - and posts the
answer to the requested channel, chunked to Discord's limit. request_type "NONE"
(default) is a clean one-shot: no persona system prompt, no memory write. The
worker starts before the OpenAI guard in cog_load, so it also runs on a
Claude-only box; cog_unload cancels it.
2) Librarian AI review. New command $wyszukaj_z_recenzja mirrors
$wyszukaj_linki_do_dokumentow but sets ai_review=True on the QueryControl, which
rides the round-trip and is matched back by UUID. When the hits return,
check_data_q sends the raw list as before, then - if flagged - hands the same
list (already in Crossref-relevance order) plus the search phrase to the AI
queue for a weighted re-rank and per-source review, delivered to the same
channel. QueryControl gains an ai_review flag (default False, so the orphan path
and all existing callers are unaffected).
Confirmed separately (and noted in the docs): the DOI list the AI receives is
pre-sorted by Crossref relevance - the librarian pipeline only filters (drops
title-less items) and splits (in-db / not-in-db), never re-sorts, and relies on
insertion-ordered dicts (Py 3.7+).
Verified: /ai_query auth (401/200/400/open), submit_ai_query and the queued
dict shape, and the QueryControl flag - via a Flask test client and
tests/integration/test_ai_query_endpoint.py (5 tests, all pass; integration
suite 11 passed, the 3 failures are the pre-existing /clear_pr_pls musician
tests fixed on a separate branch). Full first-party compile clean.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
search_bot conflated MAXTHREADS into two jobs at once - how many chunk files to
read (files 0..MAXTHREADS-1) AND how many producer sentinels to wait for - so
the two had to match exactly. Set too low it silently skipped trailing chunks;
set too high (or with any chunk missing/unreadable) a producer crashed before
emitting its sentinel, the consumers' count never reached the threshold, and
search_for_doi hung on join() forever. The idle-timeout failsafe that was meant
to break a starved consumer was dead code: `if empty_counter > 5: ... elif
empty_counter > 10: break` - >10 implies >5, so the elif never ran.
Fix, three layers:
* auto-discover the chunk files present (discover_chunk_files: <n>_chunk.txt in
numeric order) instead of range(0, MAXTHREADS). All files are read regardless
of count, and no producer is ever pointed at a missing file;
* the sentinel threshold is now the number of producers actually started, so it
can't drift from what's emitted;
* producers emit their sentinel in a finally, so even a crash (missing/unreadable
chunk) can't starve the count; and the idle backstop is reordered so it can
actually fire (>EMPTY_LIMIT seconds) as a last resort.
MAXTHREADS is deprecated and unused (kept only so old env files don't break);
docs/env updated to say chunk files are auto-discovered.
For the reported case (MAXTHREADS=40, files 0..43): before, files 40-43 were
silently never searched, and any run that referenced a missing chunk hung
forever. After, all 44 are searched and it always terminates.
Verified in a pytest-only venv (tests/unit/test_search_bot.py): DOI in a
trailing chunk is found; an unreadable chunk still terminates; empty dir returns
at once; discovery is numeric-sorted. Full unit job 27 passed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Reading radio_conjurer.liq shows request.playlist is not a Liquidsoap playlist
at all - it is a drop box. queue_processing() runs every 60s, reads every line,
pushes each into request.queue() as a URI, then deletes the file and recreates
it empty. That changes what writing to it means, so the mode is now named and
documented for what it is.
BRIDGE_MODE=dropbox (with "playlist" kept as an alias) appends the exact
translated path to that file. Because the lines are pushed as URIs it is an
exact hand-off - no keyword search - and it involves neither betoniarka nor the
bot, just the file and Liquidsoap. The liq also puts requests_queue first in the
fallback and does not apply the check_next replay guard to it, so a request
interrupts the rotation and plays even if the track ran recently; both are now
documented rather than left to be discovered.
Fixes found while wiring this up:
* the existence check was unconditional while the library mount was documented
as optional, so a drop-box-only setup could never queue anything. It is now
BRIDGE_VERIFY_FILE_EXISTS (auto|true|false), defaulting to checking only when
the library is actually visible;
* a missing parent dir was silently created, which would swallow requests into
the container's own filesystem when the radio's data dir was not mounted. It
is now a hard error naming the likely cause.
Verified: alias resolves; append/drain/append cycle against a simulation of the
liq's read-remove-recreate; both misconfigurations raise instead of silently
succeeding; drop-box-only path works with the library absent; api mode
unchanged, still prepending the sentinel that wyszukaj() discards.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Star a track in Navidrome and it lands in betoniarka's request queue.
EXTERNAL component: imports nothing from Conjurer and bypasses the Discord bot
and its API entirely. It speaks only to Navidrome's Subsonic API and to the
radio operator, so the bot can be down and this keeps working.
Not shipped as a .ndp plugin, deliberately. Navidrome's plugin capabilities are
MetadataAgent, Scrobbler, Lyrics, SonicSimilarity, TaskWorker, Lifecycle,
SchedulerCallback and WebSocketCallback - there is no UI extension capability
(a plugin cannot add a button) and no star/love event to react to. The only
plugin-shaped alternative, Scrobbler, sees every track played, which is a
firehose rather than a one-click request. So the trigger is Navidrome's own
star control, which also works from any Subsonic client including phones.
The README documents this with sources.
Both sides index the same library under different mount points, so every path
is translated between the two roots; paths reported relative to Navidrome's
library are handled as well as absolute ones.
Two delivery modes: "api" (default) posts to /request_radio_file and needs no
shared filesystem, at the cost of the radio re-finding the track by keywords;
"playlist" appends the exact translated path and is exact but needs the radio's
data dir mounted. The api path prepends a sentinel token because betoniarka's
wyszukaj() drops lista_slow[0], where the Discord command word normally sits.
Verified: path translation for both absolute and relative forms; keyword
extraction (no regex metacharacters, extension dropped, Polish characters
intact); and an end-to-end check feeding the generated keywords through
betoniarka's real wyszukaj() against a library seeded with near-miss traps
(same artist, live version of the same title, same album name under another
artist) - it resolved to the exact intended file.
Docker/compose/systemd install paths documented. Not run against a live
Navidrome or radio - no instance reachable from here.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The Python half of the short-lived share links lived in the repo; the Apache
config and the cron entries that make it work were hand-placed on the host, so
the feature could not be rebuilt from a checkout. This adds the missing half.
Scripts (defaults unchanged, so existing bare-metal cron keeps working):
* scan_shares.py / revoke_shares.py take their paths from CONJURER_SHARE_*
instead of hardcoding the Pi layout, and create their parent dirs;
* the revoke TTL is now CONJURER_SHARE_TTL_SECONDS. Its --help claimed "2min"
while the code used a hardcoded 3600 - the help text now reports the real,
configured value.
New share service:
* docker/Dockerfile.share - Apache + the two jobs, reusing the same scripts
rather than forking copies;
* docker/share-vhost.conf.tpl - the previously undocumented Apache half.
Three settings are load-bearing and commented as such: +FollowSymLinks (the
shares ARE symlinks), -Indexes (a listing would expose every live token), and
a deny rule for dotfiles (revoke_shares.py keeps .downloads.json - a map of
every live token - inside the served directory);
* docker/entrypoint.share.sh - renders the vhost, seeds the index on first run,
then runs scanner/revoker in sleep loops beside Apache (no cron, so their
output shows up in docker logs);
* compose + env example, incl. the two volumes that MUST be shared with the
musician (it creates the links and reads the index).
docs/deployment/FILE_SHARING.md documents the mechanism, both deployment routes
(docker and existing host Apache), the Discord command, why each Apache setting
matters, verification commands, and the known limitations - notably that a link
nobody ever downloads is never revoked, since the TTL starts at first download.
Verified: scanner and revoker exercised end-to-end against temp dirs (index
built; token recorded from a combined-format log line and the symlink unlinked).
Compose file not validated - no docker on this machine.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
file_search_functions was the only bot-side service client that hardcoded the
musician's LAN address and sent no auth header. Both musician endpoints it
calls (/get_share_list, /get_share_links) run _authorize_request(), so the
whole file-share feature 401'd on any deployment that set CONJURER_API_KEY -
which the deployment guides instruct you to do - and the hardcoded IP made a
musician on another host (the Proxmox/docker split) unreachable regardless.
Use FILE_SERVICE_ADDRESS + service_headers() like music_functions does, add
the two endpoint paths to constants alongside the existing ones, and set an
explicit timeout so a wedged musician can't hang the calling cog.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Lets CI be triggered on demand from the Actions tab or `gh workflow run`,
instead of pushing a commit just to exercise the pipeline.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Quality-of-life: checking which AI backend is live no longer requires
switching to it. Bare `$gadaj_teraz` replies with the active config and the
selectable ones; that read-only path is open to everyone, while switching
stays Vykidailo-gated.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>