13d2a040525add23f108032c403f0961ef633694
12 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ac16b77f56 |
Librarian: graceful shutdown with resumable search state
CI / compile (pull_request) Successful in 10s
CI / unit (pull_request) Successful in 28s
CI / integration (pull_request) Successful in 27s
build / build (push) Successful in 42s
CI / compile (push) Successful in 10s
CI / unit (push) Successful in 26s
CI / integration (push) Successful in 25s
A restart of the librarian used to throw away an in-flight search (and any
searches still queued). Now search state survives a restart:
* Resumable DB scan (search_bot): each producer records a tell()-cookie
watermark per chunk file as it goes (safe because search_for_doi drains
the work queue before returning), and can seek back to it. search_for_doi
now takes stop_event + resume and returns (result_list, positions,
interrupted).
* Persisted requests: /query writes the accepted request to a disk queue
before enqueuing; replay_requests re-enqueues unfinished ones on startup.
So even a search still waiting in the queue survives a restart.
* Checkpoints: when a graceful shutdown interrupts a scan, the librarian
writes {dois, found-so-far, per-file offsets}. On restart answer_query
loads it, skips the (already done) Crossref+refine, and continues the
scan from the saved offsets with the found DOIs pre-marked - no line is
read twice and none is missed. A finished or crashed search forgets its
request+checkpoint (no poison-pill replay).
* Graceful shutdown: SIGTERM/SIGINT set a shutdown event; the running scan
checkpoints and the worker stops. The main thread then exits within a
BOUNDED window (CONJURER_LIBRARIAN_GRACEFUL_TIMEOUT, default 45s) so the
pod can never become an un-killable zombie. Needs terminationGracePeriod
>= that in the deploy (separate PR).
Tests: search_bot resume correctness (seek past scanned, don't miss/re-scan;
stop_event -> interrupted) and librarian state mechanics (request replay,
forget, checkpoint round-trip, poison-pill drop). Suite: 58 unit + 49
integration green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
40605b959f |
Librarian: bound the DOI search work queue to stop OOM-killing the pod
CI / compile (pull_request) Successful in 9s
CI / unit (pull_request) Successful in 17s
CI / integration (pull_request) Successful in 26s
build / build (push) Successful in 23s
CI / compile (push) Successful in 7s
CI / unit (push) Successful in 22s
CI / integration (push) Successful in 26s
The pod restarted spontaneously mid-search (no liveness probe is set, so it was the kernel OOM-killer against the 1Gi limit). Cause: search_bot built its work queue with maxsize 35_500_000. The producers stream the WHOLE DOI database (tens of millions of lines across chunks) into it while a few consumers drain, so the queue could buffer gigabytes of lines - blowing the 1Gi container and taking the whole in-flight search with it. Bound the queue (default 100k lines, env CONJURER_LIBRARIAN_WORKQ_SIZE), so producers backpressure to consumers and RAM stays in the low MB. Because a bounded queue means a producer can now block on a FULL queue, make the producer's put timeout-poll the TERM sentinel, so a full queue whose consumers have already finished (all DOIs found) can never deadlock it. New test pins that: tiny queue + target on line 1 + thousands of trailing decoys still terminates and finds the target. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
44b7298a15 |
Durable result delivery: OUTBOX + idempotent INBOX so results never die
An 8h search result must survive a transient bot outage, an api/address misroute, or a restart of either side. Make the librarian->bot result path durably at-least-once with idempotent rendering: Shared: durable_queue.DiskQueue - a dependency-free, atomically-written, one-file-per-key disk queue (unit-tested), shared by both images (added to Dockerfile.librarian; the bot already COPYs *.py). Librarian (sender): finished results go to a persistent OUTBOX before sending; delivery retries with backoff; an entry is removed only on a positive ACK; a resender thread keeps flushing the OUTBOX, so a result survives a bot outage AND a librarian restart (OUTBOX is on the state volume) - it simply keeps trying until acked. Bot (receiver): /conjurer is now idempotent and durable - each result is persisted to an INBOX before acking and only queued if its uuid was not already delivered (dropped as a duplicate) or already pending. Once the cog actually renders it, mark_delivered() records the uuid and clears the inbox, so the librarian's resends become no-ops. On startup the bot replays any accepted-but-unrendered result from the INBOX, so a bot crash mid-flight doesn't lose it. Pongs stay ephemeral. Together: the librarian keeps a result until the bot confirms it; the bot keeps it until it is on screen; duplicates never double-render. Combined with the deploy return-path fix, an expensive result no longer vanishes. Tests: unit test_durable_queue; integration test_librarian_outbox (retry/backoff, resend survives outage) and test_result_durable_delivery (persist, dedup pending, dedup delivered, replay, pong not persisted). Suite: 55 unit + 39 integration green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
5d321f2f5b |
Librarian: busy-aware ping + per-query lost-result watchdog
CI / compile (pull_request) Successful in 10s
CI / unit (pull_request) Successful in 20s
CI / integration (pull_request) Successful in 20s
CI / compile (push) Successful in 10s
CI / unit (push) Successful in 21s
CI / integration (push) Successful in 20s
build / build (push) Successful in 57s
Two refinements to the librarian health/delivery story, matching how it actually behaves under load: 1. Busy-aware ping (case b - broken return path). A ping arriving while the worker is grinding a search no longer queues behind it (which made a healthy-but-busy librarian time out and look dead). The librarian tracks worker_busy and, when set, pongs back IMMEDIATELY without touching the queue. Being busy is fine - you can keep piling searches on. The ping still travels the librarian->bot return path, so it keeps catching the one thing it must: a disrupted/incompatible return path where queries vanish. Idle pings still go through the internal queue. 2. Per-query watchdog (case a - finished but result lost). The librarian now tracks every search uuid's lifecycle (queued -> processing -> gone) in active_queries, exposed via a new POST /query_status. After dispatching a search the bot records it in self.pending; watch_pending polls /query_status for each. While the librarian still knows the uuid the search is progressing - left alone. The moment a uuid VANISHES there while still pending on the bot, its result was computed but never delivered: after a grace window (to rule out an in-flight result) the bot posts a notice to the channel - but ONLY then. A normally delivered result is popped from self.pending by check_data_q and never flagged. Hardening: the worker's search body is now wrapped in try/except/finally so a crashing search can't kill the worker thread (which would freeze the queue), and worker_busy / active_queries are always cleared. The grace logic lives in a dependency-free librarian_watchdog.pending_verdict so it is unit-testable without discord/pdf libs. /ping and /query_status are plain (sync) views so they run without flask[async]. Tests: unit test_librarian_watchdog (verdict transitions); integration test_librarian_query_lifecycle (query_status known/unknown + auth, idle-ping-queues, busy-ping-pongs-directly). Suite: 28 integration + 48 unit green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
defc482a22 |
fix: batch A - crash bugs, a leak, and startup/edge fragility
CI / compile (pull_request) Successful in 10s
CI / unit (pull_request) Successful in 19s
CI / integration (pull_request) Successful in 10s
build / build (push) Failing after 46m40s
CI / compile (push) Successful in 20s
CI / unit (push) Successful in 19s
CI / integration (push) Successful in 16s
Seven confirmed defects from the code audit, each small and low-risk.
* ai_functions.get_random_cyclic_message: random.randint(0, len(CYCLIC_WORDS))
is inclusive -> could return len -> IndexError. Now randrange(len) + guard on
an empty CYCLIC_WORDS.
* librarian_commands.get_image_sadox: random.randrange(0, len(res)-1) never
picked the last comic and raised ValueError('empty range') on a single file.
Now randrange(len) + an empty-dir guard.
* ai_commands image generation: every DALL-E error branch replied but did not
return, so control fell through to `if response:` with response unbound ->
UnboundLocalError right after the friendly message. Each branch now returns;
response is pre-initialised; and PermissionDeniedError no longer passes a
(message, text) tuple as a single arg.
* search_bot DOI match: `item["DOI"] in data` was a substring test, so a DOI
that is a prefix of a longer one (10.1/1 vs 10.1/12) produced a false 'exists'
hit. Now matches the line's first whitespace token exactly, via an O(1) dict
index built once per consumer (also removes the O(queried-DOIs) per-line scan
- a real win for large databases).
* communication_subroutine.scan_incoming: matched records were never removed
from awaiting_q, so it grew unbounded over uptime and a reused UUID could
re-match a stale record. Matched records are now dropped after dispatch.
* communication_subroutine.id3: (resp.headers.get("icy-name") or "").title()
guards against a stream that omits headers (was AttributeError on None,
500-ing the /prepped_tracks "next" handler).
* betoniarka.scan_tracks: waits for the radio logs to exist instead of dying
with FileNotFoundError on a fresh deploy (which silently killed the
now-playing forwarder until a restart).
Verified: tests/unit/test_search_bot.py gains exact-match and trailing-metadata
cases; full unit job 43 passed. Remaining observations (image-gen stale
/home/pi fallback paths + dead FileNotFoundError-after-OSError branch; tailer
still vulnerable to mid-run log rotation; DOI-first-token assumption) noted for
follow-up - none are crashes on the normal path.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
e1fca864d8 |
oracle: $runy / $runa_dnia - Elder Futhark rune readings in-character
CI / compile (pull_request) Successful in 10s
CI / unit (pull_request) Successful in 14s
CI / integration (pull_request) Successful in 10s
build / build (push) Successful in 38s
CI / compile (push) Successful in 10s
CI / unit (push) Successful in 17s
CI / integration (push) Successful in 11s
Fits the mythology pillar of the persona (Slavic/Norse/Celtic, Old Norse phrases). New always-loaded cog oracle_commands: * $runy [pytanie] draws three Elder Futhark runes (past/present/future, with upright/reversed orientation - the 8 symmetric runes are never reversed) and asks the ACTIVE AI backend to read the spread in Conjurer's voice. If the AI is down it still shows the drawn runes with their own meanings, so the command always answers. * $runa_dnia gives one rune, deterministic per user per day (sha256 seed), so it's stable if asked repeatedly - no AI call, no state file. The full 24-rune Futhark, the draw logic and the reversal rules are pure and unit-tested (distinct draw, non-invertible never reversed, per-day stability, meaning fallback). Also fixes a pre-existing unit-job breakage: test_bar_commands and test_lore_commands each stubbed `discord` with different completeness and shared sys.modules, so once both landed on main the one lacking `discord.ext.tasks` shadowed the one needing it and collection failed order-dependently. A new tests/unit/conftest.py stubs discord once, completely, before any test module - the per-file stubs then skip. Full unit job: 41 passed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
3f4a1d5083 |
lore: bound pamiec.json by summarising old memory into "Legendy Baru"
CI / compile (pull_request) Successful in 8s
CI / unit (pull_request) Failing after 10s
CI / integration (pull_request) Successful in 12s
build / build (push) Successful in 37s
CI / compile (push) Successful in 38s
CI / unit (push) Failing after 1m19s
CI / integration (push) Successful in 9s
The conversation memory file grows forever (every chat appends a user+assistant pair), so startup load gets slower and the disk fills. New always-loaded cog lore_commands turns that growth into content: a background task summarises the oldest slice into one in-character "legend" via the ACTIVE AI backend, replaces those old messages with the summary (bounding the file, keeping continuity for the next startup's context), archives it to legendy.json, and announces it on Safety: the compaction transforms (build_transcript, apply_compaction) are pure and unit-tested. The file rewrite is re-read -> back up -> atomic write with no await in between, so a handle_response append that lands while the summary is being generated can neither be lost (it's in the preserved tail) nor corrupt the file (single-threaded, no interleave). A .bak is kept. Scope note: this bounds the on-disk file (startup/disk); the in-RAM MESSAGE_TABLE is a separate concern left untouched to avoid yanking context from a live conversation. Commands: $zapisz_legende (Vykidailo) forces a compaction now; $legendy recalls a random past legend. All thresholds env-overridable (CONJURER_MEMORY_COMPACT_*, CONJURER_LEGENDS_CHANNEL). constants gains LEGENDS_FILE + config + seed; bot.py registers the cog. Verified: tests/unit/test_lore_commands.py covers prefix-replace/tail-keep, preservation of appends made during summarisation, and transcript formatting + head/tail truncation. Unit job 32 passed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e9e731e2bd |
bar: $nalej invents cocktails, $menu keeps the bar's growing lore
CI / compile (pull_request) Successful in 10s
CI / unit (pull_request) Successful in 16s
CI / integration (pull_request) Successful in 11s
build / build (push) Successful in 42s
CI / compile (push) Successful in 9s
CI / unit (push) Successful in 14s
CI / integration (push) Successful in 11s
The most in-character capability the bot has: the persona is literally a 200kg bartender who mixes strong drinks with intriguing names. New always-loaded cog bar_commands: * $nalej [motyw] asks the ACTIVE AI backend (whatever $gadaj_teraz selects) to invent one themed cocktail in Conjurer's voice - persona reused from GPT_SETTINGS[0] as a system message, instructions as the user turn, via handle_response request_type NONE so it never pollutes the bar's conversation memory. Empty motyw = a surprise; "radio"/"pod muzykę" themes the drink on the track currently playing (PREPPED_TRACKS["now_playing"]). * every drink is appended to menu.json (new seeded state file, CONJURER_MENU_FILE overridable) with name/theme/author/timestamp/full text - emergent bar lore. * $menu lists the invented drinks and pours one at random from the archive. Text-only for now; a DALL-E drink image is an easy follow-up (the render path already exists in ai_commands, OpenAI-only). constants gains MENU_FILE (next to pamiec.json by default) + its seed; bot.py registers bar_commands as a core cog. Verified: tests/unit/test_bar_commands.py covers name extraction (markers/markdown/fallback) and the menu round-trip incl. corrupt-file tolerance; unit job 32 passed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
031c1f8aea |
librarian: tolerate bad bytes in chunks; log what the result-send does
CI / compile (pull_request) Successful in 10s
CI / unit (pull_request) Successful in 20s
CI / integration (pull_request) Successful in 15s
CI / compile (push) Successful in 29s
CI / unit (push) Successful in 30s
CI / integration (push) Successful in 21s
build / build (push) Failing after 7s
Two field-reported robustness gaps on top of the hang fix. 1) A stray non-UTF-8 byte in a chunk (0x96 in the report) raised UnicodeDecodeError from readline() - which is a ValueError, so the earlier `except OSError` did NOT catch it. The finally-sentinel meant no hang, but the producer died mid-file with a loud traceback and every DOI after the bad byte went unsearched. Now chunks are opened with errors="replace" (bad bytes become U+FFFD; DOIs are ASCII so a match is never affected) so the read runs to EOF, and the producer's except is broadened from OSError to Exception so no per-file error can ever crash the thread - it's logged and the sentinel still fires. 2) The result-send back to the bot (BackgroundTaskSearch._run) now logs exactly what goes out - target URL, uuid, DOI count and the DOI list - so the librarian log plainly shows a result was sent and what was in it. And a failed POST is no longer fatal: a RequestException used to propagate out of the worker loop and kill the thread, stalling every future query until restart; it's now caught and logged, and a non-200 from the bot is logged as a warning. Verified: tests/unit/test_search_bot.py gains a case writing a chunk with a 0x96 byte before a valid DOI and asserting that DOI is still found (file read to completion, not aborted). All 5 search_bot unit tests pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1a59c9f6c5 |
librarian: stop the DOI search from hanging on a chunk-count mismatch
CI / compile (pull_request) Successful in 1m26s
CI / unit (pull_request) Successful in 1m8s
CI / integration (pull_request) Failing after 10h21m8s
CI / compile (push) Successful in 12m43s
build / build (push) Failing after 13m30s
CI / unit (push) Successful in 2m32s
CI / integration (push) Failing after 1h54m10s
search_bot conflated MAXTHREADS into two jobs at once - how many chunk files to read (files 0..MAXTHREADS-1) AND how many producer sentinels to wait for - so the two had to match exactly. Set too low it silently skipped trailing chunks; set too high (or with any chunk missing/unreadable) a producer crashed before emitting its sentinel, the consumers' count never reached the threshold, and search_for_doi hung on join() forever. The idle-timeout failsafe that was meant to break a starved consumer was dead code: `if empty_counter > 5: ... elif empty_counter > 10: break` - >10 implies >5, so the elif never ran. Fix, three layers: * auto-discover the chunk files present (discover_chunk_files: <n>_chunk.txt in numeric order) instead of range(0, MAXTHREADS). All files are read regardless of count, and no producer is ever pointed at a missing file; * the sentinel threshold is now the number of producers actually started, so it can't drift from what's emitted; * producers emit their sentinel in a finally, so even a crash (missing/unreadable chunk) can't starve the count; and the idle backstop is reordered so it can actually fire (>EMPTY_LIMIT seconds) as a last resort. MAXTHREADS is deprecated and unused (kept only so old env files don't break); docs/env updated to say chunk files are auto-discovered. For the reported case (MAXTHREADS=40, files 0..43): before, files 40-43 were silently never searched, and any run that referenced a missing chunk hung forever. After, all 44 are searched and it always terminates. Verified in a pytest-only venv (tests/unit/test_search_bot.py): DOI in a trailing chunk is found; an unreadable chunk still terminates; empty dir returns at once; discovery is numeric-sorted. Full unit job 27 passed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
3d9d47aa90 |
ai: single-switch GPT/Claude backend for the chat cog
Wire the bot's AI chat pipeline (ai_functions.handle_response) to talk to either OpenAI or the Anthropic Messages API, chosen by one active-config switch. Behaviour on the default "gpt" config is unchanged. constants.py: * guarded `import anthropic` + CLAUDECLIENT (mirrors OPENAICLIENT), netrc machine 'anthropic' / ANTHROPIC_API_KEY; * CLAUDE_LATEST_MODEL / CLAUDE_CHEAP_MODEL (opus-4-8 / haiku-4-5); * AI_CONFIGS + DEFAULT_AI_CONFIG loaded from an optional 3rd element of system_gpt_settings.json (backward compatible - a 2-element file falls back to built-in defaults, active "gpt"). Single switch: CONJURER_AI_CONFIG env > settings "active" > "gpt". ai_functions.py: * provider_generate() dispatches to OpenAI (unchanged openai_call) or the new _anthropic_call() (splits system out, alternating messages, max_tokens, temperature omitted - Opus 4.8 rejects sampling params); * AIError normalises both SDKs' exceptions into one category set so handle_response keeps its single set of in-character error replies; * select_model() reads the active config; legacy "gpt-4o" default auto-maps to the active provider's model so the switch actually changes the backend; * set_active_ai_config()/list_ai_configs() with best-effort persistence back into system_gpt_settings.json index 2. ai_commands.py: * $gadaj_teraz <config> hybrid command (Vykidailo-gated) switches backend at runtime; * graceful guards when OPENAICLIENT is None: personal assistants (OpenAI Assistants API) and DALL-E image gen degrade instead of crashing, so a Claude-only deployment boots. system_gpt_settings.json: add the configs block (gpt/claude/_template) as the collection point for future backends. requirements_bot.txt: add anthropic. bot.env.example: ANTHROPIC_API_KEY + CONJURER_AI_CONFIG. Unit tests cover the message splitter, model selection, config listing, and error mapping. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a6c20a0054 |
ci: replace broken default workflows with compile/unit/integration CI
The two scaffold workflows (Python application / Python package) failed on
every PR: they installed deps from a non-existent requirements.txt, ran
flake8/pytest over the vendored yt_dlp fork (new syntax under the 3.8/3.9
matrix), and collected ad-hoc root scripts — notably test_ai.py, which is
an invalid pasted object dump (not Python).
- Remove python-app.yml / python-package.yml and the junk root scripts
(test.py, test_ai.py, test_time.py)
- Add .github/workflows/ci.yml with three PR-check jobs:
* compile — py_compile every first-party .py (no deps)
* unit — pytest on pure logic (conanjurer_functions, constants)
* integration — boot the Flask services and assert the X-Conjurer-Api-Key
auth contract (communication_subroutine + conjurer_musician)
- Add tests/ suite, pytest.ini (testpaths=tests) and conftest.py (sys.path)
Fixes surfaced by the compile gate / needed for the integration job:
- conjurer_librarian/search_bot.py + search_bot2.py: f-string reused the
same quote ({item["exists"]}) -> SyntaxError on Python < 3.12
- conjurer_musician/media_search_functions.py: made import-safe
(env-overridable paths, lazy DB load / mkdir) so the service can be
imported and tested off the Pi
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|