a802b90617
Probing the real server (192.168.1.72:11434) after it was opened to the LAN: /v1/models returns exactly one model, gemma4:e2b. Without pinning it the bot would request the built-in default llama3.1:8b and every reply would fail with "model not found", so set it explicitly on both bots. Also raise CONJURER_AI_TIMEOUT_SECONDS to 240. Self-hosted generation is far slower than a hosted API, particularly the first request after the model is evicted from VRAM. It applies to every backend, so it is deliberately not set higher than needed. Caveat recorded honestly: at the time of writing, generation on that server does not complete. /v1/models answers instantly, but both /v1/chat/ completions (180s) and native /api/generate with num_predict=5 (60s) return nothing, and /api/ps shows no model ever becomes resident - so the model never finishes loading. Ollama is 0.32.14 and gemma4:e2b is 5.1B Q4_K_M (~3.5GB) despite the "e2b" name. That is a server-side problem, not a configuration one; these values are correct and take effect once it loads. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
94 lines
5.5 KiB
YAML
94 lines
5.5 KiB
YAML
# Production ("deploy") Conjurer bot.
|
|
#
|
|
# Same infrastructure as the test `bot` (shares librarian / musician / radio and
|
|
# the CONJURER_API_KEY), with three deliberate differences:
|
|
# 1. its Discord token comes from the `deploy-conjurer-netrc` secret,
|
|
# 2. distinct names/labels so it coexists with the test bot in this namespace,
|
|
# 3. /data is an NFS volume (RWX) instead of a block PVC - so config/state can
|
|
# be uploaded while the bot runs (just copy onto the share) and backed up
|
|
# concurrently (see deploy-bot-backup.yaml). The test bot keeps its RWO PVC.
|
|
#
|
|
# ┌─ IMPORTANT: librarian & musician call ONE bot back ─────────────────────────┐
|
|
# │ The librarian (CONJURER_MAIN_BOT) and musician push search results and │
|
|
# │ "now playing" events to a SINGLE bot address - currently the test bot's │
|
|
# │ NodePort 192.168.1.73:32442. Only one bot can receive them. To make the │
|
|
# │ PRODUCTION bot the one that gets librarian results / radio events, repoint │
|
|
# │ those services' CONJURER_MAIN_BOT to this bot's NodePort (…:32443 below). │
|
|
# └─────────────────────────────────────────────────────────────────────────────┘
|
|
apiVersion: apps/v1
|
|
kind: Deployment
|
|
metadata:
|
|
name: deploy-bot
|
|
namespace: conjurer
|
|
spec:
|
|
replicas: 1 # NIGDY więcej — jedna sesja gateway na token
|
|
strategy: { type: Recreate } # NIE RollingUpdate — dwa pody = wojna o sesję Discord
|
|
selector: { matchLabels: { app: deploy-bot } }
|
|
template:
|
|
metadata: { labels: { app: deploy-bot } }
|
|
spec:
|
|
imagePullSecrets: [{ name: gitea-registry }]
|
|
containers:
|
|
- name: bot
|
|
# SEPARATE image from the test bot: conjurer-bot-deploy only gets a new
|
|
# tag when a commit message contains [deploy] (see the conjurer CI), so
|
|
# this bot updates only on versions you explicitly promote. Tag is
|
|
# managed by the image-updater (kustomization images:).
|
|
image: gitea.czernobog.pl/gitea/conjurer-bot-deploy:c8aae106
|
|
ports: [{ containerPort: 5000 }]
|
|
env:
|
|
- { name: CONJURER_DATA_DIR, value: "/data" }
|
|
- { name: CONJURER_DISCORD_HOST, value: "0.0.0.0" }
|
|
- { name: CONJURER_DISCORD_PORT, value: "5000" }
|
|
- { name: CONJURER_API_KEY, value: "d97008b3-7a5a-11f1-acd1-000b0e0f00ed" }
|
|
- { name: CONJURER_NETRC_FILE, value: "/secrets/.netrc" }
|
|
- { name: CONJURER_LIBRARIAN_SERVICE, value: "http://librarian:5001" }
|
|
- { name: CONJURER_MUSICIAN_SERVICE, value: "http://192.168.1.89:5000" }
|
|
- { name: CONJURER_RADIO_SERVICE, value: "http://192.168.1.79:5005" }
|
|
- { name: CONJURER_FILE_SERVICE, value: "http://192.168.1.89:5000" }
|
|
- { name: CONJURER_RADIO_HARBOR, value: "http://192.168.1.79:54321" }
|
|
# This bot's own callback (its NodePort). The librarian records it per
|
|
# query and answers results/pongs HERE - so it serves this bot AND the
|
|
# test bot from one instance, no CONJURER_MAIN_BOT repointing needed.
|
|
- { name: CONJURER_SELF_CALLBACK, value: "http://192.168.1.73:32443" }
|
|
# Self-hosted models (Ollama). No API key - the endpoint IS the
|
|
# configuration, and the backend stays unselectable while unset.
|
|
# Pick a model at runtime with: $gadaj_teraz ollama <model>
|
|
# ($modele_ai lists what the server actually has pulled).
|
|
- { name: CONJURER_OLLAMA_URL, value: "http://192.168.1.72:11434" }
|
|
# The server currently has exactly one model pulled (verified via
|
|
# /v1/models): gemma4:e2b. Without this the built-in default
|
|
# (llama3.1:8b) would be requested and every reply would fail.
|
|
- { name: CONJURER_OLLAMA_MODEL, value: "gemma4:e2b" }
|
|
# Self-hosted generation is far slower than a hosted API, especially
|
|
# the first request after the model is evicted from VRAM. Applies to
|
|
# every backend, so keep it only as high as you actually need.
|
|
- { name: CONJURER_AI_TIMEOUT_SECONDS, value: "240" }
|
|
volumeMounts:
|
|
- { name: data, mountPath: /data }
|
|
- { name: netrc, mountPath: /secrets, readOnly: true }
|
|
resources:
|
|
requests: { cpu: "200m", memory: "256Mi" }
|
|
limits: { cpu: "2", memory: "1Gi" }
|
|
volumes:
|
|
# /data on NFS (RWX): lets you seed/upload files while the bot runs (copy
|
|
# straight onto the export from the PVE box that holds /srv/data) and lets
|
|
# the backup CronJob read it concurrently. ADJUST server/path to your
|
|
# actual export - mirrors the librarian's 192.168.1.34 Tank1 layout.
|
|
- name: data
|
|
nfs:
|
|
server: 192.168.1.34
|
|
path: /mnt/Tank1/conjurer_swap/deploy-bot/data
|
|
- name: netrc
|
|
secret: { secretName: deploy-conjurer-netrc }
|
|
---
|
|
apiVersion: v1
|
|
kind: Service
|
|
metadata: { name: deploy-bot, namespace: conjurer }
|
|
spec:
|
|
type: NodePort # so librarian/musician COULD reach it if repointed
|
|
selector: { app: deploy-bot }
|
|
# Distinct pinned nodePort (test bot owns 32442). Point librarian/musician
|
|
# CONJURER_MAIN_BOT at 192.168.1.73:32443 to route their callbacks here.
|
|
ports: [{ port: 5000, targetPort: 5000, nodePort: 32443 }]
|