Compare commits

..

2 Commits

Author SHA1 Message Date
gitea a802b90617 conjurer: pin the model Ollama actually has, and raise the AI timeout
Probing the real server (192.168.1.72:11434) after it was opened to the LAN:
/v1/models returns exactly one model, gemma4:e2b. Without pinning it the
bot would request the built-in default llama3.1:8b and every reply would
fail with "model not found", so set it explicitly on both bots.

Also raise CONJURER_AI_TIMEOUT_SECONDS to 240. Self-hosted generation is far
slower than a hosted API, particularly the first request after the model is
evicted from VRAM. It applies to every backend, so it is deliberately not
set higher than needed.

Caveat recorded honestly: at the time of writing, generation on that server
does not complete. /v1/models answers instantly, but both /v1/chat/
completions (180s) and native /api/generate with num_predict=5 (60s) return
nothing, and /api/ps shows no model ever becomes resident - so the model
never finishes loading. Ollama is 0.32.14 and gemma4:e2b is 5.1B Q4_K_M
(~3.5GB) despite the "e2b" name. That is a server-side problem, not a
configuration one; these values are correct and take effect once it loads.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-24 14:54:32 +02:00
gitea aeb9bb5940 conjurer: point both bots at the Ollama server (192.168.1.72:11434)
Wires CONJURER_OLLAMA_URL into the test bot and the deploy bot so the
self-hosted backend from conjurer#25 is selectable. No API key exists for
Ollama - the endpoint is the whole configuration - and until it is set the
backend refuses to be selected, so this is what turns it on.

NOTE, verified from the LAN before committing: 192.168.1.72 answers ping
(0.4ms) but only port 22 is open - 11434 refuses. Ollama binds to
127.0.0.1:11434 by default, so it is not reachable off-host yet. This env
var is correct but inert until the server listens on the network:

  sudo systemctl edit ollama.service
    [Service]
    Environment="OLLAMA_HOST=0.0.0.0:11434"
  sudo systemctl daemon-reload && sudo systemctl restart ollama

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-24 14:43:01 +02:00
3 changed files with 35 additions and 37 deletions
+9 -37
View File
@@ -100,43 +100,15 @@ unset APP_PASSWORD
wstanie**: usługa z kontami, ale bez klucza, nie odróżniłaby ważnej sesji od
podrobionej. Nikt go nigdy nie musi oglądać.
### ⚠️ Masz już sekret sprzed LOG-34? Dołóż klucz, nie twórz od nowa
Komenda wyżej zakłada sekret **od zera**. Jeśli `astrololo-auth` już istnieje,
merge manifestu sam klucza nie dołoży — pod zgłosi wtedy:
```
Error: couldn't find key SESSION_SECRET in Secret astrololo/astrololo-auth
```
Dokładamy klucz, nie ruszając pozostałych:
```bash
kubectl -n astrololo patch secret astrololo-auth --type=merge \
-p "{\"stringData\":{\"SESSION_SECRET\":\"$(openssl rand -hex 32)\"}}"
kubectl -n astrololo rollout restart deploy/presentation
```
Sprawdzenie, że komplet kluczy jest na miejscu (bez pokazywania wartości):
```bash
kubectl -n astrololo get secret astrololo-auth -o jsonpath='{.data}' \
| tr ',' '\n' | grep -o '"[A-Z_]*"'
```
Oczekiwane: `APP_PASSWORD`, `INTERNAL_TOKEN`, `SESSION_SECRET`.
> `--type=merge` ze `stringData` **dokłada albo nadpisuje** i nie wymaga, żeby
> klucz wcześniej istniał — w odróżnieniu od JSON Patch z `op: replace`, który
> na brakującej ścieżce po prostu odmawia. Ta sama komenda służy więc i do
> dołożenia, i do rotacji.
### Rotacja klucza sesji
**Wylogowuje WSZYSTKICH.** To nie usterka, tylko awaryjny wyłącznik: gdy
podejrzewasz, że ktoś przechwycił cudzą sesję, podmiana klucza unieważnia je
wszystkie naraz. Komenda ta sama, co dołożenie wyżej.
> **Rotacja tego klucza wylogowuje WSZYSTKICH.** To nie usterka, tylko awaryjny
> wyłącznik: gdy podejrzewasz, że ktoś przechwycił cudzą sesję, podmiana klucza
> unieważnia je wszystkie naraz.
>
> ```bash
> kubectl -n astrololo patch secret astrololo-auth --type=json \
> -p="[{\"op\":\"replace\",\"path\":\"/data/SESSION_SECRET\",\"value\":\"$(openssl rand -hex 32 | base64 | tr -d '\n')\"}]"
> kubectl -n astrololo rollout restart deploy/presentation
> ```
`INTERNAL_TOKEN` jest losowany i **nikt go nigdy nie musi oglądać** — służy tylko
usługom do rozmowy między sobą. `APP_PASSWORD` wpisujesz w przeglądarce
+13
View File
@@ -31,6 +31,19 @@ spec:
# Where the librarian sends THIS bot's results/pongs back to (its own
# NodePort). Lets one librarian serve both bots - see deploy-bot.yaml.
- { name: CONJURER_SELF_CALLBACK, value: "http://192.168.1.73:32442" }
# Self-hosted models (Ollama). No API key - the endpoint IS the
# configuration, and the backend stays unselectable while unset.
# Pick a model at runtime with: $gadaj_teraz ollama <model>
# ($modele_ai lists what the server actually has pulled).
- { name: CONJURER_OLLAMA_URL, value: "http://192.168.1.72:11434" }
# The server currently has exactly one model pulled (verified via
# /v1/models): gemma4:e2b. Without this the built-in default
# (llama3.1:8b) would be requested and every reply would fail.
- { name: CONJURER_OLLAMA_MODEL, value: "gemma4:e2b" }
# Self-hosted generation is far slower than a hosted API, especially
# the first request after the model is evicted from VRAM. Applies to
# every backend, so keep it only as high as you actually need.
- { name: CONJURER_AI_TIMEOUT_SECONDS, value: "240" }
volumeMounts:
- { name: data, mountPath: /data }
- { name: netrc, mountPath: /secrets, readOnly: true }
+13
View File
@@ -51,6 +51,19 @@ spec:
# query and answers results/pongs HERE - so it serves this bot AND the
# test bot from one instance, no CONJURER_MAIN_BOT repointing needed.
- { name: CONJURER_SELF_CALLBACK, value: "http://192.168.1.73:32443" }
# Self-hosted models (Ollama). No API key - the endpoint IS the
# configuration, and the backend stays unselectable while unset.
# Pick a model at runtime with: $gadaj_teraz ollama <model>
# ($modele_ai lists what the server actually has pulled).
- { name: CONJURER_OLLAMA_URL, value: "http://192.168.1.72:11434" }
# The server currently has exactly one model pulled (verified via
# /v1/models): gemma4:e2b. Without this the built-in default
# (llama3.1:8b) would be requested and every reply would fail.
- { name: CONJURER_OLLAMA_MODEL, value: "gemma4:e2b" }
# Self-hosted generation is far slower than a hosted API, especially
# the first request after the model is evicted from VRAM. Applies to
# every backend, so keep it only as high as you actually need.
- { name: CONJURER_AI_TIMEOUT_SECONDS, value: "240" }
volumeMounts:
- { name: data, mountPath: /data }
- { name: netrc, mountPath: /secrets, readOnly: true }