conjurer: pin the model Ollama actually has, and raise the AI timeout
Probing the real server (192.168.1.72:11434) after it was opened to the LAN: /v1/models returns exactly one model, gemma4:e2b. Without pinning it the bot would request the built-in default llama3.1:8b and every reply would fail with "model not found", so set it explicitly on both bots. Also raise CONJURER_AI_TIMEOUT_SECONDS to 240. Self-hosted generation is far slower than a hosted API, particularly the first request after the model is evicted from VRAM. It applies to every backend, so it is deliberately not set higher than needed. Caveat recorded honestly: at the time of writing, generation on that server does not complete. /v1/models answers instantly, but both /v1/chat/ completions (180s) and native /api/generate with num_predict=5 (60s) return nothing, and /api/ps shows no model ever becomes resident - so the model never finishes loading. Ollama is 0.32.14 and gemma4:e2b is 5.1B Q4_K_M (~3.5GB) despite the "e2b" name. That is a server-side problem, not a configuration one; these values are correct and take effect once it loads. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit was merged in pull request #23.
This commit is contained in:
@@ -36,6 +36,14 @@ spec:
|
||||
# Pick a model at runtime with: $gadaj_teraz ollama <model>
|
||||
# ($modele_ai lists what the server actually has pulled).
|
||||
- { name: CONJURER_OLLAMA_URL, value: "http://192.168.1.72:11434" }
|
||||
# The server currently has exactly one model pulled (verified via
|
||||
# /v1/models): gemma4:e2b. Without this the built-in default
|
||||
# (llama3.1:8b) would be requested and every reply would fail.
|
||||
- { name: CONJURER_OLLAMA_MODEL, value: "gemma4:e2b" }
|
||||
# Self-hosted generation is far slower than a hosted API, especially
|
||||
# the first request after the model is evicted from VRAM. Applies to
|
||||
# every backend, so keep it only as high as you actually need.
|
||||
- { name: CONJURER_AI_TIMEOUT_SECONDS, value: "240" }
|
||||
volumeMounts:
|
||||
- { name: data, mountPath: /data }
|
||||
- { name: netrc, mountPath: /secrets, readOnly: true }
|
||||
|
||||
@@ -56,6 +56,14 @@ spec:
|
||||
# Pick a model at runtime with: $gadaj_teraz ollama <model>
|
||||
# ($modele_ai lists what the server actually has pulled).
|
||||
- { name: CONJURER_OLLAMA_URL, value: "http://192.168.1.72:11434" }
|
||||
# The server currently has exactly one model pulled (verified via
|
||||
# /v1/models): gemma4:e2b. Without this the built-in default
|
||||
# (llama3.1:8b) would be requested and every reply would fail.
|
||||
- { name: CONJURER_OLLAMA_MODEL, value: "gemma4:e2b" }
|
||||
# Self-hosted generation is far slower than a hosted API, especially
|
||||
# the first request after the model is evicted from VRAM. Applies to
|
||||
# every backend, so keep it only as high as you actually need.
|
||||
- { name: CONJURER_AI_TIMEOUT_SECONDS, value: "240" }
|
||||
volumeMounts:
|
||||
- { name: data, mountPath: /data }
|
||||
- { name: netrc, mountPath: /secrets, readOnly: true }
|
||||
|
||||
Reference in New Issue
Block a user