PROTOCOL 04 / INTELLIGENCE

Local AI with Ollama and LiteLLM

1–2 sessions Foundation Eden research series

Learning outcomes

  • Define each AI component
  • Connect to local Ollama
  • Call a model through LiteLLM
  • Validate structured decisions
  • Handle timeout and malformed output
  • Keep AI outside the domain

1. Assign the correct role to every component

Ollama serves the local language model over HTTP. LiteLLM gives the Python application a relatively consistent provider interface. Our adapter transforms a garden observation into model messages and receives a response. A validator converts that response into a known Decision. The simulation engine remains the only component allowed to apply a world change.

WORLDObservationAuthorised facts
ADAPTERLiteLLM requestPrompt + configuration
OLLAMAModel responseUntrusted proposal
VALIDATORDecision contractKnown structure
ENGINEAccept or rejectWorld authority

LiteLLM is not the agent, memory or orchestrator. Ollama does not create autonomy merely by running a model. Those behaviours belong to our application.

2. Prepare the local service

First decide where Python and Ollama run. localhost always means the current machine or container. When both run together, the default API is usually http://localhost:11434. Across two devices, use a trusted private network and protect the service rather than exposing an unauthenticated endpoint to the Internet.

# On the Ollama host ollama list ollama pull qwen2.5:3b # From the machine running Python curl http://localhost:11434/api/tags # In the project virtual environment python -m pip install litellm pydantic

3. Make the first isolated request

from litellm import completion response = completion( model="ollama_chat/qwen2.5:3b", api_base="http://localhost:11434", messages=[ {"role": "system", "content": "You are Adam in a virtual garden."}, {"role": "user", "content": "You are thirsty. What would you seek?"}, ], temperature=0.2, timeout=30, ) print(response.choices[0].message.content)

Run this in the console before involving Tkinter. It proves network, model name and adapter configuration separately from UI concurrency.

4. Design a bounded observation

Do not dump the whole world into every prompt. Provide the agent identity, tick, visible resources, current needs, current goal and permitted actions. The observation is a filtered view created by Python, not arbitrary access to internal state.

OBSERVATION CONTRACT

Example: Adam, tick 12, position (4, 4), thirst 8/10, visible water at (7, 2), allowed actions: move, wait. The model needs enough evidence to propose; it does not need the entire database.

5. Parse a structured decision

from typing import Literal from pydantic import BaseModel, ConfigDict, Field class Decision(BaseModel): """A model proposal; validation does not execute it.""" model_config = ConfigDict(extra="forbid") goal: str = Field(min_length=1, max_length=120) action: Literal["move", "wait"] arguments: dict reason: str = Field(min_length=1, max_length=180)

Schema validation proves shape, not truth. A valid move request can still target a wall, refer to stale state or exceed the agent's permissions. The engine will decide later.

6. Async, failure and observability

Use LiteLLM's async interface inside the established worker. Apply timeouts and return a typed failure instead of silently inventing a decision. Record model name, request ID, latency, outcome and error class; never log credentials or unnecessary private content.

EXPERIMENT A

Verify Ollama

Confirm the exact model and HTTP endpoint independently.

EXPERIMENT B

Validate output

Parse one good JSON response and reject extra fields.

EXPERIMENT C

Test async

Return the decision without blocking Tkinter.

EXPERIMENT D

Break it safely

Stop Ollama, force a timeout and send malformed JSON.

Protocol completion

You can change the configured model without changing the agent or world code. An unavailable model produces a visible, bounded error and never causes an action to execute.