PROTOCOL 08 / QUALITY

Contracts, Standards and Reliability

1–2 sessions Foundation Eden research series

Learning outcomes

  • Define module boundaries
  • Version data contracts
  • Correlate a full decision
  • Classify errors and retries
  • Test without Ollama
  • Design idempotent effects

1. Standards remove invisible assumptions

The garden now observes, consults a model, uses tools, stores memories and transports messages. Without stable contracts, one module eventually returns a slightly different dictionary and failures become difficult to locate. Standardisation means agreeing on inputs, outputs and failure rules; it does not mean building a large framework.

2. Keep contracts narrow

The central contracts are Observation, Decision, ToolCall, ToolResult, WorldEvent, AgentMessage and MemoryItem. Each expresses one responsibility. Avoid a universal dictionary that contains every possible field.

OBSERVATIONPermitted factsAgent's view
DECISIONValidated intentModel proposal
TOOL CALLNamed requestID + arguments
TOOL RESULTAccepted or rejectedExplicit outcome
WORLD EVENTConfirmed effectEngine authority

3. Version structure and world state separately

schema_version changes when a data format changes. world_revision changes when the world mutates and lets the engine reject a decision based on stale information. A tick represents simulated time but may not uniquely identify a mutation.

from typing import Any, Literal from pydantic import BaseModel, ConfigDict, Field class WorldEvent(BaseModel): """An immutable event created by the simulation engine.""" model_config = ConfigDict(extra="forbid", frozen=True) schema_version: Literal[1] = 1 event_id: str = Field(min_length=1) request_id: str = Field(min_length=1) actor_id: str = Field(min_length=1) tick: int = Field(ge=0) kind: Literal["MOVE_OK", "MOVE_REJECTED"] payload: dict[str, Any]

4. Trace one decision end to end

A request_id relates observation, model call, decision, tool request and result. The confirmed world change receives its own event_id. Structured logs should capture timestamp, tick, agent, request, component, operation, outcome, duration and error code.

OBSERVABILITY LIMIT

Record evidence needed to investigate behaviour, but never secrets. Large private reasoning traces are not required to prove that a tool was called and validated.

5. Timeouts and retries are policy decisions

Separate invalid input, rejected action, model timeout, network failure, database error and internal bug. Not every failure should be retried. A malformed contract needs correction; a temporary connection may justify bounded exponential backoff.

A retry must not repeat a confirmed side effect. Store or recognise request IDs so move(req-123) cannot move twice after a lost response.

6. Test the agent without a live model

class FakeModel: """Return deterministic decisions for integration tests.""" async def decide(self, observation: Observation) -> Decision: return Decision( goal="Reach visible water", action="move", arguments={"dx": 1, "dy": 0}, reason="Water is east", )

A fake proves orchestration, tool validation and event creation without network latency or probabilistic output. Keep a smaller set of live-model tests for adapter compatibility.

7. LiteLLM and future MCP integration

LiteLLM standardises model access. It does not replace internal contracts or permissions. MCP may later expose tools across process boundaries, but the internal tool service should remain valid without it. Add an integration only when it solves a concrete boundary.

EXPERIMENT A

Freeze a contract

Reject unknown fields and accidental mutation.

EXPERIMENT B

Correlate IDs

Follow one request through every layer.

EXPERIMENT C

Use FakeModel

Test the whole loop without Ollama.

EXPERIMENT D

Retry safely

Repeat a request ID and prove one effect.

Protocol completion

You can replace Ollama with FakeModel without changing world rules, diagnose one failed decision from structured evidence and retry transient work without duplicating a confirmed action.