PROTOCOL 05 / AGENCY

The First Agent, Tools and Safe Execution

1–2 sessions Foundation Eden research series

Learning outcomes

  • Differentiate agent, model and tool
  • Build an autonomous loop
  • Register tools explicitly
  • Validate arguments and permission
  • Return a standard ToolResult
  • Reconsider after failure

1. Agent, model and tool are not synonyms

An agent owns a goal and repeatedly observes, decides, acts and evaluates. A model proposes a decision from context. A tool is a controlled capability implemented by Python. The model may ask to use a tool; it never becomes that tool.

01PerceiveBuild observation
02DecideModel proposal
03ValidateSchema + permission
04ExecutePython handler
05EvaluateRead ToolResult

2. Design the smallest useful tool catalogue

Begin with observe, move and wait. Later chapters add drink, eat and talk. A tool has two layers: a public schema available to the model and a private Python implementation that owns the real operation.

from typing import Literal from pydantic import BaseModel, ConfigDict class MoveArgs(BaseModel): """One orthogonal grid step.""" model_config = ConfigDict(extra="forbid") dx: Literal[-1, 0, 1] dy: Literal[-1, 0, 1] def valid_orthogonal_step(dx: int, dy: int) -> bool: """Reject diagonal movement and no-op movement.""" return abs(dx) + abs(dy) == 1 assert valid_orthogonal_step(1, 0) assert not valid_orthogonal_step(1, 1)

Type validation alone does not reject (1, 1) or (0, 0). Domain rules complete the validation.

3. Build a controlled execution pipeline

  1. Confirm that the named tool is registered.
  2. Confirm that the requesting agent is authorised.
  3. Validate arguments against the schema.
  4. Confirm that the referenced world revision is still current.
  5. Ask the world whether the action is possible.
  6. Execute the handler once and record the resulting event.
SECURITY BOUNDARY

Never run arbitrary model-generated code and never use unrestricted getattr(world, model_selected_name). An explicit registry is an allowlist.

4. Return one standard result

ToolResult should include success, tool name, request ID, agent ID, tick, reason and bounded result data. A rejected action is not an infrastructure crash; it is evidence the agent can use for a new decision.

TOOL_REGISTRY = { "move": move_handler, "wait": wait_handler, } def execute_tool(call: ToolCall, context: ToolContext) -> ToolResult: """Validate every boundary before dispatching a handler.""" handler = TOOL_REGISTRY.get(call.name) if handler is None: return ToolResult.failure(call, reason="unknown_tool") if not context.permissions.allows(call.agent_id, call.name): return ToolResult.failure(call, reason="not_authorised") return handler(call, context)

5. Tool calling through LiteLLM

Some compatible models can return native tool_calls; others can return a JSON object containing a name and arguments. In both cases our execution pipeline remains the same. Verify the capability of the exact local model instead of assuming that support in LiteLLM guarantees support in Ollama.

6. Guided laboratory

EXPERIMENT A

Manual call

Execute one typed move without a model.

EXPERIMENT B

Registry and permission

Reject unknown tools and an unauthorised agent.

EXPERIMENT C

Model proposal

Parse a real tool call but keep execution separate.

EXPERIMENT D

Replanning

Return a wall collision and let the agent choose again.

Protocol completion

A valid tool may still be rejected by current world state. Every rejection leaves state unchanged, returns an explicit reason and can be correlated with the decision that caused it.