Learning outcomes
- Form a testable hypothesis
- Control experimental variables
- Use deterministic seeds
- Choose meaningful metrics
- Compare against a baseline
- Interpret without overclaiming
1. Turn a working program into a laboratory
The system can now represent a world, advance ticks, consult Ollama, execute tools, retain memories and deliver messages. The next goal is not simply to watch Adam move; it is to ask a question that evidence can answer.
This artificial-life simulation does not create living beings or prove consciousness. It creates observable behaviour from rules, goals, memory and probabilistic decisions.
2. Define scenarios with consequences
Start with 10 × 10 before scaling to 60 × 60. Water, food and obstacles create environmental pressure. Thirst, hunger and energy change through deterministic rules. Movement, eating and drinking need explicit costs; free actions make strategy difficult to compare.
Useful scenarios include known versus unknown water, disappearing resources, two agents competing for food and Eve receiving incorrect information.
3. Be precise about autonomy and emergence
An autonomous agent observes, decides when deliberation is needed, uses a tool and evaluates the result without continuous user commands. Autonomy still operates within strict permissions.
An emergent pattern is not directly scripted as a sequence—for example, agents sharing water locations and consequently exploring less. One striking run is not evidence of general learning.
Following rules, adapting with a memory and learning a general strategy are different claims. Each requires stronger evidence than the previous one.
4. Design one controlled experiment
Every experiment records a question, hypothesis, changed variable, controlled conditions, baseline, metrics, number of runs, termination rule and known limitations.
Example hypothesis: enabling memory retrieval reduces ticks to first drink. Change memory only. Hold map, positions, rules, action budget, model and generation settings constant.
5. Reproducibility and random seeds
import random
def resource_positions(seed: int, count: int = 4) -> list[tuple[int, int]]:
"""Create a reproducible deterministic resource layout."""
generator = random.Random(seed)
return [
(generator.randrange(10), generator.randrange(10))
for _ in range(count)
]
assert resource_positions(42) == resource_positions(42)A seed makes the deterministic part reproducible. It does not guarantee identical LLM responses: model version, temperature, concurrency and inference implementation also matter.
6. Measure both behaviour and infrastructure
World and agent metrics
- Ticks and steps until first drink.
- Percentage of runs that reach the goal.
- Invalid or rejected tool calls.
- Useful memories retrieved and messages acted upon.
- Explored cells, energy consumed and resources shared.
Technical metrics
- Model calls, latency and timeout rate.
- Tokens or estimated local inference cost.
- Queue depth, database duration and UI responsiveness.
- Failures grouped by component and error code.
7. Compare with a baseline
A deterministic RuleBasedAgent is valuable: it shows whether the LLM actually improves the target outcome. Compare no memory versus memory, no communication versus communication and scripted policy versus model proposal. Change one variable at a time.
8. Guided research programme
Controlled scenario
Fix map, needs, action budget and termination tick.
Rule baseline
Record success without an LLM.
Memory A/B
Run multiple seeds with memory off and on.
Message quality
Compare correct, stale and false reports.
Research completion
You can state exactly what changed, reproduce the deterministic conditions, compare several runs against a baseline and separate measured results from interpretation. The next experiment should emerge from evidence, not from adding complexity for its own sake.