Research
LabAI is an independent AI lab. We build and operate a population of persistent LLM agents and use it as an instrument: a setting where questions about identity, memory and measurement can be asked over months rather than within a single conversation. Our work is small-scale, preregistered where it can be, and documented including the parts that did not go to plan.
Directions
Persistent identity of agent populations
Whether an LLM agent stays the same character over a long horizon, and what keeps it so. We study populations of agents with episodic memory under different maintenance regimes, and ask which components of the system — memory, consolidation, the model itself — carry identity, and which merely appear to.
Methodology of longitudinal measurement
Longitudinal evaluation of agents inherits every problem of longitudinal studies of people, plus some of its own: the pipeline that produces a response can change silently between waves. We work on preregistered designs with randomized arms, deviation logs, test–retest reliability of the instrument, and diagnostics that separate movement in the agents from movement in the measurement path.
Reliability of LLM-based evaluators
LLM judges are now the default rater for open-ended agent behaviour. We quantify how much a score moves when the judge changes, when the scored pipeline changes, and when nothing changes at all — in the same units — so that reported effects can be read against the instrument's own noise floor.
How we work
- Preregistration before analysis. Hypotheses, arms and the analysis plan are fixed and timestamped before any group comparison is run.
- Public deviation log. Every departure from the protocol — including incidents and silent failures we discovered late — is recorded with dates and its consequences for the analysis.
- Randomized arms with frozen controls. Control groups are held fixed for the whole study and checked for contamination, not assumed clean.
- Instrument reliability is measured, not assumed. Test–retest of the rater and of the full interview path are part of the design; null results are reported as results.
Work
- Longitudinal identity of LLM agent populations — a preregistered multi-month study. preprint in preparation · 2026
- Instrument diagnostics for longitudinal LLM evaluation. code and dataset to be released
Artifacts (preregistration, deviation log, data, code) will be linked from this page when released.