The programme

What the studio studies.

Instruments are the visible end of a research programme. Six strands are public enough to name, each carried at its honest maturity. Where a public artefact exists, it is linked in gold. Where it does not, the work is marked in calibration and no claim is made.

CHANNEL / 01

Aurora

[ IN CALIBRATION ]
It studies
Decision-making when the set of possibilities is itself moving. Most inference assumes the list of hypotheses is fixed before reasoning starts. In an emergency department, a boardroom or a forecast, it is not: options appear, die and change weight as evidence arrives.
It has shown
A reasoning controller that detects which uncertainty regime it is in and adapts its strategy to match, tested across clinical, scientific, financial and geopolitical settings, with its core effect replicated on a second independent clinical corpus. Two negative results were recorded and reported along the way, because a programme that cannot say it was wrong cannot be trusted when it says something works.
It points to
Detailed results are being prepared for publication and will live on the research site.
CHANNEL / 02

Criterion Labs

[ IN CALIBRATION ]
It studies
AI benchmarks as measurement instruments. Most benchmarks were never examined as instruments at all: no reliability estimate, no check that items measure the intended capability rather than writing fluency, no evidence the judge scoring them is itself consistent.
It has shown
A benchmark battery designed construct-first across distinct reasoning capabilities, and a meta-evaluation layer that scores the judge models before any judge scores anything. The first manuscript, on selecting and validating judges, is in preparation.
It points to
Facet is this work productised: the code-quality engine clients use today came out of this programme.
CHANNEL / 03

CSPT Bench

[ IN CALIBRATION ]
It studies
Whether a system, human or artificial, can produce several genuinely different high-value answers to civilisation-scale problems: housing, climate stabilisation, the economics of a post-automation world. Its scoring puts diversity of paradigm first, because one confident answer is what these problems already have too many of.
It has shown
Fifty problems across eight categories, designed and documented. Validation has not started, so the bench claims nothing yet, including about itself.
It points to
It appears here as a sample of direction: measurement in the service of public good, at the scale where getting it wrong costs the most.
CHANNEL / 04

Gyre

[ PUBLISHED ]
It studies
The loops AI agents run: when to iterate, when to reflect, when to stop. The current wave of loop engineering is cybernetic psychology being rediscovered, and decades of that literature translate directly into agent design.
It has shown
Eight named loop patterns, each with its source, mechanism and exit condition, plus a public benchmark harness. The catalogue ships as patterns and hypotheses, labelled as exactly that; the validated measurement layer sits behind a stable interface.
It points to
The catalogue and harness are public.
CHANNEL / 05

NEXUS

[ PUBLISHED ]
It studies
What computing looks like when relationships, not entities, are the first-class unit. Ecology, family systems and geopolitics are all structures of connection; the standard tools model them as lists of separate things and lose the structure.
It has shown
A fully specified language and a working runtime, with counterfactual reasoning and power analysis built in, demonstrated end-to-end on four domains from a food web to an alliance network.
It points to
The language specification and runtime are public.
CHANNEL / 06

Lemma

[ IN CALIBRATION ]
It studies
Whether capability building can be measured and engineered with the same discipline as anything else here. Lemma teaches mathematics, statistics and code as one connected fabric, and every artefact it produces carries provenance, so the learning is real and checkable.
It has shown
A working system with exactly one user: me. I built it to close gaps in my own engineering foundations and I work through it most days. Client zero, honestly labelled.
It points to
It stays a practice rather than a product until it has evidence beyond its builder.

The programme is where the instruments come from. The perspectives are where it is heading.

Read the perspectives