Perspectives

What the studio believes.

Six theses on measurement, uncertainty and organisations in the AI transition. They are arguments and informed bets rather than findings, and they are written down so you can hold the studio to them.

Decisions inherit the quality of their instruments.

Organisations are moving to evidence-based decisions faster than they are validating the evidence. Dashboards, engagement scores, performance ratings, AI benchmarks: each is a measurement instrument, and most were never tested as one.

When a flawed instrument feeds a decision, the flaw does not stay in the spreadsheet. Hiring drifts, strategy tilts, and the error hardens into organisational structure. The studio calls this measurement debt · the accumulating cost of decisions made on numbers whose instruments were never validated.

Psychology spent a century learning how to ask whether an instrument measures what it claims: reliability, validity, stated limits. That discipline transfers, and it is the foundation everything here stands on.

Consequence · every Firasa instrument publishes what was tested, what was not, and where its limits sit.

Uncertainty is navigated, never abolished.

Everything that survives uncertainty does so by reading it and moving well inside it. Every mind and every organisation faces the same standing trade: keep exploring for better options, or commit to what is already found. Move too early and you lock in a mediocre answer. Move too late and the window closes.

Most decision processes handle this badly: single-point forecasts, confident plans, and the quiet assumption that the list of possibilities was complete when the planning started. It never is. Options appear, die and change weight while you deliberate.

The studio treats the moving possibility space as the normal case, and studies reasoning that adapts to the uncertainty it is actually in rather than the uncertainty the plan assumed.

Consequence · Aurora, in the programme, is this thesis running as code.

What matters most is latent.

Capability, quality, trust, judgement: none of these can be read off a surface. They are latent variables, inferred from signs, and inference is a signal problem: separating structure from noise without mistaking one for the other.

Two failures follow. The first is measuring the proxy because it is convenient · activity instead of capability, fluency instead of understanding. The second is discarding information because it looked like noise. Often the error in a measurement has structure, and the structure is the finding.

Reading latent qualities well from observable evidence is the studio's one move, applied everywhere. It is also, precisely, what the name means.

Consequence · every instrument scores latent constructs from behaviour, cites its evidence, and treats residuals as data.

Unvalidated evaluation is opinion with formatting.

As AI takes on real work, evaluation becomes the load-bearing wall: of models, of agents, of the people working alongside them, and of the organisations deploying both. Very little of it currently meets the standard psychology sets for a routine personnel decision.

A benchmark score with no reliability estimate is an anecdote. A judge model nobody validated is a rater with unknown biases. Passing a test is not the same as having a property, and confusing the two is how systems end up certified on vibes.

Human assessment learned these lessons through a century of measurement theory and some expensive public failures. AI evaluation can inherit the lessons, or repeat the failures with better hardware.

Consequence · Criterion Labs, in the programme, exists to do the inheriting: validate the judge before the judge scores anything.

The bottleneck moves up. Move with it.

When a capability gets cheap, the constraint on value does not vanish; it relocates one rung up the stack. When calculation became free, the scarce skill became knowing which model to build and whether to trust it. Execution is becoming cheap now, and what grows scarce is knowing what is worth building, orchestrating work you no longer do by hand, verifying what machines produce, and taste.

Organisations can ride this deliberately: measure a capability honestly, hand it to machines once it is stable, and reinvest the freed human capacity one rung up. A capability ratchet, turning automation into development instead of hollowing-out.

The catch is atrophy. Judgement stands on a substrate of practice, and when the substrate quietly erodes, an organisation keeps its confidence while losing its ability to check. The ratchet only works with instruments watching both directions.

Consequence · this is the work the studio most wants to do with organisations over the coming decade: measuring capability well enough to develop it, not just automate it.

A team of agents is an organisation.

The hard problems in multi-agent AI are not model problems. They are coordination problems: role clarity, division of labour, information flow, productive disagreement, leadership across levels. Organisational psychology has studied exactly these for a century, and almost none of it has crossed into agent design.

The studio builds and tests multi-agent systems grounded in that literature: teams whose structure adapts to the work, where disagreement between agents surfaces as information rather than being smoothed over, and where capability is measured at the level of role, team and system, not just the individual model.

The transfer runs both ways. Building organisations out of machines is proving to be a sharp new lens on the human ones.

Consequence · the studio treats organisation design and agent design as one discipline with two substrates.

The open questions

The questions underneath.

Under the theses are the questions that drive them. Most come back to one problem, the one living things solve constantly: how to act well without enough to go on.

  1. How does a mind know when to keep looking, and when to commit to what it has already found?
  2. How should something act when it genuinely does not know, instead of pretending it does?
  3. When a group decides something, how do you tell whether the thinking was any good, or whether everyone just agreed quickly?
  4. Can the principles that let people and other organisms coordinate solve the problems that show up when machines have to work together?
  5. Which parts of an expert's judgement can a machine take over, and which parts stay irreducibly human?
  6. As more of that work shifts to machines, which skills do we lose for lack of practice, and which higher-order ones might the freed-up capacity let us build?
  7. What if the error in a measurement is not noise to clean away, but the most useful thing in it?
  8. Why is the hardest problem in one field so often a solved problem in another, and who sees that first, the expert inside the field or the outsider fluent in a different one?

Writing

Selected public writing behind these theses. The research itself lives on the research site.

More at kareemsoliman.org