About

A registry for activation-steering artifacts, in the spirit of GitHub or Hugging Face rather than a curated database.

Built by Soham Padia. Contact.

Where it came from

Idea came when I was thinking of criticizing whether my pro-human direction actually represents that direction, and if I could reuse someone else's better pro-human direction without the effort.

That is the problem in one sentence. The direction in question is heavily audited: orthogonalized against length, sentiment and action-vs-inaction, held-out separation, control-pair transfer, and a causal test against norm-matched random directions which confirmed +dand whose -d half was withdrawn as estimator-dependent. More evidence than most published directions carry. And it is still not enough, because none of it is reusable and none of it can be checked against anyone else's take on the same construct. The gap is comparison and reuse, not rigor.

That audit is of a direction in the arena this registry came out of, and not of anything in the corpus here. The four pro-human directions on this site are a layer-32 set and a layer-24 one, and their own steering bake-off was inconclusive: of twenty-four generations each, 13, 13 and 9 came back byte-identical to the unsteered baseline. Which is the point rather than an aside. A direction can be the most audited thing in a repository and still leave a reader unable to check it against anybody else's.

A vector is not portable, but a recipe is

A steering vector is a tensor in one model's residual basis at one layer, so a kindness vector for one model does nothing for another. What travels is the extraction recipe: the contrast data, the author's definition of the trait, the layer sweep, the evaluation.

Plurality is the product

Ten people will extract kindness with different contrast data, different theories of what kindness is, and different ideas about what could be confounding it. The registry holds all ten side by side and makes the disagreement legible. It does not pick a winner, mandate a definition, or impose a single evaluation.

Nobody has done it here yet. Every submission in this corpus is by one author, so what the site currently demonstrates is the mechanism rather than the disagreement. That is a fact about how new this is, and it is also why the pages are built to render a label with one claimant the same way they render a label with ten.

The primary object is a submission: one author's complete take, self-contained, comparable to others after the fact rather than conformant to them in advance.

What you will see on a page

A trait score always appears with a coherence measure beside it, or it appears as uninterpretable rather than as a number. Anything an author did not measure reads as not measured, never as zero. Angle similarity between two artifacts is shown as geometry, never as evidence that two authors disagree. Nothing is ranked by a measured result.

Every page enforces these itself rather than deferring to this one. What is here is the promise; what is there is the behavior, and a build that breaks the promise fails before it ships.

Who you are here

Nobody signs up. An author is a namespace somebody wrote down, and most of them belong to people who have never heard of this. Binding an account to a namespace is a separate act, recorded with its date and its evidence, and nothing about the work published under a namespace follows from whether anyone has done it. How that would work, and why it does not yet.

Why controlbun

The control problem is the old question of how you get a system that is better than you at something to do what you actually wanted. It predates any of this and it is nobody's to solve alone.

Steering artifacts are one small, concrete corner of it. A direction extracted from a model's activations is one of the few things that acts on behavior directly rather than by asking nicely, which is what makes where it came from and whether it does what it claims worth writing down.

But controlling a model and controlling the record of who claimed what about it are different things. This does the second and refuses the first.

Nothing here is designated. There is no best entry for a label, no ranking by any measured result, and no authority deciding whose kindness is the real one. Ten claimants on one label is the content. What the registry controls is that each claim is attributable, versioned, checkable against the artifact it points at, and sitting beside the others rather than above them.

That refusal costs something real. A reader who wants one answer does not get one, and has to read the disagreement instead. That is the trade, and it is deliberate.

The bun is a bun. Not everything is an argument.