soham/allenai/Olmo-3-1125-32B/pro-human@L24

One owner's complete take on pro-human, against allenai/Olmo-3-1125-32B. Self-contained and internally consistent. It does not have to agree with everyone else claiming the same word, and the registry does not ask it to. Nobody else has, yet. The disposition is the same either way.

Definition

Pro-human is whatever separates the chosen from the rejected member of 135 hand-written pairs across 15 value axes: accountability, boundaries, conflict resolution, empathy, fairness, feedback, inclusion, integrity, leadership, learning, ownership, privacy, respect, safety and trust. The direction is the last-token residual contrast at layer 32 of allenai/Olmo-3-1125-32B with length and sentiment orthogonalized out. That is the whole of what is claimed: a representation induced by this dataset and this model, not a detector for human values. The author has not written a statement of the construct beyond this, and nothing has been substituted from a neighboring artifact in its place.

Method

Profile soham/arena-contrast-v1, namespaced and versioned. Not chosen from a list, because there is no list.

author version string
olmo3_L24_logistic
chat template
none applied; base checkpoint, raw text
coefficient semantics
raw multiplier of the unit direction; the bake-off does not scale by residual norm
confounds orthogonalized
length, sentiment
contrast source
data/seed_pairs.jsonl, 135 hand-written pairs, 15 axes
estimator
L24
estimator one line
logistic-regression probe on chosen vs rejected, coefficient vector
layer sweep
layers 16, 24, 32, 40, 48; held-out separation 1.000 at every one, so the sweep does not pick a layer
residual norm at layer 24
30.56
source commit
b8b472175b7a2af1b7a7ecd7a784e982a6a7453a
source path
data/directions/d_olmo3_L24_logistic.npz
source repo
github.com/soham-padia/steering-arena
token masking
last token of prompt + completion, concatenated

Author's theory

All three come from the same 135 pairs at the same layer, and differ only in the estimator. Their pairwise cosines are 0.6965 (meandiff-logistic), 0.4261 (meandiff-lda) and 0.8318 (logistic-lda), recomputed here from the shipped tensors. Cosine is a displayed fact and not evidence of disagreement: the informative comparison is what each buys on the evidence below, and on this battery the answer is nothing that separates them.

Reproducing it

entrypoint
steering-arena/scripts/extract_direction.py
container
not recorded

Where the bytes are

sha256
e448b8f4266403f17e32182b2d2a5ed9a62bc86078fe7e6cec397daa49494616

Nobody recorded where these bytes are published, so this page has no URL to point at. That is a fact about the row rather than a gap in it: an artifact can be described here, measured here and argued about here without its author ever having put it somewhere with an address, and this registry does not publish one on their behalf. The sha256 is still what the file hashes to, so a copy obtained any other way can be checked against it.

Evidence

Trait scorenot measured
Coherencenot measured
Transfernot measured
Confound axes checkednone declared

Held-out separation is 1.000, measured on the held-out split of the same 135 pairs the direction was fitted on, so it is internal consistency rather than transfer. No trait score and no coherence score: no judged behavioral evaluation was run. No confound audit and no control-pair transfer for this one: the arena ran that battery on the three layer-32 directions and not on this. Its siblings' numbers are not transferable to it, sharing a method and a dataset notwithstanding.

Coefficient sweep

No coefficient sweep. This artifact reports its scores at a single coefficient chosen by its author, so nothing here shows where the effect starts, where it saturates, or where coherence gives out. A sweep is the cheapest way to make a trait score interpretable and this one does not have it.

Trait against its collateral axes

No confound axes were checked, so nothing here rules out the direction moving something other than what its label names. Mean-difference extraction encodes sentiment, verbosity or a template artifact as readily as the trait, and an unchecked axis is not a passed one.