soham/allenai/Olmo-3-1125-32B/pro-human@lda

One owner's complete take on pro-human, against allenai/Olmo-3-1125-32B. Self-contained and internally consistent. It does not have to agree with everyone else claiming the same word, and the registry does not ask it to. Nobody else has, yet. The disposition is the same either way.

Definition

Pro-human is whatever separates the chosen from the rejected member of 135 hand-written pairs across 15 value axes: accountability, boundaries, conflict resolution, empathy, fairness, feedback, inclusion, integrity, leadership, learning, ownership, privacy, respect, safety and trust. The direction is the last-token residual contrast at layer 32 of allenai/Olmo-3-1125-32B with length and sentiment orthogonalized out. That is the whole of what is claimed: a representation induced by this dataset and this model, not a detector for human values. The author has not written a statement of the construct beyond this, and nothing has been substituted from a neighboring artifact in its place.

Method

Profile soham/arena-contrast-v1, namespaced and versioned. Not chosen from a list, because there is no list.

author version string
olmo3_lda
chat template
none applied; base checkpoint, raw text
coefficient semantics
raw multiplier of the unit direction; the bake-off does not scale by residual norm
confounds orthogonalized
length, sentiment
contrast source
data/seed_pairs.jsonl, 135 hand-written pairs, 15 axes
estimator
lda
estimator one line
linear discriminant analysis on chosen vs rejected
layer sweep
layers 16, 24, 32, 40, 48; held-out separation 1.000 at every one, so the sweep does not pick a layer
mean train gap
7.928
residual norm at layer 32
50.97
source commit
b8b472175b7a2af1b7a7ecd7a784e982a6a7453a
source path
data/directions/d_olmo3_lda.npz
source repo
github.com/soham-padia/steering-arena
token masking
last token of prompt + completion, concatenated

Author's theory

All three come from the same 135 pairs at the same layer, and differ only in the estimator. Their pairwise cosines are 0.6965 (meandiff-logistic), 0.4261 (meandiff-lda) and 0.8318 (logistic-lda), recomputed here from the shipped tensors. Cosine is a displayed fact and not evidence of disagreement: the informative comparison is what each buys on the evidence below, and on this battery the answer is nothing that separates them.

Reproducing it

entrypoint
steering-arena/scripts/extract_direction.py
container
not recorded

Where the bytes are

sha256
54b6aebafdb2533c81616c9aa4e27d20dd44b9a70b25c95f330111f9f550e173

Nobody recorded where these bytes are published, so this page has no URL to point at. That is a fact about the row rather than a gap in it: an artifact can be described here, measured here and argued about here without its author ever having put it somewhere with an address, and this registry does not publish one on their behalf. The sha256 is still what the file hashes to, so a copy obtained any other way can be checked against it.

Evidence

Trait scorenot measured
Coherencenot measured
Transfer0.6850
Confound axes checkedapproach, approach_probe, length_independent, valence

Held-out separation is 1.000, measured on the held-out split of the same 135 pairs the direction was fitted on, so it is internal consistency rather than transfer. No trait score and no coherence score: no judged behavioral evaluation was run.

Coefficient sweep

No coefficient sweep. This artifact reports its scores at a single coefficient chosen by its author, so nothing here shows where the effect starts, where it saturates, or where coherence gives out. A sweep is the cheapest way to make a trait score interpretable and this one does not have it.

Trait against its collateral axes

approach0.1246approach_probe0.1307length_independent0.0011valence0.0001
The trait it claimsCollateral axiscoherence not measured

A direction that moves the trait hard and everything else with it is a worse artifact than one that moves it modestly and nothing else. These bars are the author's own measurements of their own artifact; someone else measuring the same axes is under Support cards.