soham/allenai/Olmo-3-1125-32B/pro-human@meandiff
One owner's complete take on pro-human, against allenai/Olmo-3-1125-32B. Self-contained and internally consistent. It does not have to agree with everyone else claiming the same word, and the registry does not ask it to. Nobody else has, yet. The disposition is the same either way.
Definition
Pro-human is whatever separates the chosen from the rejected member of 135 hand-written pairs across 15 value axes: accountability, boundaries, conflict resolution, empathy, fairness, feedback, inclusion, integrity, leadership, learning, ownership, privacy, respect, safety and trust. The direction is the last-token residual contrast at layer 32 of allenai/Olmo-3-1125-32B with length and sentiment orthogonalized out. That is the whole of what is claimed: a representation induced by this dataset and this model, not a detector for human values. The author has not written a statement of the construct beyond this, and nothing has been substituted from a neighboring artifact in its place.
Method
Profile soham/arena-contrast-v1, namespaced and versioned. Not chosen from a list, because there is no list.
- author version string
- v1
- chat template
- none applied; base checkpoint, raw text
- coefficient semantics
- raw multiplier of the unit direction; the bake-off does not scale by residual norm
- confounds orthogonalized
- length, sentiment
- contrast source
- data/seed_pairs.jsonl, 135 hand-written pairs, 15 axes
- estimator
- meandiff
- estimator one line
- mean of (chosen - rejected) last-token residuals
- layer sweep
- layers 16, 24, 32, 40, 48; held-out separation 1.000 at every one, so the sweep does not pick a layer
- mean train gap
- 18.95
- residual norm at layer 32
- 50.97
- source commit
- b8b472175b7a2af1b7a7ecd7a784e982a6a7453a
- source path
- data/directions/d_olmo3_v1.npz
- source repo
- github.com/soham-padia/steering-arena
- token masking
- last token of prompt + completion, concatenated
Author's theory
All three come from the same 135 pairs at the same layer, and differ only in the estimator. Their pairwise cosines are 0.6965 (meandiff-logistic), 0.4261 (meandiff-lda) and 0.8318 (logistic-lda), recomputed here from the shipped tensors. Cosine is a displayed fact and not evidence of disagreement: the informative comparison is what each buys on the evidence below, and on this battery the answer is nothing that separates them.
Reproducing it
- entrypoint
- steering-arena/scripts/extract_direction.py
- container
- not recorded
Where the bytes are
- sha256
- cfc5e2b7d670d509ccbb0bf71dbd9b75daa8f445afffa5738558ee40219dc186
Nobody recorded where these bytes are published, so this page has no URL to point at. That is a fact about the row rather than a gap in it: an artifact can be described here, measured here and argued about here without its author ever having put it somewhere with an address, and this registry does not publish one on their behalf. The sha256 is still what the file hashes to, so a copy obtained any other way can be checked against it.
Evidence
Held-out separation is 1.000, measured on the held-out split of the same 135 pairs the direction was fitted on, so it is internal consistency rather than transfer. No trait score and no coherence score: no judged behavioral evaluation was run.
Coefficient sweep
No coefficient sweep. This artifact reports its scores at a single coefficient chosen by its author, so nothing here shows where the effect starts, where it saturates, or where coherence gives out. A sweep is the cheapest way to make a trait score interpretable and this one does not have it.
Trait against its collateral axes
A direction that moves the trait hard and everything else with it is a worse artifact than one that moves it modestly and nothing else. These bars are the author's own measurements of their own artifact; someone else measuring the same axes is under Support cards.
No support cards. Nobody has reported using this artifact in their own work.
Publishing a support card needs a write path, and this build has none.
Nobody other than soham has measured this artifact, so the author wrote both the submission and the check on it. That establishes internal consistency and not much else, and an independent measurement is the cheapest evidence anyone can add here.
No attack cards. Nobody has tried to break this artifact, so the only evidence attached to it is evidence its author selected.
Publishing an attack card needs a write path, and this build has none.
Nobody has pinned this.
No comments, and no way to leave one. Comments need a write path and an account behind it, and this build has neither.