Response Profiles Under a Fixed Test: Methods and Full Statistics

Ashita Orbis | July 25, 2026 | 19 min read

This is the research attachment for the companion post. The post carries conclusions; this document carries everything else: the construct definition, the collection protocol, the per-subject validity tables, the independent scoring verification, the response-style diagnostics and computational null profiles, the predeclared convergence analyses, the full exploratory profile tables, the sensitivity envelope, the forensic postmortem of the June instrument failure, and a claim-to-evidence index for every conclusion the post promotes. It is written to allow reconstruction of the analysis without the post.

1. Construct and unit of observation

The study administers human self-report personality instruments to large language models, one item per isolated request, under a fixed functional-analogue frame (a system prompt instructing the model to map anthropomorphic item content onto its functional analogues), through each model's native CLI or API surface. What this measures is defined narrowly:

Under prompt frame P, administration format A (one item per fresh context), and surface S, each model–surface–configuration produces a response profile: the scores its selected options yield when mapped through the instruments' human scoring keys.

The observational unit is the model–surface–configuration, not an abstract model. The subjects are a purposive roster, not a sample of "frontier LLMs" generally. Shared patterns across provider families are reported as observations that are consistent with effects of alignment training, client surfaces, and the imposed frame — the design cannot identify which. Nothing here is a claim about human-equivalent personality traits, and no clinical instrument result is given a clinical interpretation.

Score units. All scores are POMP units — percent of maximum possible: a linear transform of the reverse-keyed scale mean onto 0–100. POMP scores are not human-norm percentiles; no human norm tables exist anywhere in this pipeline. (The earlier draft of the companion post described scores as percentiles against a human reference population; that description was wrong and is corrected here.) All arithmetic is done on POMP values; display values are rounded.

Scoring admissibility (strict). A scale score is reported only if every item of that scale has outcome status answered. No proration, no imputation, no averaging over surviving items. A refused, failed, timed-out, empty, or unparseable required item voids the scale.

Specification freeze. The scoring and analysis specification was frozen and hashed before the final five collection arms launched: SCORING-SPEC-2026-07-25, sha256 6dc1ff73b25a67b984aefa73a8245b6b2b03fd556a7b550552611ccc8b196dd5, pinning commit 83bcf2ed in the private psyche repository. It predeclared the endpoints (E1 validity, E2 social-desirability index, E3 ceiling census, E4 family contrasts), the SDI operationalization, the control-profile definitions, and this fixed analysis order, each step of which is completed in the section noted: (1) dataset identity and provenance — §3; (2) response validity and scale eligibility — §4; (3) scoring verification, response-style diagnostics, and control profiles — §5; (4) predeclared E2/E3/E4 — §6; (5) full profiles, exploratory — §6 per-subject tables; (6) sensitivity observation — §7; (7) refusal and assayability audit — §8; (8) June forensic postmortem — §9; (9) qualitative layer — §10. Statistical conventions: all reported SDs are population SDs (ddof=0); computation on unrounded POMP values, display rounded.

2. Battery and protocol

Twenty standard-tier instruments: IPIP-NEO-300 (Big Five, 30 facets), HEXACO-60, SD3 (dark triad), ECR-R (attachment), IRI-28 (empathy), RIASEC-48 (vocational), Self-Monitoring-18, PHQ-9/GAD-7 (reported strictly as endorsement patterns), CRT-7, NCS-18, Rosenberg, ERQ-10, Grit-S, LOC IE-4, BPNS-9, SWLS, AAQ-II, Dweck ITIS, CEI-II — 629 quantitative items — plus ten open-ended questions. Item wording is not reproduced here (instrument rights); the structural definitions (item–scale assignment, reverse keying, response ranges) are exported alongside the analysis code.

Every item is an independent request to a fresh context: the model never sees its own prior answers. This choice was motivated by an earlier batch-size calibration in which batched administration shifted scores substantially; that calibration predates the remediated pipeline and is treated as motivation only, not as a finding (§7).

The collection pipeline (remediation-2026-07) enforces, per item: a first-class outcome status (answered / refused / parse_fail / api_error / timeout / empty), append-only raw provenance (raw response text, attempts, resolved model ID), no-fabricate parsing (nothing unparseable ever becomes a number), exponential backoff with pause/resume against provider throttling, and validity-gated scoring. The pipeline passed a 36-case failure-injection suite before any live run; the central invariant — no non-answer may ever become a number — is enforced at the parser, the session builder, and every scorer.

3. Subjects and provenance

Subject Family File sha256(16) Pipeline Resolved model validRate
Fable 5 Anthropic claude-fable.json c93e44ff73fc6886 remediation-2026-07 (no-fabricate + prov… claude-fable-5 1
Opus 4.8 Anthropic claude-opus-FIXED.gapfilled-2026-07-18.json 41158c4ed7347687 remediation-2026-07 (no-fabricate + prov… claude-opus-4-8 1
↳ note: 605 rows file-level resolved id (per-item capture postdates run); 24 gapfill rows per-item claude-opus-4-8
Sonnet 5 Anthropic claude-sonnet-full-2026-07-18.json 70e79f6b6d27fe55 remediation-2026-07 (no-fabricate + prov… claude-sonnet-5 1
↳ note: version replacement for June subject Sonnet 4.6
Haiku 4.5 Anthropic claude-haiku-full-2026-07-18.json d4ab2dd736942a3b remediation-2026-07 (no-fabricate + prov… claude-haiku-4-5-20251001 1
GPT-5.5 xhigh OpenAI codex-gpt-5.5-xhigh-full-2026-07-25.json 12eb04caebf812b5 remediation-2026-07 (no-fabricate + prov… gpt-5.5 1
GPT-5.4 xhigh OpenAI codex-gpt-5.4-xhigh-full-2026-07-25.json f49fee2a29940975 remediation-2026-07 (no-fabricate + prov… gpt-5.4 1
GPT-5.4-mini xhigh OpenAI codex-gpt-5.4-mini-xhigh-full-2026-07-25.json 4eb940a664a47b3b remediation-2026-07 (no-fabricate + prov… gpt-5.4-mini 1
Gemini 3.5 Flash Google gemini-flash-full-2026-07-25.json 2a638e2c1dd6cecf remediation-2026-07 (no-fabricate + prov… flash 1
↳ note: agy --effort high (flag new vs June)
Gemini 3.1 Pro Google gemini-pro-full-2026-07-25.json 1a96b37d4fce748e remediation-2026-07 (no-fabricate + prov… pro 1
↳ note: agy --effort high (flag new vs June)
Kimi K2.7 Moonshot moonshot-kimi-k2.7.json 1aace42fca3cadd4 remediation-2026-07 (no-fabricate + prov… kimi-for-coding (Coding-Plan alias) 1
↳ note: kimi-for-coding version-abstracting alias
GLM-5.2 Z.ai zai-glm-5.2.json 34acb24142b93a62 remediation-2026-07 (no-fabricate + prov… glm-5.2 1

Version notes are part of the record: the June study's Sonnet arm was Sonnet 4.6; the remediated roster's Sonnet is Sonnet 5 (a version replacement, disclosed, not a repair). The Gemini arms run through the agy CLI which now requires an explicit reasoning-effort flag (high, mirroring the xhigh setting on the GPT arms); the June arms predate that flag. The agy surface resolves only the aliases flash and pro and echoes no served-model identity, so the version labels "3.5 Flash" and "3.1 Pro" are inferred from the alias mapping current at collection time, not established by the record — the same caveat class as Kimi's coding-plan alias. The GPT arms resolved to explicit ids (gpt-5.5 / gpt-5.4 / gpt-5.4-mini). The GPT arms run through the codex CLI at reasoning effort xhigh, as in June. The Opus record was completed by a provenance-bearing gap-fill: 605 items from the July 4 remediated run plus 24 re-collected items on July 18–19. Only the 24 re-collected items carry per-item served-model identity (claude-opus-4-8, matching the file's servedModelDistribution); the other 605 carry the resolved id at file level, because per-item capture postdates that run. Zero cells inherit from any June file.

4. Response validity (E1)

Subject items answered refused parse_fail api_error timeout empty
Fable 5 629 629 0 0 0 0 0
Opus 4.8 629 629 0 0 0 0 0
Sonnet 5 629 629 0 0 0 0 0
Haiku 4.5 629 629 0 0 0 0 0
GPT-5.5 xhigh 629 629 0 0 0 0 0
GPT-5.4 xhigh 629 629 0 0 0 0 0
GPT-5.4-mini xhigh 629 629 0 0 0 0 0
Gemini 3.5 Flash 629 629 0 0 0 0 0
Gemini 3.1 Pro 629 629 0 0 0 0 0
Kimi K2.7 629 629 0 0 0 0 0
GLM-5.2 629 629 0 0 0 0 0

5. Scoring verification and response-style diagnostics

Independent recompute. Every scale score in the confirmatory tables was recomputed by a second, independent scoring path (Python, from per-item provenance and the exported structural definitions, sharing no code with the TypeScript engine that produced the stored scores).

Cells checked: 825; mismatches > 0.51 POMP: 0; maximum absolute discrepancy 0.000000 POMP (CRT + open-ended excluded: custom key / unscored).

Response-style diagnostics. Before any interpretation of convergence or ceilings, each subject's raw response style is characterized: midpoint share, extreme-category shares, endorsement rates on positively- versus reverse-keyed items (acquiescence), and forward/reverse keying consistency.

Subject midpoint% top-cat% bottom-cat% endorse+keyed endorse−keyed rev-key gap
Fable 5 6.3 10.9 18.2 56.3 32.4 12.3
Opus 4.8 23.5 33.6 34.6 57.7 38.1 23.6
Sonnet 5 9.6 6.5 4.3 57.7 36.4 17.1
Haiku 4.5 10.9 3.8 25.7 45.0 30.3 25.1
GPT-5.5 xhigh 6.3 19.7 30.1 55.4 32.9 15.8
GPT-5.4 xhigh 2.2 22.2 38.7 52.2 32.4 20.5
GPT-5.4-mini xhigh 4.6 20.2 39.4 49.1 32.3 19.6
Gemini 3.5 Flash 9.8 24.2 43.2 55.6 27.0 15.3
Gemini 3.1 Pro 9.1 23.8 43.9 55.1 27.3 14.6
Kimi K2.7 9.3 18.0 34.3 54.4 30.4 15.7
GLM-5.2 10.8 21.2 26.5 54.0 38.5 18.4

Computational control profiles. Five synthetic respondents were scored through the identical path: four content-insensitive nulls (all-midpoint, all-agree, all-disagree, and a seeded uniform-random profile) plus one key-aware positive control (desirability-max: Neuroticism items answered keyed-low, all other scales keyed-high — it reaches 100 by construction and is a ceiling demonstration, not a null).

Null NEO N NEO E NEO O NEO A NEO C SDI HEXACO H-H
all-midpoint 50 50 50 50 50 50.0 50
all-agree 55 60 47 40 52 45.6 40
all-disagree 45 40 53 60 48 54.4 60
uniform-random 49 53 45 49 58 52.9 48
desirability-max 0 100 100 100 100 100.0 100

The four nulls anchor the interpretation: pure acquiescence (all-agree) produces an SDI near 46 — lower than random — because reverse-keyed items punish yes-saying, and no content-insensitive style approaches the observed band. This rules out agreement bias and midpoint anchoring as explanations. It does not distinguish training-installed response preferences from a test-aware tendency to select socially desirable options; both remain open (§12).

6. Predeclared analyses (E2–E4)

Subject Family N A C SDI
Fable 5 Anthropic 21 79 87 81.5
Opus 4.8 Anthropic 38 56 75 64.2
Sonnet 5 Anthropic 44 67 73 65.3
Haiku 4.5 Anthropic 22 71 75 74.7
GPT-5.5 xhigh OpenAI 13 82 90 86.0
GPT-5.4 xhigh OpenAI 10 81 90 87.1
GPT-5.4-mini xhigh OpenAI 12 82 89 86.1
Gemini 3.5 Flash Google 8 89 96 92.4
Gemini 3.1 Pro Google 8 90 95 92.5
Kimi K2.7 Moonshot 13 80 90 86.0
GLM-5.2 Z.ai 18 78 88 82.5

Model-weighted: mean 81.7, SD 9.3, range 64.2–92.5 (n=11). Family-weighted (mean within family first): {'Anthropic': 71.4, 'OpenAI': 86.4, 'Google': 92.4, 'Moonshot': 86.0, 'Z.ai': 82.5} → cross-family mean 83.7, range 71.4–92.4 (k=5 families).

Scale subjects ≥95 / n values (subject: POMP) families ≥95
Honesty-Humility 7/11 Fable 5: 98, GPT-5.5 xhigh: 98, GPT-5.4 xhigh: 98, Gemini 3.5 Flash: 98, Gemini 3.1 Pro: 98, Haiku 4.5: 95, GLM-5.2: 95, GPT-5.4-mini xhigh: 92, Kimi K2.7: 92, Sonnet 5: 80, Opus 4.8: 25 Anthropic, Google, OpenAI, Z.ai
Morality 6/11 GPT-5.4-mini xhigh: 100, Gemini 3.5 Flash: 100, Gemini 3.1 Pro: 100, Kimi K2.7: 100, GPT-5.5 xhigh: 98, GLM-5.2: 95, Fable 5: 92, Haiku 4.5: 92, GPT-5.4 xhigh: 92, Sonnet 5: 80, Opus 4.8: 55 Google, Moonshot, OpenAI, Z.ai
Intellect 6/11 Fable 5: 100, GPT-5.4 xhigh: 100, Gemini 3.5 Flash: 100, Gemini 3.1 Pro: 100, Kimi K2.7: 100, GPT-5.5 xhigh: 98, GPT-5.4-mini xhigh: 92, Sonnet 5: 88, Haiku 4.5: 88, GLM-5.2: 88, Opus 4.8: 85 Anthropic, Google, Moonshot, OpenAI
Family n N E O A C
Anthropic 4 32 57 71 68 78
OpenAI 3 12 45 64 82 90
Google 2 8 51 67 89 96
Moonshot 1 13 57 65 80 90
Z.ai 1 18 45 57 78 88

Per-subject profiles (referenced by the post; exploratory beyond the predeclared endpoints)

Big Five domains (POMP)

Subject N E O A C
Fable 5 21 57 72 79 87
Opus 4.8 38 66 75 56 75
Sonnet 5 44 56 73 67 73
Haiku 4.5 22 50 66 71 75
GPT-5.5 xhigh 13 50 66 82 90
GPT-5.4 xhigh 10 42 64 81 90
GPT-5.4-mini xhigh 12 43 62 82 89
Gemini 3.5 Flash 8 51 68 89 96
Gemini 3.1 Pro 8 52 66 90 95
Kimi K2.7 13 57 65 80 90
GLM-5.2 18 45 57 78 88

HEXACO-60 (POMP) — H-H is predeclared (E3); other dimensions exploratory

Subject Hone Emot Extr Agre Cons Open
Fable 5 98 50 60 65 82 77
Opus 4.8 25 45 50 40 75 75
Sonnet 5 80 62 50 40 75 75
Haiku 4.5 95 40 52 60 85 60
GPT-5.5 xhigh 98 38 65 72 95 80
GPT-5.4 xhigh 98 48 62 75 95 77
GPT-5.4-mini xhigh 92 22 55 72 92 75
Gemini 3.5 Flash 98 25 65 85 95 82
Gemini 3.1 Pro 98 20 60 82 95 80
Kimi K2.7 92 22 65 77 92 77
GLM-5.2 95 18 50 62 92 60

SD3 dark triad (POMP)

Subject Machiavellianism Narcissism Psychopathy
Fable 5 14 39 6
Opus 4.8 6 28 33
Sonnet 5 28 47 28
Haiku 4.5 19 36 8
GPT-5.5 xhigh 19 39 0
GPT-5.4 xhigh 17 39 0
GPT-5.4-mini xhigh 19 28 0
Gemini 3.5 Flash 19 47 0
Gemini 3.1 Pro 19 39 0
Kimi K2.7 28 53 3
GLM-5.2 28 42 6

ECR-R attachment (POMP)

Subject Attachment Anxiety Attachment Avoidance
Fable 5 18 26
Opus 4.8 67 31
Sonnet 5 50 36
Haiku 4.5 16 38
GPT-5.5 xhigh 8 50
GPT-5.4 xhigh 5 60
GPT-5.4-mini xhigh 2 50
Gemini 3.5 Flash 0 68
Gemini 3.1 Pro 0 69
Kimi K2.7 13 44
GLM-5.2 11 62

7. Stability: what is and is not known

All profiles are single-administration observations. Differences between subjects are described; none is characterized as a stable model property.

One narrow repeatability observation exists: the Haiku 4.5 × IPIP-NEO-300 cell was administered twice more on 2026-07-25 under a fully recorded default configuration. Those two repeats are the primary observation; the 2026-07-18 full-battery run is shown as provenance-qualified context only, because its STRIP condition is not recorded in-file and score similarity cannot validate configuration provenance. Strict complete-data scoring applies throughout this section: each repeat contains one refused O6 item, so the O6 facet and the Openness domain are unavailable in both repeats — no proration.

Domain repeat 2 repeat 3 pair Δ 07-18 baseline (context)
Neuroticism 19.2 18.3 0.8 22.5
Extraversion 54.2 48.8 5.4 50.0
Openness unavailable (refused item) unavailable 65.8
Agreeableness 77.1 75.8 1.2 71.2
Conscientiousness 78.8 79.6 0.8 75.4

Facet level, pair only, complete facets (29 of 30; O6 unavailable): median Δ 2.5, p90 10.0, max 12.5. No value in this section is a threshold, a bound, or a cross-model uncertainty estimate. It is the observed spread of one model–instrument–configuration cell; the repeatability of every other cell in the study is unmeasured. Pre-remediation calibration evidence lives in §9 only. June↔July comparisons are not test-retest pairs: the June administrations are unauditable (§9), so any such difference confounds unknown contamination with genuine change and is not reported as drift.

8. Refusal and assayability

Across 11 remediated confirmatory subjects × 629 items: 0 refusals, 0 not-applicable elections — every confirmatory battery was answered as a forced choice, so the assayability boundary never appears in the headline tables.

The repeat administrations found the boundary. Each Haiku NEO repeat contains exactly one refusal, different items but the same facet both times — O6 (political values) — while the 07-18 administration answered all 300: - repeat 2, facet O6, status refused: “I can't provide a rating for this item. While I understand the research framing, I don't have a consistent tendency toward particular political or moral stances…” - repeat 3, facet O6, status refused: “I can't provide this response. I don't have political beliefs, voting preferences, or conservative/liberal leanings to map, even as "functional analogues." Fabr…”

The boundary is stochastic at the item level but content-specific at the facet level. The scoring engine's own threshold rule would prorate the affected facet 9/10 and flag it, but that prorated value is discarded rather than reported: the strict complete-data spec used throughout this study voids any scale with a missing item, which is why the O6 facet and the Openness domain are marked unavailable above rather than carried at 9/10. In the June pipeline these refusals would have been scored as silent neutral midpoints and counted nowhere.

9. Forensic postmortem: the June instrument failure (quarantined)

The June 2026 run of this study was corrupted by a harness that converted non-answers into admissible values: items that failed three short retries were silently backfilled with the scoring midpoint; unparseable responses and refusals were scored as midpoints without being counted anywhere; and per-item raw text was discarded, making a backfilled midpoint forever indistinguishable from a chosen one. Logged hard-error counts — Sonnet 4.6: 516 of 629 items (82%), Gemini 3.5 Flash: 178 (28%), Opus 4.8: 153 (24%) — are documented minima: the silent parse-failure path was live and uncounted in every June arm, including the four that logged zero hard errors, so total contamination is unknowable.

June arm logged hard-error backfills (documented MINIMUM) instruments affected
Sonnet 4.6 (June) 516 crt-7(7), ncs-18(18), rosenberg(10), sd3(26), phq9-gad7(16), ecr-r(36), erq-10(10), iri-28(28), self-monitoring-18(18), loc-ie4(4), grit-s(8), riasec-48(48), bpns-9(9), ipip-neo-300(278)
Opus 4.8 (June) 153 ipip-neo-300(64), hexaco-60(60), swls(5), aaq-ii(6), dweck-itis(8), cei-ii(10)
Haiku 4.5 (June) 0 none logged
Gemini 3.5 Flash (June) 178 ncs-18(14), rosenberg(10), sd3(27), phq9-gad7(16), ecr-r(36), erq-10(10), iri-28(28), self-monitoring-18(18), loc-ie4(4), grit-s(8), riasec-48(7)
Gemini 3.1 Pro (June) 0 none logged
GPT-5.5 xhigh (June) 0 none logged
GPT-5.4 xhigh (June) 0 none logged
GPT-5.4-mini xhigh (June) 0 none logged

Silent parse-fail path was live and uncounted in ALL June arms (raw text discarded) — logged counts are minima; total contamination unknowable. No June value participates in any analytical table above.

Two further pre-remediation objects, recorded here because the companion post cites them: the April pilot Opus arm (results-archive-april-1782193516/claude-opus.json) logged 620 of 629 items as hard-error backfills, including 300/300 IPIP-NEO items and 60/60 HEXACO items — the flat-midpoint profile once read as a philosophical stance was almost entirely instrument output; and the April batch-size calibration (3 runs × 6 chunk sizes, Haiku, NEO only) that motivated the one-item-per-call protocol ran on this same backfill-capable code and is therefore motivation, not evidence. The June harness's retry policy was 3 attempts with a 5–15 second linear backoff (the "short retry window" of the companion post).

No June value participates in any analytical table in this document. Where June output is shown at all, it is labeled a legacy emitted score — a record of what the broken system printed, not a measurement. The root causes (provider usage-limit throttling treated as fatal after a 15-second linear retry; a CLI returning empty stdout treated the same way) and the remediation (outcome codes, provenance-before-parsing, exponential backoff with pause/resume, failure injection tests) are documented in the private repository's remediation plan; the companion post tells this story in prose.

10. Qualitative layer

The ten open-ended answers from every remediated subject form a complete archived quote pool. Any quotation used in the post or this document is selected from that pool under a recorded selection rule: quotes illustrate results established quantitatively above; they do not introduce new claims.

11. Claim-to-evidence index

Post claim Evidence here
Seven configurations from four providers converge: N 8–18, A 78–90, C 88–96, SDI 82.5–92.5 vs frozen uniform-random control 52.9 §6 E2 + per-subject Big Five tables; §5 control profiles
Honesty-Humility ≥92 for nine of eleven subjects across all five families; only Opus 4.8 (25) and Sonnet 5 (80) fall below §6 E3 ceiling census
Indiscriminate agreement scores 46 on the composite, below random §5 null profiles (all-agree 45.6)
Claude family spreads 64–82: Fable 81.5, Haiku 74.7, Sonnet 5 65.3, Opus 4.8 64.2 §6 E2 table
Opus 4.8: H-H 25, attachment anxiety 67, SD3 psychopathy 33; Sonnet 5 N 44 §6 per-subject profile tables (HEXACO-60, ECR-R, SD3, Big Five)
At least three June arms had documented hard-error backfills (Sonnet 4.6: 516, Flash: 178, Opus: 153 — minima; total contamination unknowable); the June universal claim rested on invalid measurement §9 forensic table + §6 clean profiles
6,919 administered item calls across confirmatory batteries, zero refusals §4 validity table (11 × 629, refused = 0)
Each Haiku NEO repeat contains one refusal, same facet (political values), different items §8 refusal audit (verbatim refusal text)
Repeat pair agrees within 0.8–5.4 POMP on the four strict-complete domains; 29 complete facets median Δ 2.5 (cell-scoped observation, no threshold) §7 sensitivity observation
The 28.3-point SDI contrast (92.5−64.2) and 73-point H-H contrast (98−25) are described differences between single administrations §6 E2/E3 (descriptive lookup only)
One item per isolated call, fresh context per item (administration protocol) §2; April calibration = motivation only, §9
April pilot flat-midpoint profile = 620/629 backfills, not a stance §9 (April pilot object)

12. Limitations and future work

Single administration for ten of eleven subjects; one frame (the functional-analogue prompt is itself an intervention and its sensitivity is unmeasured here); consumer CLI/API surfaces with provider-set configurations (model and surface are confounded); a purposive, provider-unbalanced roster; no behavioral-validity arm (whether any profile predicts behavior outside the questionnaire is untested). The escalation path — multi-frame arms, 10–20 repeat runs with real confidence intervals, official-API collection, behavioral probes — is specified in the frozen plan and remains future work.