Research
Working papers, analyses, and technical investigations. Not peer-reviewed.
Benchmarking Structured Conversation Extraction Across Three Evaluation Layers: A Comparison of Eight Models with a Public Replication
Evaluating structured extraction from conversational data with a single quality metric conceals the failure modes that matter most to systems built on top of the extraction...
The Etymology Tax: Etymological Register Effects on LLM Multi-Step Reasoning
This study tests whether the etymological register of input text affects large language model performance on multi-step reasoning tasks...
Convergent AI-Mediated Personality Assessment: Psychometric Profiling and Narrative Inference from Digital Communication Data
This paper examines the convergence between two AI-mediated approaches to personality assessment applied to a single participant: explicit psychometric measurement through sixteen validated instruments triangulated across three inference methods, and implicit personality inference through literary narrative generation from a 267MB text messaging archive...
BullshitBench v2: A Methodology Critique
Statistical methodology critique and extended analysis of BullshitBench v2's rubric design, judge reliability, and scoring philosophy.