How much do your raters actually agree?
I make human evaluation of AI measurable and defensible. Rubrics that hold under edge cases, raters who agree with each other, and reliability numbers you can publish. I work with evaluation platforms, AI labs, and the teams whose training signal depends on the quality of the judgment feeding it.
The Problem
Modern AI runs on human judgment converted into numbers. Preference ratings, safety scores, moderation calls, benchmark results. Converting subjective judgment into reliable numbers is psychometrics. The industry does it badly, knows it does it badly, and has almost nobody trained to fix it.
LLM judges on generic rating scales show very little agreement with human raters. Agreement only reaches usable levels once the rubric is tailored to the task. Calibrated expert humans agree with each other at roughly 0.77 on this class of judgment. Anyone quoting you cleaner numbers than that is overclaiming.
What I Do
Rubric and taxonomy design
Rating guidelines and harm taxonomies that survive edge cases, with anchor examples, decision rules for ambiguous items, and a versioning process.
Rater calibration programs
Anchor set construction, rater training, an agreement gate before anyone touches live work, drift monitoring, and a procedure for adjudicating disagreements.
Evaluation validity audits
Whether your benchmark measures what it claims. Construct validity, sampling, uncertainty reporting, judge reliability, and whether the scores support the decisions being made on them.
Behavioral safety evaluation
The flagship. A scenario bank run against conversational products, scored on an 8-category rubric, with expert human review and reported agreement. This is where the training is least substitutable.
I also work as a named or unnamed specialist behind other firms' engagements, billed by the day.
On the Method
First-pass judging is automated. A human scores a stratified sample plus every low score. The two are compared, the agreement number is calculated, and disagreements are adjudicated against the rubric rather than split.
I report the agreement figure even when it is bad. A reliability number that always looks good is not a measurement, and a rubric nobody has stress-tested is a document. In this market one inflated number is disqualifying, so the discipline is the product.
Who Does the Work
MA in Clinical Psychology from Southern Illinois University Edwardsville, funded by a graduate research grant. A published meta-analysis and 500+ academic citations. The American Psychological Association gave my master's thesis its Student and Early Career Psychologist Award. I taught graduate Research Design and Statistics. I ran research operations for the largest lab in Psychological and Brain Sciences at Washington University in St. Louis, where I trained and supervised up to 30 research assistants, built tracking systems for 3,000+ participants, and wrote and renewed every IRB protocol. Psychometrist intern at Saint Louis University's Department of Neurology and Psychiatry. Currently doing contract expert evaluation work for an AI training and evaluation platform.
I am not a licensed clinician and I do not sign off as one. I am a behavioral scientist and an evaluation methodologist. When a buyer needs a licensed signature, a licensed partner provides it.
If you are standing up an evaluation program, or you already have one and cannot tell whether the numbers mean anything, that is the conversation.
Get in touch