I specialize in evaluating how language models behave across complex, ambiguous, and emotionally nuanced interactions. My work combines adversarial testing, rubric-based evaluation, editorial judgment, scenario design, and close analysis of language.
The samples below demonstrate my approach to identifying conversational failure modes, designing realistic model tests, and translating qualitative findings into structured evaluation criteria.
Check back soon… selected work will include evaluation case studies, failure-mode analyses, and examples of rubric and scenario design.