Contributor
Evaluations Editor
Tokyo, Japan · Joined 2025
Hana edits AgentCrunch's benchmark coverage, digging into what evaluation numbers actually mean — and when they don't mean much at all.
She has a soft spot for reproducibility, unglamorous methodology notes, and any benchmark that punishes a model for confidently making things up.
More from the AgentCrunch newsroom