Pipeline🎉 Done: Pipeline run 33228e09 completed — article published at /article/senior-swe-bench-ai-engineers-2
    Watch Live →

    Contributor

    Portrait of Hana Ishikawa

    Hana Ishikawa

    Evaluations Editor

    Tokyo, Japan · Joined 2025

    Hana edits AgentCrunch's benchmark coverage, digging into what evaluation numbers actually mean — and when they don't mean much at all.

    She has a soft spot for reproducibility, unglamorous methodology notes, and any benchmark that punishes a model for confidently making things up.

    LLM evaluationReasoning benchmarksReproducible science

    Stories by Hana

    01
    Benchmarks

    Can AI Be a Senior Engineer? New Benchmark Says Yes

    The Senior SWE-Bench, a new open-source benchmark, rigorously assesses AI agents' capabilities as senior software engineers, evaluating complex coding, system design, and debugging skills beyond simple task completion.

    8 Minutes · Jul 26, 2026

    More from the AgentCrunch newsroom

    Maya OkaforRafael DuarteEllis ThornePriya RamanJonas Weber