Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

(terminal-bench-science.ai)

12 points | by matt_d 56 minutes ago

1 comments

  • rubslopes 39 minutes ago
    I'm glad to see that GPT Sol beats Opus at least in Mathematical Sciences, because that's my need right now, and I much prefer GPT's prose style.