Programming Language Benchmarks

Self-invoking code benchmarks help you decide which LLMs to use for your programming tasks

As large language models (LLMs) continue to improve at coding, the benchmarks used to evaluate their performance are steadily becoming less useful. That's because though many LLMs have similar high ...

Communications of the ACM

Measuring What Matters in Large Language Model Performance

As large language models (LLMs) gain momentum worldwide, there’s a growing need for reliable ways to measure their performance. Benchmarks that evaluate LLM outputs allow developers to track ...

Some results have been hidden because they may be inaccessible to you

Show inaccessible results

Self-invoking code benchmarks help you decide which LLMs to use for your programming tasks

Measuring What Matters in Large Language Model Performance

Trending now