LiveBench Mostly Measures One General Ability

LiveBench is an LLM benchmark with tasks grouped into categories. I’ll analyze its April 7, 2025 public release, which included 18 tasks across six categories. Item-level data were available for seven tasks across three categories: Coding (LCB Generation and Coding Completion), Language (Connections, Plot Unscrambling, and Typos), and Instruction Following (Paraphrase and Story Generation). Unfortunately, I can’t analyze the other tasks because LiveBench didn’t release their item-level data, which is essential for this kind of analysis.

Here’s what each task involves:

LCB Generation (Category: Coding)

The model receives a programming problem (typically from LiveCodeBench) and must write a complete solution which is then run against test cases. The model receives a score of 1 if it passes and 0 if it fails; there is no partial credit.

Coding Completion (Category: Coding)

The model receives a programming problem and a fragment of a correct solution, which it must complete. Scoring is the same as for LCB Generation.

Connections (Category: Language)

The Connections task works much like the NYT game of the same name. The model receives a shuffled list of words and must sort them into groups of four, with each group sharing a theme. For example:

LiveBench uses 8 words (2 groups), 12 words (3 groups), or 16 words (4 groups). The score is the fraction of complete groups identified correctly.

Plot Unscrambling (Category: Language)

The model receives the sentences of a recent movie synopsis in random order and must reconstruct the original narrative. The evaluator fuzzy-matches the response to the original sentences, allowing minor transcription changes, then measures the edit distance between the proposed and correct orders:

\[\text{score} = 1 - \frac{d}{n}\]

where $d$ is the ordering distance and $n$ is the number of sentences.

Typos (Category: Language)

The model receives text (usually based on a recent arXiv abstract) with synthetic spelling errors inserted. It must correct the misspellings while leaving everything else unchanged. That means it shouldn’t rewrite the text, change punctuation, change US spelling to UK spelling or vice versa, or add stylistic “improvements”. The scorer gives 1 if the ground-truth text appears anywhere in the output and 0 otherwise. Thus, despite the instruction, extra surrounding text does not necessarily cause a failure.

Paraphrase (Category: Instruction Following)

The model receives the beginning of a recent Guardian article and is asked to paraphrase it while following several mechanically verifiable instructions. For example:

Paraphrase this article. Include a title. Use the words “course,” “media,” and “sun.” Write exactly three paragraphs. Begin the first paragraph with “hand.”

LiveBench scores compliance with the explicit instructions, not the quality of the paraphrase. In fact, it doesn’t care about the actual paraphrase at all; models can receive full credit without ever attempting to paraphrase the article. It averages two components:

Story Generation (Category: Instruction Following)

The model receives a recent news article and is asked to generate a story based on it while following mechanically verifiable instructions. LiveBench uses the same two-component scoring method as Paraphrase.

The seven tasks report different kinds of item scores. Here is how LiveBench’s published scores relate to the responses I use in this analysis:

Task LiveBench’s published item score Response used in this analysis
LCB Generation Pass/fail: 0 or 1 Binary, unchanged
Coding Completion Pass/fail: 0 or 1 Binary, unchanged
Connections Fraction of complete four-word groups identified correctly; an item has two, three, or four groups Ordered partial credit, unchanged
Typos Exact-match pass/fail: 0 or 1 Binary, unchanged
Plot Unscrambling One minus ordering edit distance divided by the number of sentences; bounded between 0 and 1 Count-adjusted logit of the score, treated as continuous
Paraphrase Average of all-instructions-correct accuracy and the fraction of individual instructions followed Instruction-level fraction only, modeled as ordered partial credit
Story Generation Same two-component score as Paraphrase Instruction-level fraction only, modeled as ordered partial credit

I don’t like how Paraphrase and Story Generation are currently graded. Their published scores average instruction-level accuracy with an all-or-nothing prompt-level component, so missing just one instruction costs more than half the grade. I therefore use instruction-level accuracy alone.

For Plot Unscrambling, I logit-transform the score using this count-adjusted formula:

\[\text{transformed score} = \log\left(\frac{n - D + \frac{1}{2}}{D + \frac{1}{2}}\right)\]

where $D$ is the edit distance and $n$ is the number of sentences.

Determining the Number of Factors

Naturally, I used parallel analysis to determine the number of factors. It yielded 19, which is a lot (see the plot).

However, some item pairs have no models in common, and others have only a few, making their correlations unavailable or imprecise. So I modeled the correlation matrix Bayesianly and repeated parallel analysis across posterior draws to see how much the recommended factor count varies. Each of the first 12 factors exceeds the chance threshold in at least 95% of posterior draws, and the 90% interval for the number retained is 12–13.

Even 12–13 factors is a lot for seven tasks. I would have expected something closer to seven, so let’s look at the tasks individually to see where the extra dimensions might be coming from.

Select a task to see its Bayesian parallel analysis. The badge counts leading factors above the chance threshold in at least 95% of posterior draws.

LCB Generation2 factors
Bayesian parallel analysis for LCB Generation
Coding Completion2 factors
Bayesian parallel analysis for Coding Completion
Connections3 factors
Bayesian parallel analysis for Connections
Plot Unscrambling2 factors
Bayesian parallel analysis for Plot Unscrambling
Typos7 factors
Bayesian parallel analysis for Typos
Paraphrase2 factors
Bayesian parallel analysis for Paraphrase
Story Generation2 factors
Bayesian parallel analysis for Story Generation

Although the Bayesian parallel analyses support more than one factor for every task, the first dimension dominates in six of them. The ratio of the first two (posterior-median) eigenvalues ranges from 3.1 to 10.6 for those tasks, compared with just 1.5 for Typos. The within-task item-correlation matrices offer another way to see this:

Items are ordered by median task-factor loading. Grey means unavailable, not zero. Scroll to compare tasks; select a matrix for full size.

Most of the matrices suggest a clear positive manifold. Typos is the exception: it still looks quite ugly after the items are sorted by loading. That makes me want to check whether the Typos items themselves are sound. Auditing every item would take too long, but CTT/IRT measures can flag suspicious ones for us to focus on…

Item Flags

In my post about flagging suspicious questions in AI benchmarks, I discussed flags based on item discrimination and distractor behavior. None of the items here is multiple choice, so there are no distractors to examine. That leaves item discrimination, which I measure using the corrected item–task score correlation: the correlation between an item and its task score calculated without that item.

Red: below 0 · Yellow: 0–0.2 · Green: above 0.2 · Grey: undefined. Choose a task, then select its histogram for full size.

Histogram of corrected item–task correlations for Typos
Histogram of corrected item–task correlations for LCB Generation
Histogram of corrected item–task correlations for Coding Completion
Histogram of corrected item–task correlations for Connections
Histogram of corrected item–task correlations for Plot Unscrambling
Histogram of corrected item–task correlations for Paraphrase
Histogram of corrected item–task correlations for Story Generation

All 494 items are counted. Percentages use each task's total, including undefined correlations, and are rounded to whole numbers; rows may not sum to 100%.

TaskItemsRed (< 0)Yellow (0–0.2)Green (> 0.2)Undefined
Typos1003 (3%)15 (15%)80 (80%)2 (2%)
LCB Generation780 (0%)1 (1%)72 (92%)5 (6%)
Coding Completion500 (0%)2 (4%)48 (96%)0 (0%)
Connections1000 (0%)2 (2%)98 (98%)0 (0%)
Plot Unscrambling900 (0%)0 (0%)90 (100%)0 (0%)
Paraphrase500 (0%)3 (6%)47 (94%)0 (0%)
Story Generation261 (4%)3 (12%)22 (85%)0 (0%)
All tasks4944 (1%)26 (5%)457 (93%)7 (1%)

The flags uncovered two Typos items with valid alternative answers, three LCB Generation items with grading problems, and three Story Generation items with prompt or checker problems. The item-by-item audit gives the archived-answer counts and rescoring checks. I could inspect only a subset of items and models, so other problems may remain.

Factor Structure

Time to look at the factor structure. I’ll exclude the items I found problems with: 832610e9 and 0becbf34 from Typos; “Wrong Answer,” “Takahashi Quest,” and “Bad Juice” from LCB Generation; and 0d828b10, 6c5eb0ac, and 230fffb5 from Story Generation. Other problematic items may remain because I couldn’t inspect them. As noted above, every task except Typos is strongly unidimensional. A separate factor analysis of Typos produced factors that were hard to interpret, so for simplicity I’ll model each task, including Typos, with a single factor.

Seven task factors each load onto their own item responses, and a heptagram-like network connects every pair of task factors with one of 21 freely estimated correlations.

The item loadings in the correlated-task model look like this:

Item loading distributions by task in the correlated-task model; points show posterior medians and vertical bars show 90% intervals, with zero-crossing intervals distinguished from positive intervals.

Fourteen items have 90% loading intervals that include zero. That doesn’t mean they’re flawed, but it does make them worth a closer look. I’ve put my notes in a collapsible section so they don’t interrupt the main discussion.

Items whose 90% loading intervals include zero (14)

Typos

Paraphrase

Story Generation

The task-factor correlations look like this:

Posterior correlations among the seven task factors; each cell shows the median correlation and its 90% interval on a vivid red-to-green scale.

There’s a clear positive manifold, and parallel analysis of the task-factor correlation matrix supports a single factor. This would imply a hierarchical model in which a higher-order factor explains the correlation among tasks. However, since LiveBench groups tasks into categories, we might instead add Coding, Language, and Instruction Following domain factors. Those domains could correlate freely or load on an even higher-order factor.

Three proposed factor structures: one general factor above all seven tasks; three freely correlated domain factors above their respective tasks; and one grand factor above the three domains, which in turn explain their respective tasks.

Unfortunately, each of these models fits worse than the correlated-task model (see the initial comparison).

The domain estimates help explain why. In the correlated-domains model, the domain correlations are quite high:

Posterior correlations among the Coding, Language, and Instruction Following domains, with medians and 90% intervals.

Even so, they can’t account for some of the task-factor correlations, as we’ll see below. In the grand-factor model, all three domain loadings are very close to 1; the lowest loading is 0.9993. That leaves little domain-specific variance, so modeling the categories doesn’t seem to add much.

That brings us back to the hierarchical model. It also fits worse than the correlated-task model, but I find it more plausible a priori. The positive manifold is what I’d expect from LLMs, and parallel analysis of the task-factor correlations supports one common factor. Its poorer fit suggests that the general factor alone misses some relationships between tasks. We can allow for those relationships by adding correlated residuals, so that selected task factors can correlate more than the general factor predicts.

To see which links might be worth including, I fit an exploratory hierarchical model with positive-only shrinkage priors on all 21 task-residual correlations:

Positive-only shrinkage fit: residual correlations among all seven task factors, with posterior medians and 90% intervals.

The largest estimated residual correlations are Plot Unscrambling–Typos (+.52), LCB Generation–Coding Completion (+.29), and Connections–Story Generation (+.27). But Coding Completion’s loading on the general factor is almost one in this exploratory fit, leaving virtually no task-specific variance. Its +.29 residual correlation therefore adds only about +.002 to the implied correlation between the two tasks. I’ll include Plot–Typos and Connections–Story, but not LCB–Coding. Because the exploratory prior rules out negative residual correlations, an interval above zero is not, by itself, a reason to include a link.

In the final hierarchical model, only those two residual correlations are estimated, with priors that allow either sign. All other residual correlations are fixed at zero. The posterior estimates are:

Task pair Median residual correlation 90% interval
Plot Unscrambling–Typos +.57 [+.52, +.61]
Connections–Story Generation +.35 [+.19, +.48]
One general factor loads on seven task factors; dashed arcs show the only two additional task-residual correlations, Connections–Story Generation and Plot Unscrambling–Typos.

Comparing the new model with the previous ones:

WAIC expected log predictive density differences from correlated task factors, including the hierarchy with Plot–Typos and Connections–Story residual links; bars show one paired pointwise standard error.

The two residual correlations improve the hierarchical model’s fit.

The loadings of the seven tasks on the general factor in this fit are:

Task factor Median loading 90% interval
LCB Generation .91 [.90, .92]
Coding Completion 1.001 [1.00, 1.00]
Connections .82 [.80, .83]
Plot Unscrambling .70 [.69, .71]
Typos .79 [.77, .81]
Paraphrase .77 [.74, .80]
Story Generation .86 [.82, .89]

We can also ask how much of each task’s total-score variance is attributable to the general factor, its task-specific factor, or item-specific variation. This decomposition is for an equal-weighted sum of underlying item responses, not the observed mixed-format LiveBench score:

Stacked bars for each task showing the posterior mean percentages of latent total-score variance attributable to the general factor, task-specific factor, and item-specific variation.

Correlations with the ECI

This compares Epoch’s ECI with posterior-mean factor scores from the selected hierarchical fit.2 The general factor correlates strongly with ECI. The task factors do too, though much of that correlation appears to come from their shared general component. Once that component is removed, only the Plot Unscrambling and Paraphrase residuals have 90% intervals entirely above zero.

Pearson correlations with ECI for the general factor and, for each of seven tasks, the full task factor and its task-specific residual, with 90% bootstrap intervals and matched-model counts.

Blue: full task factor · Orange: task residual after removing the general factor · General factor at left. Bars show 90% bootstrap intervals; counts include direct task responses only.

Takeaways

Once again, benchmark item flags proved useful for finding problems. The task factors display a positive manifold, as expected. What’s more notable is the lack of clear domain factors beyond the general factor. Human cognitive ability is well modeled by g, but not perfectly: someone may be better at spatial tasks, and someone else better at verbal tasks, than their levels of g would predict. They could have the same g, yet if you needed to navigate an unfamiliar city or write an essay, you might have a clear choice between them.

That distinction is much less apparent for the AI models and tasks tested here. LiveBench divides its tasks into categories, but I find little evidence that these categories capture distinct abilities. It doesn’t seem especially useful to say “use Model A for Coding and Model B for Language” when performance across those domains is so closely tied to general performance. Individual tasks can still differ: Plot Unscrambling and Typos, for example, are more closely related than the general factor alone predicts. But LCB Generation and Coding Completion do not show much extra association, despite both involving coding. The distinctions worth paying attention to seem to lie with particular tasks, not the broad category labels.

Appendix

Classical Parallel Analysis

Flagged-item audit

Looking at the flagged items, along with items that have constant scores across models:

Typos

LCB Generation

Coding Completion

Connections

Plot Unscrambling (No flagged items)

Paraphrase

Story Generation

I can inspect archived answers for only a subset of items and models, so I cannot determine the full scope of these problems. Nonetheless, I’ve tried my best doing what I can do.

Initial Model Comparison

WAIC expected log predictive density differences from the correlated-task reference for the single higher-order factor, correlated domains, and grand-factor-over-domains models, with bars showing one paired standard error.

Loadings vs Difficulties

These plots place each item’s task-factor loading against its estimated difficulty in the selected hierarchical fit.

Choose a task, then select its plot for full size and 90% intervals. Difficulty uses a task-specific response scale, so compare horizontal positions only within a task.

LCB Generation item loadings versus difficulty
Coding Completion item loadings versus difficulty
Connections item loadings versus difficulty
Plot Unscrambling item loadings versus difficulty
Typos item loadings versus difficulty
Paraphrase item loadings versus difficulty
Story Generation item loadings versus difficulty

Task Information Curves

The seven task-factor information curves are overlaid on the same axes, so their heights are directly comparable.

Select a task or its curve to highlight it. Hover, focus, or tap a model marker for its name, median task score, 90% interval, and information at that score. The full-size plots also show individual item curves.

Seven overlaid task-factor information curves on shared score and information axes
Loading interactive model markers…
  1. All values in the task-loading table are rounded to two decimal places. Coding Completion’s displayed 1.00 values are slightly below 1 before rounding. ↩

  2. For this figure, I include ECI models with exactly one distinct model version in their benchmark records and ECI scores dated no later than April 7, 2025. I match version names to LiveBench after ignoring case and punctuation, but not version numbers or words. If several LiveBench runs match the same ECI model, I use the one with the most item responses. This leaves 38 ECI models before task-specific response requirements. ↩