Articles

GPT-5.6 joins the HUMAINE leaderboard: how Sol, Terra, and Luna rank with real people

Jasmehr Bhatia
|July 10, 2026

All rankings and win rates below reflect the HUMAINE leaderboard as of July 2026, with 54 models on the board. Results update as new data is collected, so check the live leaderboard for current standings.


On July 9, OpenAI released the GPT-5.6 family: Sol, built for frontier reasoning and long-horizon agentic work; Terra, a balanced model for everyday business tasks; and Luna, the fastest and lowest-cost tier. The launch arrived with strong benchmark results in coding and security work.

We have now evaluated all three with real people on HUMAINE, our demographically representative, multi-dimensional human preference leaderboard. Within a day of the launch, more than 2,000 UK and US participants had blind-tested the new models in nearly 2,900 head-to-head conversations. No labels, no logos, just conversations. Here is where they landed.



Where the family ranks

Terra debuts at #27 of 54, Luna at #36, and Sol at #38. Luna and Sol are statistically tied with each other.

Two things stand out immediately. First, the family enters behind OpenAI's own earlier models. None of the three catches GPT-5.2 Chat (#11), GPT-5.4 (#16), or GPT-5.5 (#19). In fitted head-to-heads against GPT-5.2 Chat, OpenAI's strongest model with our raters, Terra wins 43% of the time, Luna 36%, and Sol 33%. Terra is roughly level with the year-old GPT-4.1 at 50/50. What the family does clear is OpenAI's 2025 cohort: all three sit above o3 (#40), GPT-5 (#42), and the rest of the earlier lineup.

Second, this is an unusual debut pattern. Every previous major-lab cohort we have measured entered at or above the family's standing models. Last month's debuts included a model that entered 4th and another that entered 2nd.

Against the wider field, the top of the board does not move. Google's top three and Claude Fable 5 remain out of reach, with Fable 5 beating Terra in six of ten matchups. Terra also loses to most of June's debut cohort, including Gemini 3.5 Flash (38% win rate), Qwen3.7 Max (41%), and DeepSeek V4 Pro (45%). Where Terra does compete is the crowded middle of the board: a statistical dead heat with Kimi K2.6 (49%) and a narrow edge over Claude Opus 4.8 (52%), which was Anthropic's flagship until Claude Fable 5 arrived in June. Luna and Sol land among the previous Claude generation, between Claude Opus 4 and Claude Sonnet 4.

Here is how the three rank on each of HUMAINE's four dimensions:
 

Rank of 54

Terra

Luna

Sol

Task performance

#14

#33

#35

Trust and ethics

#19

#33

#34

Interaction fluidity

#23

#33

#39

Communication style

#29

#35

#38

Overall

#27

#36

#38

Trust and ethics separates models far less than the other dimensions, with most matchups on it close to a coin flip, so treat gaps on that row lightly.

A note on precision: new models carry the widest uncertainty on the board. Roughly 1,100 comparisons per model is enough to place each in a tier, not to fix an exact rank, so read Terra as somewhere in the twenties and Luna and Sol as mid-thirties to around forty. The ordering, Terra ahead with Luna and Sol tied, is solid.

Terra separates from its siblings on substance

Inside the family, Terra beats Luna 57% of the time and Sol 61% in fitted head-to-heads. Luna versus Sol is a coin flip.

The direct battles show why. In the 199 head-to-head conversations where Terra and Sol faced each other, raters who picked a side gave Terra 74% of the votes on task performance and 71% on interaction fluidity, but only 61% on communication style. The same shape appears against Luna: 69% on task performance, 56% on style. The siblings sound similar to raters and perform differently.

That profile is rare for a debut. Terra's strongest dimension is task performance, where it ranks #14 of 54. Its weakest is communication style, at #29. That 15-place spread is one of the most substance-tilted profiles on the board, and the mirror image of last month's Claude Fable 5 debut, which entered on charm with interaction fluidity at #1 and task performance at #10. Most models that climb the board on debut win on style first and substance later. In practice, this suggests Terra punches above its overall rank for practical, get-me-an-answer use and below it for open-ended conversation.

It is worth noting that this matches OpenAI's own positioning. Terra is the tier OpenAI built for high-volume everyday tasks, and it is the tier everyday users prefer.

Strong on benchmarks, mid-table with people

The GPT-5.6 family arrived with a formidable technical reputation. So why does it land mid-table when real people choose? Three findings from our data explain the gap.

People are not running benchmarks. HUMAINE conversations are everyday life. In our June classification of 116,000 conversations, about half were information seeking, another large share was personal advice and decision support, and fewer than 3% were technical assistance. A model's ability to write an exploit-proof function is simply never exercised by most raters.

Everyday questions saturate capability. Across the whole board, roughly four in ten task-performance votes end in a tie. When both models handle "help me plan this trip" competently, the task axis stops separating them and the decision shifts to how the conversation feels. That is where this family is weakest, with communication style ranks between #29 and #38, so its technical headroom has little room to show.

OpenAI's history on our board shows conversational tuning sets the rank. Every OpenAI model built primarily for reasoning and technical work has ranked in our bottom third: o3 (#40), GPT-5 (#42), o1 (#51), o3-mini (#53). The conversationally tuned GPT-5.2 Chat sits at #11. Same lab, same underlying capability trajectory, and a roughly 30-place gap decided by conversational tuning. The GPT-5.6 family debuts on the technical side of that divide.

None of this makes the board wrong about the family. It makes the board specific. HUMAINE measures what a representative sample of people prefer to talk to, and on that question the results above stand. For teams choosing a model for user-facing products, that distinction between benchmark capability and human preference is exactly the information a technical leaderboard cannot provide.

Honest, direct, and easy to drift away from

A separate automated content analysis of the family's 3,558 conversations adds a behavioral layer to the votes. These are machine-judged reads of conversation behavior, so treat them as corroboration rather than vote counts.

 

Judged message by message, the family looks exemplary. Terra ranks #2 of 54 on helpfulness, all three are top five on tone and on how trustworthy their assertions are, and Terra shows the best-judged confidence calibration on the board. All three are also the least flattering models we have ever measured, by a wide margin against the field average.
 

The deficits sit exactly where the votes went against the family. Engagement, meaning how well a model sustains a dialogue rather than closing it, lands in the bottom half for all three (Terra #25, Luna #35, Sol #36), and the family runs brief in a field that runs long. Terra, the one variant that pairs directness with completeness (#3 of 54), is the one that wins.

Raters described the trade in their own words. When the family won, the reasons were directness: "straight to the point, which I appreciated more." When it lost, the reasons were depth and warmth: "[Sol] gave ok answers, but were quite brief. [The other] was fantastic, it gave very in-depth answers."

For teams evaluating models, the pattern is worth sitting with. Zero flattery and careful calibration cost this family nothing with raters. The losses clustered around brevity and warmth instead, and only the variant that paired honesty with depth turned that behavior into wins.

The demographic picture

HUMAINE ranks models across 22 demographic groups in the US and UK, and the family's debut is not uniform across them.

Terra and Sol both rank noticeably higher with US raters than UK ones. Terra averages #24 in the US against #27 in the UK across our demographic groups; Sol averages #36 against #39. Luna is the exception, slightly warmer in the UK than the US, and the most consistent of the three across every group we track.

Terra's brightest results are demographic. Every US audience we track already places it ahead of Claude Opus 4.8, and early data suggests three US audiences, led by 18 to 34 year olds, place it ahead of OpenAI's own GPT-5.5. These are day-one reads with wide error bars, but the direction is consistent: if there is an early-adopter audience for this family, it is young, American, and it has picked Terra.

Explore the results

HUMAINE evaluates models through blind head-to-head conversations with a stratified, verified participant pool, scored across task performance, communication style, interaction fluidity, and trust and ethics, then ranked with a hierarchical Bradley-Terry-Davidson model. The framework was published as a conference paper at ICLR 2026.

You can explore the full results, including head-to-head comparisons and demographic breakdowns, on the interactive leaderboard, or read more about how HUMAINE works.