To back
01 Sep 2026

AI Alliance Launches MERA Text 2.0, a New Leaderboard for Evaluating Text Language Models

The first version of MERA was launched in 2023, when existing benchmarks still provided strong differentiation between models across core capabilities such as world knowledge, logic, text comprehension, and domain-specific knowledge. Over the past three years, language models have advanced significantly. On several of the previous benchmarks, modern SOTA models now reach or exceed the Human Baseline, while the gap between the strongest models is becoming increasingly narrow.

For MERA Text 2.0, the team therefore chose not simply to expand the existing task set, but to redesign the text evaluation around capabilities that still provide meaningful insight into the limits of modern models.

The update introduces 12 new private tests grouped into four categories: Human-Centric, Culture-Specific, Agentic, and Reasoning.

Human-Centric evaluates capabilities related to understanding people and communication, including metaphors, humor, character traits, and conversational context. The category includes Riddles, Humor, and Characters.

Culture-Specific focuses on the Russian language and cultural context. RussianRegions evaluates the understanding of regional vocabulary; SAGE measures the ability to correct real-world Russian texts while preserving their meaning and style; and RUBIN evaluates knowledge of Russian cultural context, including memes, slang, cinema, songs, folklore, and idiomatic expressions.

Agentic focuses on the foundational model capabilities required for building agentic systems. IFHardBench evaluates complex instruction following under multiple simultaneous constraints, SOBHard tests structured output and document transformation, and GorillaHard assesses tool selection and the correct construction of tool calls.

Reasoning brings together tasks that require models to handle ambiguity and reason over the structure of a problem. Enantiosemy evaluates the understanding of words and expressions that can carry opposite meanings depending on context. NewReasoning covers different types of logical reasoning and robustness to familiar solution patterns, while LIMUR evaluates reference resolution as well as syntactic and linguistic ambiguity.

All MERA Text 2.0 test sets are private: models do not have access to the evaluation examples in advance. This approach helps reduce the risk of contamination and preserve the benchmark’s discriminative power over time.

Results are calculated separately for each of the four capability categories. Within each category, task scores are aggregated using a geometric mean. This makes the final score more sensitive to weaknesses in individual capabilities and prevents strong performance on one task from fully compensating for poor performance on another.

All models are evaluated using the same fixed protocol. MERA does not use model-specific prompt optimization or post-processing, while evaluation parameters and logs are retained to ensure reproducibility.

Text 2.0 also focuses primarily on the capabilities of the model itself. Additional agentic, reasoning, or tool-use pipelines that could solve part of the task on the model’s behalf are not included in the evaluation. Full agentic and environment-based scenarios will be covered by separate directions within the MERA ecosystem.

At launch, the team has already evaluated a number of current open and proprietary models. Their results are available on the leaderboard, including separate breakdowns across the four capability categories.

MERA Text 2.0 will continue to evolve. The leaderboard architecture allows new tests to be added to existing categories and new evaluation categories to be introduced as language model capabilities develop.