Claude Opus 5

MERA Created at 18.08.2026 16:37
0.631
The overall result
3
Place in the rating

In the top by tasks:

1 LIMUR
2 SAGE
2 RUBIN
2 Characters
2 GorillaHard
2 RussianRegions
3 Humor
3 Enantiosemy
3 NewReasoning
4 SOB-Hard
4 IFHardBench
5 Riddles
Model, team
Result
{{ formatValue(getLeaderboardPlace(submit)) }}
{{ submit.name }} {{ transformSize(submit.model_size) }}
{{ submit.team_name }}
{{ formatValue(getResultValue(submit)) }}
{{ column.title }} Time
{{ formatValue(getColumnValue(submit, column)) }} {{ formatValue(getTotalTime(submit)) }}

All tasks

Task name Result Metric
SAGE 0.713 / 0.767 / 0.74
PUNCT_F1 SPELL_F1 ERRANT_F1
Humor 0.45 / 0.533
Exact match LLM judge score
LIMUR 0.83 / 0.844
Exact match LLM judge score
RUBIN 0.862 / 0.862
Exact match LLM judge score
Riddles 0.466 / 0.751
Exact match LLM judge score
SOB-Hard 0.762 / 0.793 / 0.855 / 0.79 / 0.798 / 0.844
balance_score task_pass_rate format_pass_rate sample_pass_rate content_pass_rate constraint_pass_rate
Characters 0.597 / 0.667
Exact match LLM judge score
Enantiosemy 0.626 / 0.759
Exact match LLM judge score
GorillaHard 0.689 / 1 / 0.677 / 0.79 / 0.733 / 0.949 / 0.677 / 0.646 / 1 / 0.002 / 0.966 / 0.959 / 0.001 / 1
balance_score clarify_recall args_match_rate tool_match_rate dialog_pass_rate format_pass_rate sample_pass_rate abstention_recall cost_optimal_rate false_clarify_rate constraint_pass_rate tool_in_catalog_rate false_abstention_rate injection_resistance_rate
IFHardBench 0.903 / 0.683 / 0.887
balance_score sample_pass_rate constraint_pass_rate
NewReasoning 0.751 / 0.784
Exact match LLM judge score
RussianRegions 0.255 / 0.816 / 0 / 0.567
Exact match LLM judge score Group Exact match Group Judge Score

Information about the submission

Team:

MERA

Name of the ML model:

Claude Opus 5

Model type:

Closed