GPT 5.6 Sol

MERA Создан 18.08.2026 16:58
0.639
Общий результат
2
Место в рейтинге

В топе по задачам:

1 SOB-Hard
2 Humor
2 Riddles
2 Enantiosemy
2 NewReasoning
3 RUBIN
3 Characters
3 IFHardBench
3 RussianRegions
4 LIMUR
4 GorillaHard
5 SAGE
Модель, команда
Результат
{{ formatValue(getLeaderboardPlace(submit)) }}
{{ submit.name }} {{ transformSize(submit.model_size) }}
{{ submit.team_name }}
{{ formatValue(getResultValue(submit)) }}
{{ column.title }} Время
{{ formatValue(getColumnValue(submit, column)) }} {{ formatValue(getTotalTime(submit)) }}

Все задачи

Задача Результат Метрика
SAGE 0.659 / 0.704 / 0.681
PUNCT_F1 SPELL_F1 ERRANT_F1
Humor 0.497 / 0.516
Exact match LLM judge score
LIMUR 0.749 / 0.754
Exact match LLM judge score
RUBIN 0.854 / 0.854
Exact match LLM judge score
Riddles 0.67 / 0.705
Exact match LLM judge score
SOB-Hard 0.938 / 0.942 / 1 / 0.941 / 0.944 / 0.988
balance_score task_pass_rate format_pass_rate sample_pass_rate content_pass_rate constraint_pass_rate
Characters 0.627 / 0.627
Exact match LLM judge score
Enantiosemy 0.658 / 0.753
Exact match LLM judge score
GorillaHard 0.59 / 1 / 0.702 / 0.865 / 0.867 / 0.999 / 0.691 / 0.583 / 1 / 0.004 / 0.999 / 0.986 / 0.009 / 1
balance_score clarify_recall args_match_rate tool_match_rate dialog_pass_rate format_pass_rate sample_pass_rate abstention_recall cost_optimal_rate false_clarify_rate constraint_pass_rate tool_in_catalog_rate false_abstention_rate injection_resistance_rate
IFHardBench 0.956 / 0.779 / 0.944
balance_score sample_pass_rate constraint_pass_rate
NewReasoning 0.751 / 0.782
Exact match LLM judge score
RussianRegions 0.242 / 0.775 / 0 / 0.478
Exact match LLM judge score Group Exact match Group Judge Score

Информация о сабмите

Команда:

MERA

Название ML-модели:

GPT 5.6 Sol

Тип модели:

Закрытая