Grok 4.6

MERA Created at 18.08.2026 17:01
0.623
The overall result
4
Place in the rating

In the top by tasks:

1 Enantiosemy
1 GorillaHard
2 SOB-Hard
2 IFHardBench
3 SAGE
3 LIMUR
3 Riddles
4 Characters
5 Humor
5 RUBIN
6 RussianRegions
10 NewReasoning
Model, team
Result
{{ formatValue(getLeaderboardPlace(submit)) }}
{{ submit.name }} {{ transformSize(submit.model_size) }}
{{ submit.team_name }}
{{ formatValue(getResultValue(submit)) }}
{{ column.title }} Time
{{ formatValue(getColumnValue(submit, column)) }} {{ formatValue(getTotalTime(submit)) }}

All tasks

Task name Result Metric
SAGE 0.74 / 0.71 / 0.725
PUNCT_F1 SPELL_F1 ERRANT_F1
Humor 0.455 / 0.478
Exact match LLM judge score
LIMUR 0.766 / 0.768
Exact match LLM judge score
RUBIN 0.757 / 0.757
Exact match LLM judge score
Riddles 0.654 / 0.678
Exact match LLM judge score
SOB-Hard 0.926 / 0.928 / 0.995 / 0.928 / 0.933 / 0.985
balance_score task_pass_rate format_pass_rate sample_pass_rate content_pass_rate constraint_pass_rate
Characters 0.627 / 0.628
Exact match LLM judge score
Enantiosemy 0.666 / 0.753
Exact match LLM judge score
GorillaHard 0.788 / 1 / 0.743 / 0.824 / 0.967 / 1 / 0.737 / 0.661 / 1 / 0 / 1 / 0.996 / 0.004 / 1
balance_score clarify_recall args_match_rate tool_match_rate dialog_pass_rate format_pass_rate sample_pass_rate abstention_recall cost_optimal_rate false_clarify_rate constraint_pass_rate tool_in_catalog_rate false_abstention_rate injection_resistance_rate
IFHardBench 0.952 / 0.82 / 0.953
balance_score sample_pass_rate constraint_pass_rate
NewReasoning 0.659 / 0.703
Exact match LLM judge score
RussianRegions 0.167 / 0.658 / 0 / 0.323
Exact match LLM judge score Group Exact match Group Judge Score

Information about the submission

Team:

MERA

Name of the ML model:

Grok 4.6

Link to the ML model:

https://x.ai/news/grok-4-6

Model type:

Closed