Gemini 3.7 Flash

MERA Created at 18.08.2026 16:42
0.678
The overall result
1
Place in the rating

In the top by tasks:

1 SAGE
1 Humor
1 RUBIN
1 Riddles
1 Characters
1 IFHardBench
1 NewReasoning
1 RussianRegions
2 LIMUR
3 SOB-Hard
3 GorillaHard
4 Enantiosemy
Model, team
Result
{{ formatValue(getLeaderboardPlace(submit)) }}
{{ submit.name }} {{ transformSize(submit.model_size) }}
{{ submit.team_name }}
{{ formatValue(getResultValue(submit)) }}
{{ column.title }} Time
{{ formatValue(getColumnValue(submit, column)) }} {{ formatValue(getTotalTime(submit)) }}

All tasks

Task name Result Metric
SAGE 0.801 / 0.77 / 0.786
PUNCT_F1 SPELL_F1 ERRANT_F1
Humor 0.58 / 0.592
Exact match LLM judge score
LIMUR 0.827 / 0.83
Exact match LLM judge score
RUBIN 0.884 / 0.884
Exact match LLM judge score
Riddles 0.706 / 0.732
Exact match LLM judge score
SOB-Hard 0.868 / 0.884 / 0.931 / 0.882 / 0.886 / 0.928
balance_score task_pass_rate format_pass_rate sample_pass_rate content_pass_rate constraint_pass_rate
Characters 0.641 / 0.645
Exact match LLM judge score
Enantiosemy 0.563 / 0.697
Exact match LLM judge score
GorillaHard 0.609 / 1 / 0.764 / 0.857 / 0.9 / 1 / 0.763 / 0.74 / 1 / 0 / 1 / 0.987 / 0.013 / 1
balance_score clarify_recall args_match_rate tool_match_rate dialog_pass_rate format_pass_rate sample_pass_rate abstention_recall cost_optimal_rate false_clarify_rate constraint_pass_rate tool_in_catalog_rate false_abstention_rate injection_resistance_rate
IFHardBench 0.979 / 0.937 / 0.982
balance_score sample_pass_rate constraint_pass_rate
NewReasoning 0.811 / 0.842
Exact match LLM judge score
RussianRegions 0.282 / 0.891 / 0 / 0.711
Exact match LLM judge score Group Exact match Group Judge Score

Information about the submission

Team:

MERA

Name of the ML model:

Gemini 3.7 Flash

Model type:

Closed