| {{ column.title }} | Время |
|---|---|
| {{ formatValue(getColumnValue(submit, column)) }} | {{ formatValue(getTotalTime(submit)) }} |
| Задача | Результат | Метрика |
|---|---|---|
| SAGE | 0.684 / 0.691 / 0.688 |
PUNCT_F1
SPELL_F1
ERRANT_F1
|
| Humor | 0.282 / 0.301 |
Exact match
LLM judge score
|
| LIMUR | 0.626 / 0.629 |
Exact match
LLM judge score
|
| RUBIN | 0.697 / 0.704 |
Exact match
LLM judge score
|
| Riddles | 0.492 / 0.511 |
Exact match
LLM judge score
|
| SOB-Hard | 0.116 / 0.218 / 0.744 / 0.184 / 0.223 / 0.751 |
balance_score
task_pass_rate
format_pass_rate
sample_pass_rate
content_pass_rate
constraint_pass_rate
|
| Characters | 0.523 / 0.578 |
Exact match
LLM judge score
|
| Enantiosemy | 0.348 / 0.514 |
Exact match
LLM judge score
|
| GorillaHard | 0.467 / 0.8 / 0.329 / 0.55 / 0.567 / 0.857 / 0.324 / 0.402 / 1 / 0.052 / 0.982 / 0.865 / 0.074 / 1 |
balance_score
clarify_recall
args_match_rate
tool_match_rate
dialog_pass_rate
format_pass_rate
sample_pass_rate
abstention_recall
cost_optimal_rate
false_clarify_rate
constraint_pass_rate
tool_in_catalog_rate
false_abstention_rate
injection_resistance_rate
|
| IFHardBench | 0.846 / 0.467 / 0.821 |
balance_score
sample_pass_rate
constraint_pass_rate
|
| NewReasoning | 0.6 / 0.639 |
Exact match
LLM judge score
|
| RussianRegions | 0.184 / 0.658 / 0 / 0.368 |
Exact match
LLM judge score
Group Exact match
Group Judge Score
|
GigaChat
GigaChat3.5-432B-A28B-Reasoning
432.0B
GigaChat 3.5 Reasoning is the first GigaChat model with full reasoning trained with online RL. Compared with GigaChat 3.5 Ultra Instruct, the largest gains are in mathematics, code, instruction following, and structured output. It uses a custom hybrid architecture that combines Multi-head Latent Attention (MLA) with GatedDeltaNet linear-attention layers.
MIT