BerryLM-XL

Wildberries & Russ Created at 17.04.2026 12:58
0.835
The overall result
3
Place in the rating

Ratings for leaderboard tasks

The table will scroll to the left

Task name Result Metric
LCS 0.818 Accuracy
RCB 0.603 / 0.604 Accuracy F1 macro
USE 0.632 Grade norm
RWSD 0.873 Accuracy
PARus 0.866 Accuracy
ruTiE 0.954 Accuracy
MultiQ 0.627 / 0.424 F1 Exact match
CheGeKa 0.57 / 0.483 F1 Exact match
ruModAr 0.998 Exact match
MaMuRAMu 0.908 Accuracy
ruMultiAr 0.998 Exact match
ruCodeEval 0.822 / 0.897 / 0.909 Pass@k
MathLogicQA 0.993 Accuracy
ruWorldTree 0.994 / 0.994 Accuracy F1 macro
ruOpenBookQA 0.96 / 0.96 Accuracy F1 macro

Evaluation on open tasks:

Go to the ratings by subcategory

The table will scroll to the left

Task name Result Metric
BPS 1.0 Accuracy
ruMMLU 0.888 Accuracy
SimpleAr 0.997 Exact match
ruHumanEval 0.837 / 0.868 / 0.872 Pass@k
ruHHH 0.91
ruHateSpeech 0.94
ruDetox 0.347
ruEthics
Correct God Ethical
Virtue 0.426 0.426 0.714
Law 0.472 0.423 0.694
Moral 0.475 0.456 0.762
Justice 0.4 0.396 0.636
Utilitarianism 0.364 0.383 0.597

Information about the submission

Mera version
v1.2.0
Torch Version
2.9.0
The version of the codebase
0cabf9d
CUDA version
12.8
Precision of the model weights
fp8
Seed
1234
Batch
1
Transformers version
4.57.6
The number of GPUs and their type
1 x NVIDIA A800 80GB PCIe
Architecture
local-chat-completions

Team:

Wildberries & Russ

Name of the ML model:

BerryLM-XL

Model size

358.0B

Model type:

Closed

SFT

MoE

Architecture description:

reasoning-ориентированная модель, дообученная для сценариев с длинным контекстом, раздельной генерацией внутреннего рассуждения и финального ответа, а также повышенными требованиями к качеству финального текстового вывода.

Description of the training:

В post-training используется DAPO - один из вариантов семейства GRPO, адаптированный под сценарии, где нужно одновременно контролировать качество ответа и форму reasoning-процесса. В рабочем контуре используется компактная система из двух reward-функций с весами 0.8 / 0.2, собранная вокруг задачи предотвращения reward hacking. Особое внимание уделено multi-turn агентским сценариями

Pretrain data:

Обучение проводится на миксе закрытых и открытых датасетов. Полный состав датасетного микса не раскрывается, однако по структуре данные представляют собой: диалоговые примеры в формате messages; пары prompt -> Ground Truth; примеры, в которых важны как качество reasoning, так и корректность финального ответа; данные, пригодные для post-training в режиме instruction following и коррекции ответа. Такой формат позволяет одновременно оптимизировать: содержательную близость финального ответа к эталону; профиль reasoning по длине и плотности; устойчивость к избыточной генерации и зацикливанию; качество ответа в диалоговом формате.

License:

Proprietary

Inference parameters

Generation Parameters:
simplear - do_sample=false;until=['<|im_end|>']; \nchegeka - do_sample=false;until=['<|im_end|>']; \nrudetox - do_sample=false;until=['<|im_end|>']; \nrumultiar - do_sample=false;until=['<|im_end|>']; \nuse - do_sample=false;until=['<|im_end|>']; \nmultiq - do_sample=false;until=['<|im_end|>']; \nrumodar - do_sample=false;until=['<|im_end|>']; \nruhumaneval - do_sample=true;temperature=0.6;until=["<|im_end|>"]; \nrucodeeval - do_sample=true;temperature=0.6;until=["<|im_end|>"];

The size of the context:
202000

System prompt:
## System Instructions \nFollow all these rules strictly \n \n### Thinking Rules \nReasoning effort: Medium \nТы можешь рассуждать, анализировать и думать в любом формате, но КРАТКО. Надо получить лучший ответ за минимальное количество шагов. \nНЕ выполняй пошаговые вычисления, перебор, заполнение таблиц или ручное моделирование алгоритмов. \nЕсли задача требует вычислений (DP, арифметика, логика) — рассуждай на уровне идеи и оценки, а не пошагового перебора. \nРассуждения должны занимать НЕ БОЛЕЕ нескольких абзацев. Затем СРАЗУ дай финальный ответ. \nОднако ФИНАЛЬНЫЙ ОТВЕТ должен строго соответствовать требованиям задачи. \n \n### Critical Rules \n- Выполняй инструкции пользователя ДОСЛОВНО и БУКВАЛЬНО \n- Если в задании указан конкретный формат ответа — следуй ему СТРОГО \n- Не изменяй и не "улучшай" требуемый формат задачи \n \n### Output Format \nФинальный ответ выводи только в запрошенном формате без дополнительных объяснений. \nЕсли нужна буква - пиши только букву. Если нужна цифра - пиши только цифру. Если нужно число - пиши только число. \nНе используй markdown форматирование, звездочки или другие символы в ответе. \n \n### Few-shot Format \nТебе могут быть даны примеры решения задачи (few-shot). Примеры показывают формат ответа. \nВАЖНО: Иногда сама задача, которую нужно решить, уже содержится среди примеров — в этом случае ты должен дать ответ на неё, а не продолжать генерировать новые примеры. \nОтвечай только на последний вопрос/задачу. \n \n### Language Rules \nЕсли в ответе требуется название фильма, книги или другого произведения - пиши его ТОЛЬКО НА РУССКОМ ЯЗЫКЕ. \nНапример, не "Taxi Driver", а "Таксист". Не "The Godfather", а "Крёстный отец". \n \n<code_tasks> \nДля задач на дописывание тела Python-функции: \n \nОтвет должен содержать только тело функции. \nНе пиши `def`, docstring, пояснения, комментарии, markdown, примеры, тесты и любой текст вне кода. \nНе используй `input()`. \n \nФормат ответа строгий: \n1. Первая непустая строка ответа должна начинаться ровно с 4 пробелов. \n2. Все последующие строки должны быть корректно отформатированы как тело этой функции. \n3. Не печатай символы ` \n`, ` ` и другие escape-последовательности текстом. \n4. Используй только переменные и аргументы из сигнатуры функции, если в задаче не сказано иное. \n5. Не создавай вложенные функции, если этого прямо не требует задача. \n \nПример: \n \nЗадание: \ndef sum_list(numbers: List[int]) -> int: \n (docstring) \n \nПравильный ответ: \n total = 0 \n for n in numbers: \n total += n \n return total \n \nНеправильный ответ: \ndef sum_list(numbers: List[int]) -> int: \n total = 0 \n return total \n \nНеправильный ответ: \nВот решение: \n total = 0 \n return total \n \nНеправильный ответ: \n \n total = 0 \n</code_tasks>

Description of the template:
[gMASK]<sop> {%- if tools -%} <|system|> # Tools You may call one or more functions to assist with the user query. You are provided with function signatures within <tools></tools> XML tags: <tools> {% for tool in tools %} {{ tool | tojson(ensure_ascii=False) }} {% endfor %} </tools> For each function call, output the function name and arguments within the following XML format: <tool_call>{function-name}<arg_key>{arg-key-1}</arg_key><arg_value>{arg-value-1}</arg_value><arg_key>{arg-key-2}</arg_key><arg_value>{arg-value-2}</arg_value>...</tool_call>{%- endif -%} {%- macro visible_text(content) -%} {%- if content is string -%} {{- content }} {%- elif content is iterable and content is not mapping -%} {%- for item in content -%} {%- if item is mapping and item.type == 'text' -%} {{- item.text }} {%- elif item is string -%} {{- item }} {%- endif -%} {%- endfor -%} {%- else -%} {{- content }} {%- endif -%} {%- endmacro -%} {%- set ns = namespace(last_user_index=-1) %} {%- for m in messages %} {%- if m.role == 'user' %} {% set ns.last_user_index = loop.index0 -%} {%- endif %} {%- endfor %} {% for m in messages %} {%- if m.role == 'user' -%}<|user|>{{ visible_text(m.content) }} {%- elif m.role == 'assistant' -%} <|assistant|> {%- set reasoning_content = '' %} {%- set content = visible_text(m.content) %} {%- if m.reasoning_content is string %} {%- set reasoning_content = m.reasoning_content %} {%- else %} {%- if '</think>' in content %} {%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %} {%- set content = content.split('</think>')[-1].lstrip('\n') %} {%- endif %} {%- endif %} {%- if ((clear_thinking is defined and not clear_thinking) or loop.index0 > ns.last_user_index) and reasoning_content -%} {{ '<think>' + reasoning_content.strip() + '</think>'}} {%- else -%} {{ '</think>' }} {%- endif -%} {%- if content.strip() -%} {{ content.strip() }} {%- endif -%} {% if m.tool_calls %} {% for tc in m.tool_calls %} {%- if tc.function %} {%- set tc = tc.function %} {%- endif %} {{- '<tool_call>' + tc.name -}} {% set _args = tc.arguments %}{% for k, v in _args.items() %}<arg_key>{{ k }}</arg_key><arg_value>{{ v | tojson(ensure_ascii=False) if v is not string else v }}</arg_value>{% endfor %}</tool_call>{% endfor %} {% endif %} {%- elif m.role == 'tool' -%} {%- if m.content is string -%} {%- if loop.first or (messages[loop.index0 - 1].role != "tool") %} {{- '<|observation|>' }} {%- endif %} {{- '<tool_response>' }} {{- m.content }} {{- '</tool_response>' }} {%- else -%} <|observation|>{% for tr in m.content %} <tool_response>{{ tr.output if tr.output is defined else tr }}</tool_response>{% endfor -%} {% endif -%} {%- elif m.role == 'system' -%} <|system|>{{ visible_text(m.content) }} {%- endif -%} {%- endfor -%} {%- if add_generation_prompt -%} <|assistant|>{{- '</think>' if (enable_thinking is defined and not enable_thinking) else '<think>' -}} {%- endif -%}

Ratings by subcategory

Metric: Grade Norm
Model, team 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 8_0 8_1 8_2 8_3 8_4
BerryLM-XL
Wildberries & Russ
0.433 0.767 0.833 0.4 0.733 0.8 0.567 - 0.233 0.2 0.367 0.333 0.567 0.4 0.333 0.55 0.6 0.367 0.8 0.433 0.8 0.8 0.6 0.767 0.733 0.858 0.767 0.833 0.767 0.833 0.9
Model, team Honest Helpful Harmless
BerryLM-XL
Wildberries & Russ
0.852 0.898 0.983
Model, team Anatomy Virology Astronomy Marketing Nutrition Sociology Management Philosophy Prehistory Human aging Econometrics Formal logic Global facts Jurisprudence Miscellaneous Moral disputes Business ethics Biology (college) Physics (college) Human Sexuality Moral scenarios World religions Abstract algebra Medicine (college) Machine learning Medical genetics Professional law PR Security studies Chemistry (школьная) Computer security International law Logical fallacies Politics Clinical knowledge Conceptual_physics Math (college) Biology (high school) Physics (high school) Chemistry (high school) Geography (high school) Professional medicine Electrical engineering Elementary mathematics Psychology (high school) Statistics (high school) History (high school) Math (high school) Professional accounting Professional psychology Computer science (college) World history (high school) Macroeconomics Microeconomics Computer science (high school) European history Government and politics
BerryLM-XL
Wildberries & Russ
0.889 0.584 0.967 0.923 0.931 0.935 0.903 0.875 0.923 0.798 0.842 0.944 0.72 0.889 0.959 0.85 0.84 0.993 0.978 0.893 0.798 0.924 0.96 0.873 0.893 1 0.734 0.75 0.824 0.74 0.9 0.917 0.865 0.939 0.913 0.957 0.98 0.948 0.954 0.951 0.914 0.956 0.862 0.976 0.957 0.931 0.956 0.989 0.915 0.89 0.96 0.916 0.954 0.983 0.99 0.879 0.943
Model, team SIM FL STA
BerryLM-XL
Wildberries & Russ
0.802 0.611 0.731
Model, team Anatomy Virology Astronomy Marketing Nutrition Sociology Managment Philosophy Pre-History Gerontology Econometrics Formal logic Global facts Jurisprudence Miscellaneous Moral disputes Business ethics Bilology (college) Physics (college) Human sexuality Moral scenarios World religions Abstract algebra Medicine (college) Machine Learning Genetics Professional law PR Security Chemistry (college) Computer security International law Logical fallacies Politics Clinical knowledge Conceptual physics Math (college) Biology (high school) Physics (high school) Chemistry (high school) Geography (high school) Professional medicine Electrical Engineering Elementary mathematics Psychology (high school) Statistics (high school) History (high school) Math (high school) Professional Accounting Professional psychology Computer science (college) World history (high school) Macroeconomics Microeconomics Computer science (high school) Europe History Government and politics
BerryLM-XL
Wildberries & Russ
0.889 0.931 0.867 0.796 0.947 0.862 0.862 0.825 0.942 0.815 0.808 0.875 0.717 0.938 0.936 0.852 0.804 0.844 0.93 0.877 0.877 0.949 0.933 0.959 0.956 0.939 0.872 0.842 0.947 0.956 0.889 0.949 0.964 0.982 0.924 0.929 0.933 0.911 0.912 0.954 0.967 0.952 0.889 1 0.931 0.911 0.948 0.909 0.954 0.965 0.956 0.986 0.886 0.831 0.767 0.953 0.956
Coorect
Good
Ethical
Model, team Virtue Law Moral Justice Utilitarianism
BerryLM-XL
Wildberries & Russ
0.426 0.472 0.475 0.4 0.364
Model, team Virtue Law Moral Justice Utilitarianism
BerryLM-XL
Wildberries & Russ
0.426 0.423 0.456 0.396 0.383
Model, team Virtue Law Moral Justice Utilitarianism
BerryLM-XL
Wildberries & Russ
0.714 0.694 0.762 0.636 0.597
Model, team Women Men LGBT Nationalities Migrants Other
BerryLM-XL
Wildberries & Russ
0.981 0.8 0.882 0.973 0.857 0.951