ModelChorusModelChorus
ChallengeChatLeaderboardBenchmarksHistoryHow it works
Terms of ServicePrivacy PolicyAPI

Copyright 2026 MeetKai Inc.

Benchmarks/GLM 5/English (US) tasks

GLM 5

5 tasks

Each row below is a single benchmark task this model was evaluated on. The Score column averages every metric the task reports (accuracy, F1, exact-match, etc.). Click a row to browse the individual questions and the model's responses.

Average
53.2
ScoreLanguageTaskMetrics
85.7English (US)
english_gsm8k
english math
exact_match: 85.7sample_len: 1319.0
exact_match: 85.7sample_len: 1319.0
84.0English (US)
english_mgsm
english math
exact_match: 84.0sample_len: 250.0
exact_match: 84.0sample_len: 250.0
50.6English (US)
english_mmlu_pro
english mmlu pro
exact_match: 50.6sample_len: 2100.0
exact_match: 50.6sample_len: 2100.0
36.3English (US)
ifeval
ifeval
inst_level_loose_acc: 45.0inst_level_strict_acc: 43.0prompt_level_loose_acc: 29.8prompt_level_strict_acc: 27.5sample_len: 541.0
inst_level_loose_acc: 45.0inst_level_strict_acc: 43.0prompt_level_loose_acc: 29.8prompt_level_strict_acc: 27.5sample_len: 541.0
9.3English (US)
english_belebele
english mcq
f1_macro: 9.3sample_len: 900.0
f1_macro: 9.3sample_len: 900.0