Collection
Korean LLM benchmark datasets
47 catalog items · GearDel
GearDel lists 47 catalog items in this group (46 text, 1 image). Sources: 27 Hugging Face, 3 GitHub. Examples: KLUE Benchmark, KMMLU, CLIcK (Cultural & Legal Knowledge). 35 are marked commercial-use-allowed in the catalog. These are evaluation sets, not pretraining dumps. Use them to score models, then train on a separate corpus. Use this list to compare license and size, then open the original provider. GearDel does not host files.
Do not train on the test split. KLUE, KoBEST, and KMMLU families are for scoring, then go back to an instruction or corpus collection.
| Name | Use / topic | Source | Size | License | Commercial |
|---|---|---|---|---|---|
| KLUE Benchmark | Korean LLM Benchmark Dataset | Hugging Face | 62.3 MB | CC BY-SA 4.0 | Commercial use may be conditional |
| KMMLU | Korean LLM Benchmark Dataset | Hugging Face | 70.7 MB | CC BY-ND 4.0 | Commercial use restricted |
| CLIcK (Cultural & Legal Knowledge) | Korean LLM Benchmark Dataset | GitHub | 5.5MB | License unknown | Commercial use allowed (catalog) |
| KoBEST | Korean LLM Benchmark Dataset | Hugging Face | 5.1MB | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| KoBigBench Lite | Korean LLM Benchmark Dataset | Hugging Face | 12MB | License unknown | Commercial use allowed (catalog) |
| KoInFoBench | Korean Instruction Tuning Dataset | Hugging Face | Small | MIT | Commercial use allowed (catalog) |
| Korean QA Dataset | Korean QA Dataset | Hugging Face | Small | Apache 2.0 | Commercial use allowed (catalog) |
| Korean LLM Benchmark Dataset | Korean LLM Benchmark Dataset | Hugging Face (SOGANG-ISDS) | Small | Apache 2.0 | Commercial use allowed (catalog) |
| SNU Ko-MuSR | Korean LLM Benchmark Dataset | Hugging Face (thunder-research-group) | Small | MIT | Commercial use allowed (catalog) |
| KLUE | Korean NLP Dataset | Hugging Face (KLUE) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| KLUE | Korean NLP Dataset | Hugging Face (KLUE) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| KLUE | Korean LLM Benchmark Dataset | Hugging Face (KLUE) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| KLUE | Korean NER Dataset | Hugging Face (KLUE) | Medium | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| KLUE | Korean NLP Dataset | Hugging Face (KLUE) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| KLUE | Korean QA Dataset | Hugging Face (KLUE) | Medium | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| COPA | Korean LLM Benchmark Dataset | Hugging Face (SKT-brain) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| HellaSwag | Korean LLM Benchmark Dataset | Hugging Face (SKT-brain) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| BoolQ | Korean QA Dataset | Hugging Face (SKT-brain) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| WiC | Korean NLP Dataset | Hugging Face (SKT-brain) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| Korean LLM Benchmark Dataset | Korean LLM Benchmark Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| KMMLU-Pro | Korean LLM Benchmark Dataset | Hugging Face (LGAI-EXAONE) | 12MB | CC BY-NC 4.0 | Commercial use allowed (catalog) |
| Thunder-KoNUBench | Korean LLM Benchmark Dataset | Hugging Face (thunder-research-group) | 5.4MB | CC BY-NC-SA 4.0 | Commercial use allowed (catalog) |
| K-BrowseComp | Korean LLM Benchmark Dataset | Hugging Face (prometheus-eval) | 3.5MB | MIT | Commercial use allowed (catalog) |
| CSAT-KOREAN-2025 | Korean QA Dataset | Hugging Face (KKACHI-HUB) | 3.2MB | MIT | Commercial use allowed (catalog) |
| Korean LLM Benchmark Dataset | Korean LLM Benchmark Dataset | Hugging Face / GitHub | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| Korean LLM Benchmark Dataset | Korean LLM Benchmark Dataset | Hugging Face | Small | Apache 2.0 | Commercial use allowed (catalog) |
| KoGEM | Korean LLM Benchmark Dataset | Hugging Face | 1.2MB | MIT | Commercial use allowed (catalog) |
| Ko-MMLU-Pro | Korean LLM Benchmark Dataset | Hugging Face | 45MB | MIT | Commercial use allowed (catalog) |
| KoSafetyBench | Korean LLM Benchmark Dataset | Hugging Face | 15MB | CC BY 4.0 | Commercial use allowed (catalog) |
| Ko-ARC Challenge | Korean LLM Benchmark Dataset | Hugging Face | 8MB | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| HAE-RAE Bench 1.1 | Korean LLM Benchmark Dataset | Hugging Face | 2.9MB | CC BY-NC-ND 4.0 | Commercial use may be conditional |
| HAE-RAE Bench 2.0 | Korean LLM Benchmark Dataset | Hugging Face | 1.0MB | MIT | Commercial use allowed (catalog) |
| KoCommonGEN v2 | Korean LLM Benchmark Dataset | Hugging Face | 183KB | License unknown | Commercial use may be conditional |
| APEACH | Korean LLM Benchmark Dataset | Hugging Face | 1.1MB | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| HRM8K (HAE-RAE Math 8K) | Korean LLM Benchmark Dataset | Hugging Face | 4.9MB | MIT | Commercial use allowed (catalog) |
| KorNAT QA | Korean QA Dataset | Hugging Face | 2.6MB | CC BY-NC 4.0 | Commercial use may be conditional |
| HRMCR | Korean LLM Benchmark Dataset | Hugging Face | 106KB | Apache 2.0 | Commercial use allowed (catalog) |
| KoBBQ QA | Korean QA Dataset | Hugging Face | 2.4MB | MIT | Commercial use allowed (catalog) |
| KorMedMCQA QA | Korean QA Dataset | Hugging Face | 2.1MB | CC BY-NC 4.0 | Commercial use may be conditional |
| KoBALT-700 | Korean LLM Benchmark Dataset | Hugging Face | 644KB | CC BY-NC 4.0 | Commercial use may be conditional |
| KoSimpleQA QA | Korean QA Dataset | Hugging Face | 301KB | License unknown | Commercial use may be conditional |
| Ko-PIQA | Korean QA Dataset | Hugging Face | 204KB | License unknown | Commercial use may be conditional |
| K-MMBench | Korean Computer Vision Dataset | Hugging Face | 92.4MB | CC BY-NC 4.0 | Commercial use may be conditional |
| Eval | Korean Instruction Tuning Dataset | Hugging Face | 178KB | MIT | Commercial use allowed (catalog) |
| KoDialogBench | Korean LLM Benchmark Dataset | Hugging Face | 82,962 , 21 | CC BY-NC-SA 4.0 | Commercial use may be conditional |
| FunctionChat-Bench | Korean LLM Benchmark Dataset | GitHub | 500 + 45 | Apache 2.0 | Commercial use allowed (catalog) |
| KBL | Korean QA Dataset | Hugging Face | 8.1MB | CC BY-NC 4.0 | Commercial use may be conditional |