Catalog
Korean AI datasets, by name
201–240 of 286 · page 6 / 8
This is the full GearDel catalog split into short HTML pages so crawlers do not have to download every row on the homepage. License and size cells are the stored strings. Themed lists (speech, OCR, instruction, sentiment) live under Collections. GearDel does not host files.
Pick by task →·Search on the homepage →
| Name | Use / topic | Source | Size | License | Commercial |
|---|---|---|---|---|---|
| KBL | Korean QA Dataset | Hugging Face | 8.1MB | CC BY-NC 4.0 | Commercial use may be conditional |
| KBMC NER | Korean NER Dataset | Hugging Face | 1.4MB | Apache 2.0 | Commercial use allowed (catalog) |
| KBS MBC | Korean Summarization Dataset | AI Hub | 440MB | AI Hub terms | Commercial use may be conditional |
| KcBERT | Korean Pretraining Corpus | Hugging Face | 7.90 GiB | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| KIT-19 | Korean Instruction Tuning Dataset | Hugging Face | 85MB | CC BY 4.0 | Commercial use allowed (catalog) |
| KLUE | Korean NLP Dataset | Hugging Face (KLUE) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| KLUE | Korean NER Dataset | Hugging Face (KLUE) | Medium | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| KLUE | Korean NLP Dataset | Hugging Face (KLUE) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| KLUE | Korean QA Dataset | Hugging Face (KLUE) | Medium | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| KLUE | Korean LLM Benchmark Dataset | Hugging Face (KLUE) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| KLUE | Korean NLP Dataset | Hugging Face (KLUE) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| KLUE Benchmark | Korean LLM Benchmark Dataset | Hugging Face | 62.3 MB | CC BY-SA 4.0 | Commercial use may be conditional |
| KMMLU | Korean LLM Benchmark Dataset | Hugging Face | 70.7 MB | CC BY-ND 4.0 | Commercial use restricted |
| KMMLU-Pro | Korean LLM Benchmark Dataset | Hugging Face (LGAI-EXAONE) | 12MB | CC BY-NC 4.0 | Commercial use allowed (catalog) |
| KMRE | Korean NLP Dataset | GitHub | 15MB | Apache 2.0 | Commercial use allowed (catalog) |
| Ko-Agent-Trajectories-1.0 | Korean NLP Dataset | Hugging Face | Large | Apache 2.0 | Commercial use allowed (catalog) |
| Ko-ARC Challenge | Korean LLM Benchmark Dataset | Hugging Face | 8MB | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| Ko-Llama | Korean Instruction Tuning Dataset | Hugging Face | Large | Apache 2.0 | Commercial use allowed (catalog) |
| Ko-MMLU-Pro | Korean LLM Benchmark Dataset | Hugging Face | 45MB | MIT | Commercial use allowed (catalog) |
| Ko-PIQA | Korean QA Dataset | Hugging Face | 204KB | License unknown | Commercial use may be conditional |
| KoAlpaca v1.0 Alpaca | Korean Instruction Tuning Dataset | Hugging Face | 8.1MB | CC BY-NC 4.0 | Commercial use may be conditional |
| KoAlpaca-v1.1a | Korean Instruction Tuning Dataset | Hugging Face | Large | License unknown | Commercial use allowed (catalog) |
| KoBALT-700 | Korean LLM Benchmark Dataset | Hugging Face | 644KB | CC BY-NC 4.0 | Commercial use may be conditional |
| KoBBQ QA | Korean QA Dataset | Hugging Face | 2.4MB | MIT | Commercial use allowed (catalog) |
| BoolQ | Korean QA Dataset | Hugging Face (SKT-brain) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| WiC | Korean NLP Dataset | Hugging Face (SKT-brain) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| HellaSwag | Korean LLM Benchmark Dataset | Hugging Face (SKT-brain) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| KoBEST | Korean LLM Benchmark Dataset | Hugging Face | 5.1MB | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| COPA | Korean LLM Benchmark Dataset | Hugging Face (SKT-brain) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| KoBigBench Lite | Korean LLM Benchmark Dataset | Hugging Face | 12MB | License unknown | Commercial use allowed (catalog) |
| KoCommonGEN v2 | Korean LLM Benchmark Dataset | Hugging Face | 183KB | License unknown | Commercial use may be conditional |
| KoCoSa | Korean Instruction Tuning Dataset | GitHub | 18MB | Apache 2.0 | Commercial use allowed (catalog) |
| KoDialogBench | Korean LLM Benchmark Dataset | Hugging Face | 82,962 , 21 | CC BY-NC-SA 4.0 | Commercial use may be conditional |
| KoELECTRA | Korean Dialogue Dataset | GitHub (monologg) | Large (34GB) | Apache 2.0 | Commercial use allowed (catalog) |
| KoelLabs | Korean Speech Recognition Dataset | Hugging Face (KoelLabs) | 18 GB | CC BY-NC 4.0 | Commercial use allowed (catalog) |
| KoGEM | Korean LLM Benchmark Dataset | Hugging Face | 1.2MB | MIT | Commercial use allowed (catalog) |
| KoGPT2 | Korean Pretraining Corpus | Hugging Face | 3.2GB | License unknown | Commercial use allowed (catalog) |
| KoIn | Korean Computer Vision Dataset | GitHub (dukong1) | 140 GB | MIT | Commercial use allowed (catalog) |
| KoInFoBench | Korean Instruction Tuning Dataset | Hugging Face | Small | MIT | Commercial use allowed (catalog) |
| KOLD | Korean Toxicity Dataset | GitHub | 40,429 comments | License unknown | Commercial use may be conditional |