Catalog
Korean AI datasets, by name
161–200 of 286 · page 5 / 8
This is the full GearDel catalog split into short HTML pages so crawlers do not have to download every row on the homepage. License and size cells are the stored strings. Themed lists (speech, OCR, instruction, sentiment) live under Collections. GearDel does not host files.
Pick by task →·Search on the homepage →
| Name | Use / topic | Source | Size | License | Commercial |
|---|---|---|---|---|---|
| Korean Summarization Dataset | Korean Summarization Dataset | Hugging Face / AI Hub | Large | Apache 2.0 | Commercial use allowed (catalog) |
| QG | Korean Instruction Tuning Dataset | Hugging Face | Small | Apache 2.0 | Commercial use allowed (catalog) |
| KorQuAD | Korean QA Dataset | KorQuAD | Medium | CC BY-ND 4.0 | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | Hugging Face | Small | Apache 2.0 | Commercial use allowed (catalog) |
| Function Calling | Korean Instruction Tuning Dataset | Hugging Face (gyung) | Small | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 140hours | AI Hub terms | Commercial use may be conditional |
| Korean OCR Dataset | Korean OCR Dataset | AI Hub | 140GB | AI Hub terms | Commercial use may be conditional |
| Korean Pretraining Corpus | Korean Pretraining Corpus | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean LLM Benchmark Dataset | Korean LLM Benchmark Dataset | Hugging Face / GitHub | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 2.4TB | AI Hub terms | Commercial use may be conditional |
| Korean Translation Dataset | Korean Translation Dataset | AI Hub | 55MB | AI Hub terms | Commercial use may be conditional |
| Korean Translation Dataset | Korean Translation Dataset | GitHub | Medium | CC0 | Commercial use allowed (catalog) |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 420GB | AI Hub terms | Commercial use may be conditional |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 210GB | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 15GB | AI Hub terms | Commercial use may be conditional |
| APEACH | Korean LLM Benchmark Dataset | Hugging Face | 1.1MB | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| BEEP! Korean Hate Speech | Korean Toxicity Dataset | GitHub | 1.2MB | CC BY-SA 4.0 | Commercial use may be conditional |
| CCTV | Korean Computer Vision Dataset | AI Hub | 5TB | AI Hub terms | Commercial use may be conditional |
| CLIcK (Cultural & Legal Knowledge) | Korean LLM Benchmark Dataset | GitHub | 5.5MB | License unknown | Commercial use allowed (catalog) |
| ClovaCall | Korean Speech Recognition Dataset | GitHub | 11,000 | MIT | Commercial use restricted |
| COYO-700M | Korean Computer Vision Dataset | Hugging Face | Large (7 ) | CC BY 4.0 | Commercial use allowed (catalog) |
| CSAT-KOREAN-2025 | Korean QA Dataset | Hugging Face (KKACHI-HUB) | 3.2MB | MIT | Commercial use allowed (catalog) |
| DKTC (Dataset of Korean Threat Dialogue) | Korean Dialogue Dataset | GitHub | 1.8MB | CC BY-NC-SA 4.0 | Commercial use restricted |
| FunctionChat-Bench | Korean LLM Benchmark Dataset | GitHub | 500 + 45 | Apache 2.0 | Commercial use allowed (catalog) |
| GovOn Legal Response Dataset | Korean QA Dataset | Hugging Face | 267 MB | CC BY 4.0 | Commercial use allowed (catalog) |
| GSM8K | Korean Translation Dataset | Hugging Face | 4.8MB | MIT | Commercial use allowed (catalog) |
| HAE-RAE Bench 1.1 | Korean LLM Benchmark Dataset | Hugging Face | 2.9MB | CC BY-NC-ND 4.0 | Commercial use may be conditional |
| HAE-RAE Bench 2.0 | Korean LLM Benchmark Dataset | Hugging Face | 1.0MB | MIT | Commercial use allowed (catalog) |
| HAE-RAE CoT 1.5M | Korean NLP Dataset | Hugging Face | 1.05GB | CC BY 4.0 | Commercial use allowed (catalog) |
| HateScore | Korean Toxicity Dataset | GitHub | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| HRM8K (HAE-RAE Math 8K) | Korean LLM Benchmark Dataset | Hugging Face | 4.9MB | MIT | Commercial use allowed (catalog) |
| HRMCR | Korean LLM Benchmark Dataset | Hugging Face | 106KB | Apache 2.0 | Commercial use allowed (catalog) |
| Jejueo JIT JSS | Korean Translation Dataset | GitHub | 170K+ sentences, 10K | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 850GB | AI Hub terms | Commercial use may be conditional |
| K-BrowseComp | Korean LLM Benchmark Dataset | Hugging Face (prometheus-eval) | 3.5MB | MIT | Commercial use allowed (catalog) |
| K-HATERS | Korean Toxicity Dataset | Hugging Face | 50.5MB | CC BY 4.0 | Commercial use allowed (catalog) |
| K-MHaS | Korean Toxicity Dataset | Hugging Face | 9.1MB | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| K-MMBench | Korean Computer Vision Dataset | Hugging Face | 92.4MB | CC BY-NC 4.0 | Commercial use may be conditional |
| K-POP | Korean Translation Dataset | GitHub | 2.5MB | License unknown | Commercial use allowed (catalog) |
| Eval | Korean Instruction Tuning Dataset | Hugging Face | 178KB | MIT | Commercial use allowed (catalog) |