Collection
Korean AI dataset catalog
286 catalog items · GearDel
GearDel lists 286 catalog items in this group (215 text, 38 speech, 33 image). Sources: 79 Hugging Face, 92 AI Hub, 60 GitHub. Examples: Korean Summarization Dataset, KULLM-v2, GovOn Legal Response Dataset. 161 are marked commercial-use-allowed in the catalog. Use this list to compare license and size, then open the original provider. GearDel does not host files.
Compare license and source before download. 161 of 286 carry a commercial-allowed catalog mark.
| Name | Use / topic | Source | Size | License | Commercial |
|---|---|---|---|---|---|
| Korean Summarization Dataset | Korean Summarization Dataset | AI Hub | Large | AI Hub terms | Commercial use may be conditional |
| KULLM-v2 | Korean Instruction Tuning Dataset | Hugging Face | 180MB | Apache 2.0 | Commercial use allowed (catalog) |
| GovOn Legal Response Dataset | Korean QA Dataset | Hugging Face | 267 MB | CC BY 4.0 | Commercial use allowed (catalog) |
| KLUE Benchmark | Korean LLM Benchmark Dataset | Hugging Face | 62.3 MB | CC BY-SA 4.0 | Commercial use may be conditional |
| KMMLU | Korean LLM Benchmark Dataset | Hugging Face | 70.7 MB | CC BY-ND 4.0 | Commercial use restricted |
| KoAlpaca-v1.1a | Korean Instruction Tuning Dataset | Hugging Face | Large | License unknown | Commercial use allowed (catalog) |
| CLIcK (Cultural & Legal Knowledge) | Korean LLM Benchmark Dataset | GitHub | 5.5MB | License unknown | Commercial use allowed (catalog) |
| BEEP! Korean Hate Speech | Korean Toxicity Dataset | GitHub | 1.2MB | CC BY-SA 4.0 | Commercial use may be conditional |
| DKTC (Dataset of Korean Threat Dialogue) | Korean Dialogue Dataset | GitHub | 1.8MB | CC BY-NC-SA 4.0 | Commercial use restricted |
| Korean Dialogue Dataset | Korean Dialogue Dataset | NIKL | Large | NIKL terms | Commercial use may be conditional |
| Korean Dialogue Dataset | Korean Dialogue Dataset | NIKL | 3.5MB | NIKL terms | Commercial use may be conditional |
| Korean Dialogue Dataset | Korean Dialogue Dataset | NIKL | 24.5MB | NIKL terms | Commercial use may be conditional |
| Korean Dialogue Dataset | Korean Dialogue Dataset | NIKL | 1.2GB | NIKL terms | Commercial use may be conditional |
| Korean Dialogue Dataset | Korean Dialogue Dataset | NIKL | 520MB | NIKL terms | Commercial use may be conditional |
| Korean Dialogue Dataset | Korean Dialogue Dataset | NIKL | 850MB | NIKL terms | Commercial use may be conditional |
| Korean Dialogue Dataset | Korean Dialogue Dataset | NIKL | 120MB | NIKL terms | Commercial use may be conditional |
| KoBEST | Korean LLM Benchmark Dataset | Hugging Face | 5.1MB | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| GSM8K | Korean Translation Dataset | Hugging Face | 4.8MB | MIT | Commercial use allowed (catalog) |
| KoBigBench Lite | Korean LLM Benchmark Dataset | Hugging Face | 12MB | License unknown | Commercial use allowed (catalog) |
| KoGPT2 | Korean Pretraining Corpus | Hugging Face | 3.2GB | License unknown | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | AI Hub | 450MB | AI Hub terms | Commercial use may be conditional |
| Korean Summarization Dataset | Korean Summarization Dataset | AI Hub | 620MB | AI Hub terms | Commercial use may be conditional |
| Korean QA Dataset | Korean QA Dataset | AI Hub | 180MB | AI Hub terms | Commercial use may be conditional |
| Korean QA Dataset | Korean QA Dataset | AI Hub | 140MB | AI Hub terms | Commercial use may be conditional |
| Korean Summarization Dataset | Korean Summarization Dataset | AI Hub | 540MB | AI Hub terms | Commercial use may be conditional |
| Korean NLP Dataset | Korean NLP Dataset | AI Hub | 890MB | AI Hub terms | Commercial use may be conditional |
| Korean NLP Dataset | Korean NLP Dataset | AI Hub | 320MB | AI Hub terms | Commercial use may be conditional |
| Korean NLP Dataset | Korean NLP Dataset | AI Hub | 1.1GB | AI Hub terms | Commercial use may be conditional |
| SNS | Korean Dialogue Dataset | AI Hub | 980MB | AI Hub terms | Commercial use may be conditional |
| Korean Summarization Dataset | Korean Summarization Dataset | AI Hub | 750MB | AI Hub terms | Commercial use may be conditional |
| Korean QA Dataset | Korean QA Dataset | AI Hub | 45MB | AI Hub terms | Commercial use may be conditional |
| Korean Summarization Dataset | Korean Summarization Dataset | AI Hub | 410MB | AI Hub terms | Commercial use may be conditional |
| Korean Dialogue Dataset | Korean Dialogue Dataset | AI Hub | 2.4GB | AI Hub terms | Commercial use may be conditional |
| Korean Translation Dataset | Korean Translation Dataset | AI Hub | 55MB | AI Hub terms | Commercial use may be conditional |
| Korean Dialogue Dataset | Korean Dialogue Dataset | AI Hub | 3.5GB | AI Hub terms | Commercial use may be conditional |
| Korean Summarization Dataset | Korean Summarization Dataset | AI Hub | 780MB | AI Hub terms | Commercial use may be conditional |
| Korean QA Dataset | Korean QA Dataset | AI Hub | 8.5MB | AI Hub terms | Commercial use may be conditional |
| Korean QA Dataset | Korean QA Dataset | AI Hub | 120MB | AI Hub terms | Commercial use may be conditional |
| Korean NLP Dataset | Korean NLP Dataset | AI Hub | 180MB | AI Hub terms | Commercial use may be conditional |
| Korean NLP Dataset | Korean NLP Dataset | AI Hub | 340MB | AI Hub terms | Commercial use may be conditional |
| Korean NLP Dataset | Korean NLP Dataset | AI Hub | 190MB | AI Hub terms | Commercial use may be conditional |
| Korean NLP Dataset | Korean NLP Dataset | AI Hub | 480MB | AI Hub terms | Commercial use may be conditional |
| Korean NLP Dataset | Korean NLP Dataset | AI Hub | 520MB | AI Hub terms | Commercial use may be conditional |
| Korean NLP Dataset | Korean NLP Dataset | AI Hub | 90MB | AI Hub terms | Commercial use may be conditional |
| Korean NLP Dataset | Korean NLP Dataset | AI Hub | 1.5GB | AI Hub terms | Commercial use may be conditional |
| Korean NLP Dataset | Korean NLP Dataset | AI Hub | 110MB | AI Hub terms | Commercial use may be conditional |
| Korean Summarization Dataset | Korean Summarization Dataset | AI Hub | 85MB | AI Hub terms | Commercial use may be conditional |
| SUV QA | Korean QA Dataset | AI Hub | 230MB | AI Hub terms | Commercial use may be conditional |
| Korean Sentiment Analysis Dataset | Korean Sentiment Analysis Dataset | AI Hub | 40MB | AI Hub terms | Commercial use may be conditional |
| Korean NLP Dataset | Korean NLP Dataset | AI Hub | 130MB | AI Hub terms | Commercial use may be conditional |
| Korean Translation Dataset | Korean Translation Dataset | AI Hub | 1.8GB | AI Hub terms | Commercial use may be conditional |
| KBS MBC | Korean Summarization Dataset | AI Hub | 440MB | AI Hub terms | Commercial use may be conditional |
| Korean Dialogue Dataset | Korean Dialogue Dataset | AI Hub | 310MB | AI Hub terms | Commercial use may be conditional |
| Korean QA Dataset | Korean QA Dataset | AI Hub | 85MB | AI Hub terms | Commercial use may be conditional |
| KSS Dataset | Korean Speech Synthesis Dataset | Hugging Face | 12hours+ | CC BY-NC-SA 4.0 | Commercial use restricted |
| ClovaCall | Korean Speech Recognition Dataset | GitHub | 11,000 | MIT | Commercial use restricted |
| KsponSpeech | Korean Speech Recognition Dataset | AI Hub | 1000hours | AI Hub terms | Commercial use may be conditional |
| Speech Emotion Dataset (ko) | Korean Speech Dataset | GitHub | 3.2GB | License unknown | Commercial use allowed (catalog) |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 500hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Recognition Dataset | Korean Speech Recognition Dataset | AI Hub | 120hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Recognition Dataset | Korean Speech Recognition Dataset | AI Hub | 150hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 200hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Synthesis Dataset | Korean Speech Synthesis Dataset | AI Hub | 100hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 80hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 800hours | AI Hub terms | Commercial use may be conditional |
| TTS | Korean Speech Synthesis Dataset | AI Hub | 300hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 600hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 240GB | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 400hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Synthesis Dataset | Korean Speech Synthesis Dataset | AI Hub | 180hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 4TB | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 50hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 40GB | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 15GB | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 30GB | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 85GB | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 120GB | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 60hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 150GB | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 8GB | AI Hub terms | Commercial use may be conditional |
| TTS | Korean Speech Synthesis Dataset | AI Hub | 90hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 220hours | AI Hub terms | Commercial use may be conditional |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub | 140hours | AI Hub terms | Commercial use may be conditional |
| COYO-700M | Korean Computer Vision Dataset | Hugging Face | Large (7 ) | CC BY 4.0 | Commercial use allowed (catalog) |
| skt KVQA | Korean Computer Vision Dataset | Hugging Face | 100,445 | Korean VQA License | Commercial use restricted |
| KRETA (Korean Rich Text Visual QA) | Korean OCR Dataset | Hugging Face | 125MB | License unknown | Commercial use allowed (catalog) |
| Korean Tourist Spot Dataset | Korean Computer Vision Dataset | GitHub | 8.4GB | License unknown | Commercial use allowed (catalog) |
| Korean-Light-OCR | Korean OCR Dataset | GitHub | 4.2GB | License unknown | Commercial use allowed (catalog) |
| Korean Receipts OCR | Korean OCR Dataset | Hugging Face | 95MB | CC BY 4.0 | Commercial use allowed (catalog) |
| Table-VQA-ko | Korean OCR Dataset | Hugging Face | 185MB | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 12TB | AI Hub terms | Commercial use may be conditional |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 850GB | AI Hub terms | Commercial use may be conditional |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 2.4TB | AI Hub terms | Commercial use may be conditional |
| OCR | Korean OCR Dataset | AI Hub | 1.5TB | AI Hub terms | Commercial use may be conditional |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 420GB | AI Hub terms | Commercial use may be conditional |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 380GB | AI Hub terms | Commercial use may be conditional |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 190GB | AI Hub terms | Commercial use may be conditional |
| CCTV | Korean Computer Vision Dataset | AI Hub | 5TB | AI Hub terms | Commercial use may be conditional |
| OCR | Korean OCR Dataset | AI Hub | 290GB | AI Hub terms | Commercial use may be conditional |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 8TB | AI Hub terms | Commercial use may be conditional |
| Korean OCR Dataset | Korean OCR Dataset | AI Hub | 140GB | AI Hub terms | Commercial use may be conditional |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 640GB | AI Hub terms | Commercial use may be conditional |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 910GB | AI Hub terms | Commercial use may be conditional |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 430GB | AI Hub terms | Commercial use may be conditional |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 150GB | AI Hub terms | Commercial use may be conditional |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 210GB | AI Hub terms | Commercial use may be conditional |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 1.2TB | AI Hub terms | Commercial use may be conditional |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 750GB | AI Hub terms | Commercial use may be conditional |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 1.1TB | AI Hub terms | Commercial use may be conditional |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 1.6TB | AI Hub terms | Commercial use may be conditional |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 320GB | AI Hub terms | Commercial use may be conditional |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 480GB | AI Hub terms | Commercial use may be conditional |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | AI Hub | 520GB | AI Hub terms | Commercial use may be conditional |
| ko | Korean NLP Dataset | GitHub | 5.1MB | Other | Commercial use may be conditional |
| ko | Korean QA Dataset | GitHub | 18.5MB | Apache 2.0 | Commercial use allowed (catalog) |
| ko-en | Korean Translation Dataset | GitHub | Large | License unknown | Commercial use allowed (catalog) |
| Tatoeba (ko-en) | Korean Translation Dataset | GitHub | 3.4MB | CC BY 2.0 | Commercial use allowed (catalog) |
| TinyStories | Korean Translation Dataset | Hugging Face | Large | MIT | Commercial use allowed (catalog) |
| KoInFoBench | Korean Instruction Tuning Dataset | Hugging Face | Small | MIT | Commercial use allowed (catalog) |
| Korean Pretraining Corpus | Korean Pretraining Corpus | Hugging Face | Medium | MIT | Commercial use allowed (catalog) |
| Korean Instruction Tuning Dataset | Korean Instruction Tuning Dataset | Hugging Face | Large | MIT | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | Hugging Face | Large | MIT | Commercial use allowed (catalog) |
| QARV | Korean Instruction Tuning Dataset | Hugging Face | Medium | MIT | Commercial use allowed (catalog) |
| Orca DPO | Korean NLP Dataset | Hugging Face | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| KOpen | Korean Translation Dataset | Hugging Face | Medium | MIT | Commercial use allowed (catalog) |
| Korean Instruction Tuning Dataset | Korean Instruction Tuning Dataset | Hugging Face | Large | MIT | Commercial use allowed (catalog) |
| Korean Instruction Tuning Dataset | Korean Instruction Tuning Dataset | Hugging Face | Small | Apache 2.0 | Commercial use allowed (catalog) |
| Korean QA Dataset | Korean QA Dataset | Hugging Face | Small | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Instruction Tuning Dataset | Korean Instruction Tuning Dataset | Hugging Face | Medium | MIT | Commercial use allowed (catalog) |
| Korean Dialogue Dataset | Korean Dialogue Dataset | Hugging Face | Large | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Sentiment Analysis Dataset | Korean Sentiment Analysis Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean Toxicity Dataset | Korean Toxicity Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean Pretraining Corpus | Korean Pretraining Corpus | Hugging Face | Large | Apache 2.0 | Commercial use allowed (catalog) |
| SQuARe | Korean NLP Dataset | GitHub | Medium | MIT | Commercial use allowed (catalog) |
| KoSBi | Korean NLP Dataset | GitHub | Large | MIT | Commercial use allowed (catalog) |
| Korean Instruction Tuning Dataset | Korean Instruction Tuning Dataset | GitHub | Small | Apache 2.0 | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| HateScore | Korean Toxicity Dataset | GitHub | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| LIMA | Korean Instruction Tuning Dataset | Hugging Face | Small | CC BY-NC-SA 4.0 | Commercial use allowed (catalog) |
| Zeroth | Korean Speech Recognition Dataset | GitHub | Large | Apache 2.0 | Commercial use allowed (catalog) |
| MeloTTS | Korean Speech Synthesis Dataset | Hugging Face | Medium | MIT | Commercial use allowed (catalog) |
| Korean LLM Benchmark Dataset | Korean LLM Benchmark Dataset | Hugging Face (SOGANG-ISDS) | Small | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Pretraining Corpus | Korean Pretraining Corpus | Hugging Face (devngho) | Large | MIT | Commercial use allowed (catalog) |
| FineWeb-Edu | Korean Pretraining Corpus | Hugging Face (eliceai) | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| SNU Ko-MuSR | Korean LLM Benchmark Dataset | Hugging Face (thunder-research-group) | Small | MIT | Commercial use allowed (catalog) |
| Korean Instruction Tuning Dataset | Korean Instruction Tuning Dataset | Hugging Face (devngho) | Large | MIT | Commercial use allowed (catalog) |
| Korean Computer Vision Dataset | Korean Computer Vision Dataset | Hugging Face (Nagase-Kotono) | Large (7.85 GB) | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Instruction Tuning Dataset | Korean Instruction Tuning Dataset | Hugging Face (coastral) | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| Function Calling | Korean Instruction Tuning Dataset | Hugging Face (gyung) | Small | Apache 2.0 | Commercial use allowed (catalog) |
| KOTE | Korean Sentiment Analysis Dataset | Hugging Face (searle-j) | Medium | MIT | Commercial use allowed (catalog) |
| NSMC | Korean Sentiment Analysis Dataset | GitHub (e9t) | Small | CC BY 2.0 | Commercial use allowed (catalog) |
| KorQuAD | Korean QA Dataset | KorQuAD | Medium | CC BY-ND 4.0 | Commercial use allowed (catalog) |
| KLUE | Korean NLP Dataset | Hugging Face (KLUE) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| KLUE | Korean NLP Dataset | Hugging Face (KLUE) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| KLUE | Korean LLM Benchmark Dataset | Hugging Face (KLUE) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| KLUE | Korean NER Dataset | Hugging Face (KLUE) | Medium | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| KLUE | Korean NLP Dataset | Hugging Face (KLUE) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| KLUE | Korean QA Dataset | Hugging Face (KLUE) | Medium | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| Smilegate UnSmile | Korean NLP Dataset | GitHub | Small | CC BY-NC-ND 4.0 | Commercial use may be conditional |
| KorQuAD | Korean QA Dataset | KorQuAD | Large ( ) | CC BY-ND 4.0 | Commercial use allowed (catalog) |
| Korean Dialogue Dataset | Korean Dialogue Dataset | GitHub | Medium | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| COPA | Korean LLM Benchmark Dataset | Hugging Face (SKT-brain) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| HellaSwag | Korean LLM Benchmark Dataset | Hugging Face (SKT-brain) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| Korean Pretraining Corpus | Korean Pretraining Corpus | Hugging Face | Small | MIT | Commercial use allowed (catalog) |
| Korean Sentiment Analysis Dataset | Korean Sentiment Analysis Dataset | GitHub (bab2min) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| Korean Sentiment Analysis Dataset | Korean Sentiment Analysis Dataset | GitHub (bab2min) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| KorNLI | Korean NLP Dataset | GitHub (kakaobrain) | Large | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| KorSTS | Korean NLP Dataset | GitHub (kakaobrain) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| BoolQ | Korean QA Dataset | Hugging Face (SKT-brain) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| WiC | Korean NLP Dataset | Hugging Face (SKT-brain) | Small | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| Korean Pretraining Corpus | Korean Pretraining Corpus | GitHub | Medium | MIT | Commercial use allowed (catalog) |
| QG | Korean Instruction Tuning Dataset | Hugging Face | Small | Apache 2.0 | Commercial use allowed (catalog) |
| Korean LLM Benchmark Dataset | Korean LLM Benchmark Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| KMMLU-Pro | Korean LLM Benchmark Dataset | Hugging Face (LGAI-EXAONE) | 12MB | CC BY-NC 4.0 | Commercial use allowed (catalog) |
| Thunder-KoNUBench | Korean LLM Benchmark Dataset | Hugging Face (thunder-research-group) | 5.4MB | CC BY-NC-SA 4.0 | Commercial use allowed (catalog) |
| KoSum | Korean Speech Dataset | Hugging Face (iontail) | 45GB | CC BY-NC 4.0 | Commercial use allowed (catalog) |
| K-BrowseComp | Korean LLM Benchmark Dataset | Hugging Face (prometheus-eval) | 3.5MB | MIT | Commercial use allowed (catalog) |
| WanJuanSiLu | Korean Dialogue Dataset | Hugging Face (opendatalab) | 124 GB | CC BY 4.0 | Commercial use allowed (catalog) |
| Role-Playing | Korean Instruction Tuning Dataset | Hugging Face (huggingface-KREW) | 28MB | Apache 2.0 | Commercial use allowed (catalog) |
| Ko-WinoGrande | Korean Instruction Tuning Dataset | Hugging Face (thunder-research-group) | 4.2MB | CC BY 4.0 | Commercial use allowed (catalog) |
| SmileStyle | Korean Pretraining Corpus | GitHub (smilegate-ai) | 35MB | MIT | Commercial use allowed (catalog) |
| KoIn | Korean Computer Vision Dataset | GitHub (dukong1) | 140 GB | MIT | Commercial use allowed (catalog) |
| Theory of Mind | Korean NLP Dataset | GitHub | 8.5MB | Apache 2.0 | Commercial use allowed (catalog) |
| KoCoSa | Korean Instruction Tuning Dataset | GitHub | 18MB | Apache 2.0 | Commercial use allowed (catalog) |
| KPoEM | Korean Sentiment Analysis Dataset | GitHub | 4.2MB | MIT | Commercial use allowed (catalog) |
| KMRE | Korean NLP Dataset | GitHub | 15MB | Apache 2.0 | Commercial use allowed (catalog) |
| Entailment | Korean NLP Dataset | GitHub | 6.2MB | Apache 2.0 | Commercial use allowed (catalog) |
| Korean SAT Math | Korean NLP Dataset | Hugging Face (Quadyun) | 8.5MB | MIT | Commercial use allowed (catalog) |
| KoelLabs | Korean Speech Recognition Dataset | Hugging Face (KoelLabs) | 18 GB | CC BY-NC 4.0 | Commercial use allowed (catalog) |
| CSAT-KOREAN-2025 | Korean QA Dataset | Hugging Face (KKACHI-HUB) | 3.2MB | MIT | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | Hugging Face (AFOS-Analytics1) | 45MB | CC BY 4.0 | Commercial use allowed (catalog) |
| Nemotron | Korean NLP Dataset | Hugging Face (nvidia) | 120MB | CC BY 4.0 | Commercial use may be conditional |
| Korean Translation Dataset | Korean Translation Dataset | Hugging Face (cfpark00) | 5.5MB | CC BY-NC-SA 4.0 | Commercial use allowed (catalog) |
| Korean Translation Dataset | Korean Translation Dataset | GitHub | Medium | CC0 | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | Hugging Face | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | Hugging Face | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| NLP | Korean NER Dataset | GitHub (Naver-NLP) | Medium | CC0 | Commercial use allowed (catalog) |
| IoT | Korean Speech Dataset | GitHub | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean Speech Dataset | Korean Speech Dataset | AI Hub / GitHub | Large | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Summarization Dataset | Korean Summarization Dataset | Hugging Face / AI Hub | Large | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Sentiment Analysis Dataset | Korean Sentiment Analysis Dataset | Hugging Face | Small | MIT | Commercial use allowed (catalog) |
| SMS | Korean NLP Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean LLM Benchmark Dataset | Korean LLM Benchmark Dataset | Hugging Face / GitHub | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Translation Dataset | Korean Translation Dataset | GitHub / AI Hub | Large | Apache 2.0 | Commercial use allowed (catalog) |
| KoELECTRA | Korean Dialogue Dataset | GitHub (monologg) | Large (34GB) | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Dialogue Dataset | Korean Dialogue Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean Sentiment Analysis Dataset | Korean Sentiment Analysis Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean Sentiment Analysis Dataset | Korean Sentiment Analysis Dataset | GitHub | Small | Apache 2.0 | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | GitHub | Medium | MIT | Commercial use allowed (catalog) |
| Korean Sentiment Analysis Dataset | Korean Sentiment Analysis Dataset | GitHub | Medium | MIT | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | GitHub | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| Korean QA Dataset | Korean QA Dataset | Hugging Face | Medium | MIT | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Meme | Korean NLP Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean Speech Synthesis Dataset | Korean Speech Synthesis Dataset | AI Hub / GitHub | Large | Apache 2.0 | Commercial use allowed (catalog) |
| Korean LLM Benchmark Dataset | Korean LLM Benchmark Dataset | Hugging Face | Small | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Summarization Dataset | Korean Summarization Dataset | Hugging Face | Medium | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| Korean Translation Dataset | Korean Translation Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Ko-Llama | Korean Instruction Tuning Dataset | Hugging Face | Large | Apache 2.0 | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | Hugging Face | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| Gold Standard | Korean NER Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean Pretraining Corpus | Korean Pretraining Corpus | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean Pretraining Corpus | Korean Pretraining Corpus | Hugging Face | Small | Apache 2.0 | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean Sentiment Analysis Dataset | Korean Sentiment Analysis Dataset | GitHub | Medium | MIT | Commercial use allowed (catalog) |
| Korean Pretraining Corpus | Korean Pretraining Corpus | GitHub | Medium | Apache 2.0 | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | GitHub | Small | MIT | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | Hugging Face | Small | Apache 2.0 | Commercial use allowed (catalog) |
| Korean Translation Dataset | Korean Translation Dataset | AI Hub | 140MB | AI Hub terms | Commercial use may be conditional |
| Korean Translation Dataset | Korean Translation Dataset | AI Hub | 320MB | AI Hub terms | Commercial use may be conditional |
| Korean Translation Dataset | Korean Translation Dataset | AI Hub | 120MB | AI Hub terms | Commercial use may be conditional |
| Korean Translation Dataset | Korean Translation Dataset | AI Hub | 410MB | AI Hub terms | Commercial use may be conditional |
| Korean Translation Dataset | Korean Translation Dataset | AI Hub | 620MB | AI Hub terms | Commercial use may be conditional |
| K-POP | Korean Translation Dataset | GitHub | 2.5MB | License unknown | Commercial use allowed (catalog) |
| Korean Translation Dataset | Korean Translation Dataset | GitHub | 18MB | License unknown | Commercial use allowed (catalog) |
| Korean Translation Dataset | Korean Translation Dataset | AI Hub | 2.2GB | AI Hub terms | Commercial use may be conditional |
| Korean Translation Dataset | Korean Translation Dataset | AI Hub | 180MB | AI Hub terms | Commercial use may be conditional |
| Korean Translation Dataset | Korean Translation Dataset | AI Hub | 240MB | AI Hub terms | Commercial use may be conditional |
| Ko-Agent-Trajectories-1.0 | Korean NLP Dataset | Hugging Face | Large | Apache 2.0 | Commercial use allowed (catalog) |
| KoGEM | Korean LLM Benchmark Dataset | Hugging Face | 1.2MB | MIT | Commercial use allowed (catalog) |
| KIT-19 | Korean Instruction Tuning Dataset | Hugging Face | 85MB | CC BY 4.0 | Commercial use allowed (catalog) |
| Ko-MMLU-Pro | Korean LLM Benchmark Dataset | Hugging Face | 45MB | MIT | Commercial use allowed (catalog) |
| Korean NLP Dataset | Korean NLP Dataset | Hugging Face | 120MB | License unknown | Commercial use allowed (catalog) |
| KoSafetyBench | Korean LLM Benchmark Dataset | Hugging Face | 15MB | CC BY 4.0 | Commercial use allowed (catalog) |
| Ko-ARC Challenge | Korean LLM Benchmark Dataset | Hugging Face | 8MB | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| Korean Open-ORPO SFT+DPO | Korean Instruction Tuning Dataset | Hugging Face | 65MB | Apache 2.0 | Commercial use allowed (catalog) |
| HAE-RAE Bench 1.1 | Korean LLM Benchmark Dataset | Hugging Face | 2.9MB | CC BY-NC-ND 4.0 | Commercial use may be conditional |
| HAE-RAE Bench 2.0 | Korean LLM Benchmark Dataset | Hugging Face | 1.0MB | MIT | Commercial use allowed (catalog) |
| KoCommonGEN v2 | Korean LLM Benchmark Dataset | Hugging Face | 183KB | License unknown | Commercial use may be conditional |
| APEACH | Korean LLM Benchmark Dataset | Hugging Face | 1.1MB | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| HRM8K (HAE-RAE Math 8K) | Korean LLM Benchmark Dataset | Hugging Face | 4.9MB | MIT | Commercial use allowed (catalog) |
| K-MHaS | Korean Toxicity Dataset | Hugging Face | 9.1MB | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| K-HATERS | Korean Toxicity Dataset | Hugging Face | 50.5MB | CC BY 4.0 | Commercial use allowed (catalog) |
| KorNAT QA | Korean QA Dataset | Hugging Face | 2.6MB | CC BY-NC 4.0 | Commercial use may be conditional |
| KoAlpaca v1.0 Alpaca | Korean Instruction Tuning Dataset | Hugging Face | 8.1MB | CC BY-NC 4.0 | Commercial use may be conditional |
| KOpen-Platypus | Korean Instruction Tuning Dataset | Hugging Face | 15.9MB | CC BY 4.0 | Commercial use allowed (catalog) |
| OpenAssistant Guanaco | Korean Instruction Tuning Dataset | Hugging Face | 16.7MB | Apache 2.0 | Commercial use allowed (catalog) |
| HRMCR | Korean LLM Benchmark Dataset | Hugging Face | 106KB | Apache 2.0 | Commercial use allowed (catalog) |
| KoBBQ QA | Korean QA Dataset | Hugging Face | 2.4MB | MIT | Commercial use allowed (catalog) |
| KorMedMCQA QA | Korean QA Dataset | Hugging Face | 2.1MB | CC BY-NC 4.0 | Commercial use may be conditional |
| KoBALT-700 | Korean LLM Benchmark Dataset | Hugging Face | 644KB | CC BY-NC 4.0 | Commercial use may be conditional |
| KoSimpleQA QA | Korean QA Dataset | Hugging Face | 301KB | License unknown | Commercial use may be conditional |
| Ko-PIQA | Korean QA Dataset | Hugging Face | 204KB | License unknown | Commercial use may be conditional |
| K-MMBench | Korean Computer Vision Dataset | Hugging Face | 92.4MB | CC BY-NC 4.0 | Commercial use may be conditional |
| Eval | Korean Instruction Tuning Dataset | Hugging Face | 178KB | MIT | Commercial use allowed (catalog) |
| KcBERT | Korean Pretraining Corpus | Hugging Face | 7.90 GiB | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| Korean Summarization Dataset | Korean Summarization Dataset | Hugging Face | 44.4MB | Apache 2.0 | Commercial use allowed (catalog) |
| OpenOrca-KO | Korean Instruction Tuning Dataset | Hugging Face | 21.8MB | MIT | Commercial use allowed (catalog) |
| HAE-RAE CoT 1.5M | Korean NLP Dataset | Hugging Face | 1.05GB | CC BY 4.0 | Commercial use allowed (catalog) |
| KBMC NER | Korean NER Dataset | Hugging Face | 1.4MB | Apache 2.0 | Commercial use allowed (catalog) |
| KOLD | Korean Toxicity Dataset | GitHub | 40,429 comments | License unknown | Commercial use may be conditional |
| KorWikiTableQuestions | Korean QA Dataset | GitHub | 70K QA | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| KoDialogBench | Korean LLM Benchmark Dataset | Hugging Face | 82,962 , 21 | CC BY-NC-SA 4.0 | Commercial use may be conditional |
| FunctionChat-Bench | Korean LLM Benchmark Dataset | GitHub | 500 + 45 | Apache 2.0 | Commercial use allowed (catalog) |
| Jejueo JIT JSS | Korean Translation Dataset | GitHub | 170K+ sentences, 10K | Apache 2.0 | Commercial use allowed (catalog) |
| ParaKQC | Korean NLP Dataset | GitHub | 10,000 (v1) | CC BY-SA 4.0 | Commercial use allowed (catalog) |
| Pansori TEDxKR | Korean Speech Recognition Dataset | GitHub | ~3hours | CC BY-NC-ND 4.0 | Commercial use may be conditional |
| KBL | Korean QA Dataset | Hugging Face | 8.1MB | CC BY-NC 4.0 | Commercial use may be conditional |
| OLKAVS | Korean Speech Recognition Dataset | GitHub | 1,150hours | License unknown | Commercial use may be conditional |
| Korean Pretraining Corpus | Korean Pretraining Corpus | Hugging Face | 764MB | CC BY-SA 4.0 | Commercial use may be conditional |
| xP3x | Korean Instruction Tuning Dataset | Hugging Face | 4,642,468rows | Apache 2.0 | Commercial use may be conditional |