KoBEST
Korean LLM Benchmark Dataset · Source: Hugging Face · License: CC BY-SA 4.0 · 5.1MB · 4.8K items. Catalog only; files stay with the original provider.
KoBEST, KMMLU, OCR, TTS, instruction tuning — filter by license, size, and format, then open the original download.
GearDel is a Korean NLP, speech, and vision discovery catalog. We do not host files.
Korean AI datasets from AI Hub, Hugging Face, GitHub, and NIKL—compare source, license, and size. GearDel does not host files.
Showing 286 of 286
Verification badges show catalog status per field — not every dataset is fully verified.
Korean LLM Benchmark Dataset · Source: Hugging Face · License: CC BY-SA 4.0 · 5.1MB · 4.8K items. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 1000hours · 2,000 · WAV/. Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: Hugging Face · License: CC BY-ND 4.0 · 70.7 MB · 35,030 items. Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: Hugging Face · License: CC BY-SA 4.0 · 62.3 MB · 8. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face · License: CC BY 4.0 · 95MB · 12K. Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: Hugging Face · License: MIT · 1.2MB · 1,524 QA pairs. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face · License: CC BY-NC-SA 4.0 · 12hours+ · WAV 12,853. Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: Hugging Face (LGAI-EXAONE) · License: CC BY-NC 4.0 · 12MB · 2,822 items. Catalog only; files stay with the original provider.
Korean Sentiment Analysis Dataset · Source: GitHub (e9t) · License: CC BY 2.0 · Small · 200K reviews. Catalog only; files stay with the original provider.
Korean QA Dataset · Source: KorQuAD · License: CC BY-ND 4.0 · Medium · 70K items. Catalog only; files stay with the original provider.
Korean Instruction Tuning Dataset · Source: Hugging Face · License: MIT · Large · 250K -. Catalog only; files stay with the original provider.
Korean Toxicity Dataset · Source: Hugging Face · License: CC BY-SA 4.0 · 1.1MB · 11,666 sentences. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: Apache 2.0 · Large · 51.6hours (22.2K ). Catalog only; files stay with the original provider.
Korean Toxicity Dataset · Source: GitHub · License: CC BY-SA 4.0 · 1.2MB · 9.4K comments. Catalog only; files stay with the original provider.
Korean QA Dataset · Source: Hugging Face (SKT-brain) · License: CC BY-SA 4.0 · Small · 1.4K items. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 5TB · 85K. Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: GitHub · License: License unknown · 5.5MB · 2K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: MIT · 11,000. Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: Hugging Face (SKT-brain) · License: CC BY-SA 4.0 · Small · 1K items. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face · License: CC BY 4.0 · Large (7 ). Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: Hugging Face (KKACHI-HUB) · License: MIT · 3.2MB. Catalog only; files stay with the original provider.
Korean Dialogue Dataset · Source: GitHub · License: CC BY-NC-SA 4.0 · 1.8MB · 4.5K dialogues. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: Apache 2.0 · 6.2MB · 15K. Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: Hugging Face · License: MIT · 178KB · 522 items. Catalog only; files stay with the original provider.
Korean Pretraining Corpus · Source: Hugging Face (eliceai) · License: Apache 2.0 · Medium · 30GB. Catalog only; files stay with the original provider.
Korean Instruction Tuning Dataset · Source: Hugging Face (gyung) · License: Apache 2.0 · Small · 1.5K. Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: GitHub · License: Apache 2.0 · 500 + 45. Catalog only; files stay with the original provider.
Korean NER Dataset · Source: GitHub · License: MIT · Small · 10K. Catalog only; files stay with the original provider.
Korean QA Dataset · Source: Hugging Face · License: CC BY 4.0 · 267 MB · 269,837. Catalog only; files stay with the original provider.
Korean Translation Dataset · Source: Hugging Face · License: MIT · 4.8MB · 8.5K items. Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: Hugging Face · License: CC BY-NC-ND 4.0 · 2.9MB · 4,900 items, 13. Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: Hugging Face · License: MIT · 1.0MB · 3,841 items. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face · License: CC BY 4.0 · 1.05GB · 1,586,688. Catalog only; files stay with the original provider.
Korean Toxicity Dataset · Source: GitHub · License: Apache 2.0 · Medium · 11K. Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: Hugging Face (SKT-brain) · License: CC BY-SA 4.0 · Small · 2.5K. Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: Hugging Face · License: MIT · 4.9MB · 8,011 items. Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: Hugging Face · License: Apache 2.0 · 106KB · 100 items. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: Apache 2.0 · Medium · 10K words. Catalog only; files stay with the original provider.
Korean Translation Dataset · Source: GitHub · License: Apache 2.0 · 170K+ sentences, 10K. Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: Hugging Face (prometheus-eval) · License: MIT · 3.5MB · 300 items. Catalog only; files stay with the original provider.
Korean Toxicity Dataset · Source: Hugging Face · License: CC BY 4.0 · 50.5MB · 192,158 comments. Catalog only; files stay with the original provider.
Korean Toxicity Dataset · Source: Hugging Face · License: CC BY-SA 4.0 · 9.1MB · 109,692 comments. Catalog only; files stay with the original provider.
Korean QA Dataset · Source: Hugging Face · License: CC BY-NC 4.0 · 92.4MB · 4,329 items. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: License unknown · 2.5MB · 3.5K. Catalog only; files stay with the original provider.
Korean QA Dataset · Source: Hugging Face · License: CC BY-NC 4.0 · 8.1MB · 3,456 items. Catalog only; files stay with the original provider.
Korean NER Dataset · Source: Hugging Face · License: Apache 2.0 · 1.4MB · 6,150. Catalog only; files stay with the original provider.
Korean Summarization Dataset · Source: AI Hub · License: AI Hub terms · 440MB · 35K. Catalog only; files stay with the original provider.
Korean Pretraining Corpus · Source: Hugging Face · License: CC BY-SA 4.0 · 7.90 GiB · 86,246,285 sentences. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face · License: CC BY 4.0 · 85MB · 100K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face (KLUE) · License: CC BY-SA 4.0 · Small · 12K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face (KLUE) · License: CC BY-SA 4.0 · Small · 25K sentences. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face (KLUE) · License: CC BY-SA 4.0 · Small · 11K sentences. Catalog only; files stay with the original provider.
Korean NER Dataset · Source: Hugging Face (KLUE) · License: CC BY-SA 4.0 · Medium · 21K sentences. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face (KLUE) · License: CC BY-SA 4.0 · Small · 32K sentences. Catalog only; files stay with the original provider.
Korean QA Dataset · Source: Hugging Face (KLUE) · License: CC BY-SA 4.0 · Medium · 29K items. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: Apache 2.0 · 15MB · 12K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: Other · 5.1MB · 5K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: Apache 2.0 · 18.5MB · 11K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face · License: Apache 2.0 · Large · 561K trajectories. Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: Hugging Face · License: CC BY-SA 4.0 · 8MB · 2,590 items. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: License unknown · Large. Catalog only; files stay with the original provider.
Korean Instruction Tuning Dataset · Source: Hugging Face · License: Apache 2.0 · Large · 80K. Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: Hugging Face · License: MIT · 45MB · 12K items. Catalog only; files stay with the original provider.
Korean QA Dataset · Source: Hugging Face · License: License unknown · 204KB · 441 items. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face (thunder-research-group) · License: CC BY 4.0 · 4.2MB · WinoGrande. Catalog only; files stay with the original provider.
Korean Instruction Tuning Dataset · Source: Hugging Face · License: CC BY-NC 4.0 · 8.1MB · 49,620. Catalog only; files stay with the original provider.
Korean Instruction Tuning Dataset · Source: Hugging Face · License: License unknown · Large · 21,155. Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: Hugging Face · License: CC BY-NC 4.0 · 644KB · 700 items. Catalog only; files stay with the original provider.
Korean QA Dataset · Source: Hugging Face · License: MIT · 2.4MB · 81,128 (HF). Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: Hugging Face · License: License unknown · 12MB · 35. Catalog only; files stay with the original provider.
Korean NER Dataset · Source: Hugging Face · License: License unknown · 183KB · 852 public samples. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: Apache 2.0 · 18MB · 25K. Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: Hugging Face · License: CC BY-NC-SA 4.0 · 82,962 , 21. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub (monologg) · License: Apache 2.0 · Large (34GB). Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face (KoelLabs) · License: CC BY-NC 4.0 · 18 GB · 4.5K. Catalog only; files stay with the original provider.
Korean Pretraining Corpus · Source: Hugging Face · License: License unknown · 3.2GB · 40M tokens. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub (dukong1) · License: MIT · 140 GB · 100K. Catalog only; files stay with the original provider.
Korean Instruction Tuning Dataset · Source: Hugging Face · License: MIT · Small · 233 items (60 ). Catalog only; files stay with the original provider.
Korean Toxicity Dataset · Source: GitHub · License: License unknown · 40,429 comments. Catalog only; files stay with the original provider.
Korean Translation Dataset · Source: Hugging Face · License: MIT · Medium · 60K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face · License: CC BY 4.0 · 15.9MB · 24,926. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 12TB · 5M rows. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 850GB · 1.2M. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 2.4TB · 800K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 420GB · 220K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 380GB · 95K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 190GB · 120K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 8TB · 140K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 640GB · 110K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 910GB · 350K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 430GB · 65K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 150GB · 45K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 210GB · 55K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 1.2TB · 400K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 750GB · 180K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 1.1TB · 600K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 1.6TB · 15K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 320GB · 75K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 480GB · 210K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 520GB · 130K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face (Nagase-Kotono) · License: Apache 2.0 · Large (7.85 GB) · 120K -. Catalog only; files stay with the original provider.
Korean Dialogue Dataset · Source: NIKL · License: NIKL terms · Large · 978,342. Catalog only; files stay with the original provider.
Korean Dialogue Dataset · Source: NIKL · License: NIKL terms · 3.5MB · 691,535. Catalog only; files stay with the original provider.
Korean Dialogue Dataset · Source: NIKL · License: NIKL terms · 24.5MB · 12K sessions. Catalog only; files stay with the original provider.
Korean Dialogue Dataset · Source: NIKL · License: NIKL terms · 1.2GB · 5.5M words. Catalog only; files stay with the original provider.
Korean Dialogue Dataset · Source: NIKL · License: NIKL terms · 520MB · 1.5M sentences. Catalog only; files stay with the original provider.
Korean Dialogue Dataset · Source: NIKL · License: NIKL terms · 850MB · 3.2M sentences. Catalog only; files stay with the original provider.
Korean Dialogue Dataset · Source: NIKL · License: NIKL terms · 120MB · 1.1M. Catalog only; files stay with the original provider.
Korean Dialogue Dataset · Source: AI Hub · License: AI Hub terms · 2.4GB · 1.2M articles. Catalog only; files stay with the original provider.
Korean Dialogue Dataset · Source: AI Hub · License: AI Hub terms · 3.5GB · 2.1M. Catalog only; files stay with the original provider.
Korean Dialogue Dataset · Source: AI Hub · License: AI Hub terms · 310MB · 75K. Catalog only; files stay with the original provider.
Korean Dialogue Dataset · Source: Hugging Face · License: Apache 2.0 · Large · 80K dialogues sessions. Catalog only; files stay with the original provider.
Korean Dialogue Dataset · Source: GitHub · License: CC BY-SA 4.0 · Medium · 1.2M sentences. Catalog only; files stay with the original provider.
Korean Dialogue Dataset · Source: GitHub · License: MIT · Small · 5K. Catalog only; files stay with the original provider.
Korean Instruction Tuning Dataset · Source: Hugging Face · License: MIT · Large · 100K. Catalog only; files stay with the original provider.
Korean Instruction Tuning Dataset · Source: Hugging Face · License: Apache 2.0 · Small · 4.8K. Catalog only; files stay with the original provider.
Korean Instruction Tuning Dataset · Source: Hugging Face · License: MIT · Medium · 49K. Catalog only; files stay with the original provider.
Korean Instruction Tuning Dataset · Source: GitHub · License: Apache 2.0 · Small · 5K. Catalog only; files stay with the original provider.
Korean Instruction Tuning Dataset · Source: Hugging Face (devngho) · License: MIT · Large · 15.8 GB. Catalog only; files stay with the original provider.
Korean Instruction Tuning Dataset · Source: Hugging Face (coastral) · License: Apache 2.0 · Medium. Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: Hugging Face (SOGANG-ISDS) · License: Apache 2.0 · Small · 530 items. Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: GitHub · License: MIT · Small · 3K. Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: Hugging Face / GitHub · License: Apache 2.0 · Medium · 18K. Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: Hugging Face · License: Apache 2.0 · Small · 5K items. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 450MB · 110K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 890MB · 120K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 320MB · 250K sentences. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 1.1GB · 45K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 180MB · 45K QA pairs. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 340MB · 110K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 190MB · 35K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 480MB · 90K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 520MB · 14K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 90MB · 120K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 1.5GB · 210K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 110MB · 30K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 130MB · 28K QA pairs. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face · License: MIT · Large · 400K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: MIT · Small · 3K items. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: MIT · Small · 5K sentences. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face (AFOS-Analytics1) · License: CC BY 4.0 · 45MB · 22K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face · License: Apache 2.0 · Medium · 15K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face · License: Apache 2.0 · Medium · 20K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: MIT · Small · 150K words. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: MIT · Small · 8K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: MIT · Small · 8K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: MIT · Medium · 15K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: Apache 2.0 · Medium · 15K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: MIT · Small · 10K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: MIT · Small · 12K. Catalog only; files stay with the original provider.
Korean Dialogue Dataset · Source: Hugging Face · License: Apache 2.0 · Medium · 20K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: MIT · Small · 5K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: MIT · Small · 6K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face · License: Apache 2.0 · Small · 4K items. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face · License: License unknown · 120MB · 50K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 140GB · 1.2M. Catalog only; files stay with the original provider.
Korean Instruction Tuning Dataset · Source: Hugging Face · License: Apache 2.0 · 65MB · 40K pairs. Catalog only; files stay with the original provider.
Korean Pretraining Corpus · Source: Hugging Face · License: MIT · Medium · 5.8K. Catalog only; files stay with the original provider.
Korean Pretraining Corpus · Source: Hugging Face · License: Apache 2.0 · Large · 21K. Catalog only; files stay with the original provider.
Korean Pretraining Corpus · Source: Hugging Face (devngho) · License: MIT · Large · docs. Catalog only; files stay with the original provider.
Korean Pretraining Corpus · Source: Hugging Face · License: MIT · Small · 8K. Catalog only; files stay with the original provider.
Korean Pretraining Corpus · Source: GitHub · License: MIT · Medium · 45K. Catalog only; files stay with the original provider.
Korean Pretraining Corpus · Source: GitHub · License: MIT · Small · 7K. Catalog only; files stay with the original provider.
Korean Pretraining Corpus · Source: Hugging Face · License: Apache 2.0 · Small · 8K. Catalog only; files stay with the original provider.
Korean Pretraining Corpus · Source: GitHub · License: Apache 2.0 · Medium · 15K. Catalog only; files stay with the original provider.
Korean Pretraining Corpus · Source: Hugging Face · License: CC BY-SA 4.0 · 764MB. Catalog only; files stay with the original provider.
Korean QA Dataset · Source: AI Hub · License: AI Hub terms · 180MB · 80K. Catalog only; files stay with the original provider.
Korean QA Dataset · Source: AI Hub · License: AI Hub terms · 140MB · 55K QA pairs. Catalog only; files stay with the original provider.
Korean QA Dataset · Source: AI Hub · License: AI Hub terms · 45MB · 100K. Catalog only; files stay with the original provider.
Korean QA Dataset · Source: AI Hub · License: AI Hub terms · 8.5MB · 5.2K dialogues. Catalog only; files stay with the original provider.
Korean QA Dataset · Source: AI Hub · License: AI Hub terms · 120MB · 65K QA pairs. Catalog only; files stay with the original provider.
Korean QA Dataset · Source: AI Hub · License: AI Hub terms · 85MB · 22K. Catalog only; files stay with the original provider.
Korean QA Dataset · Source: Hugging Face · License: Apache 2.0 · Small · 2.3K. Catalog only; files stay with the original provider.
Korean QA Dataset · Source: Hugging Face · License: MIT · Medium · 40K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face (Quadyun) · License: MIT · 8.5MB · 2K. Catalog only; files stay with the original provider.
Korean Sentiment Analysis Dataset · Source: AI Hub · License: AI Hub terms · 40MB · 150K. Catalog only; files stay with the original provider.
Korean Sentiment Analysis Dataset · Source: GitHub · License: MIT · Small · 11.8K QA pairs. Catalog only; files stay with the original provider.
Korean Sentiment Analysis Dataset · Source: GitHub (bab2min) · License: CC BY-SA 4.0 · Small · 200K reviews. Catalog only; files stay with the original provider.
Korean Sentiment Analysis Dataset · Source: GitHub (bab2min) · License: CC BY-SA 4.0 · Small · 100K. Catalog only; files stay with the original provider.
Korean Sentiment Analysis Dataset · Source: Hugging Face · License: MIT · Small · 25K. Catalog only; files stay with the original provider.
Korean Sentiment Analysis Dataset · Source: GitHub · License: MIT · Small · 8K. Catalog only; files stay with the original provider.
Korean Sentiment Analysis Dataset · Source: GitHub · License: Apache 2.0 · Small · 10K articles. Catalog only; files stay with the original provider.
Korean Sentiment Analysis Dataset · Source: GitHub · License: MIT · Medium · 30K dialogues. Catalog only; files stay with the original provider.
Korean Sentiment Analysis Dataset · Source: GitHub · License: MIT · Medium · 30K reviews. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 500hours · 300K clips. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 200hours · 110K clips. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 80hours · 150K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 800hours · 500K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 600hours · 15K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 240GB · 35K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 400hours · 30K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 4TB · 12K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 50hours · 3.5K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 40GB · 22K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 15GB · 11K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 30GB · 8.5K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 85GB · 18K clips. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 120GB · 45K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 60hours · 12K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 150GB · 5.5K clips. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 8GB · 4.2K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 220hours · 14K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 140hours · 60K files. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub / GitHub · License: Apache 2.0 · Large · 150hours. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 120hours · 80K files. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 150hours · 90K files. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 100hours · 45K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 180hours · 55K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub / GitHub · License: Apache 2.0 · Large · 15K. Catalog only; files stay with the original provider.
Korean Summarization Dataset · Source: AI Hub · License: AI Hub terms · Large · 350K dialogues. Catalog only; files stay with the original provider.
Korean Summarization Dataset · Source: AI Hub · License: AI Hub terms · 620MB · 200K. Catalog only; files stay with the original provider.
Korean Summarization Dataset · Source: AI Hub · License: AI Hub terms · 540MB · 300K. Catalog only; files stay with the original provider.
Korean Summarization Dataset · Source: AI Hub · License: AI Hub terms · 750MB · 150K. Catalog only; files stay with the original provider.
Korean Summarization Dataset · Source: AI Hub · License: AI Hub terms · 410MB · 80K rows. Catalog only; files stay with the original provider.
Korean Summarization Dataset · Source: AI Hub · License: AI Hub terms · 780MB · 220K. Catalog only; files stay with the original provider.
Korean Summarization Dataset · Source: AI Hub · License: AI Hub terms · 85MB · 55K. Catalog only; files stay with the original provider.
Korean Summarization Dataset · Source: Hugging Face / AI Hub · License: Apache 2.0 · Large · 50K. Catalog only; files stay with the original provider.
Korean Summarization Dataset · Source: Hugging Face · License: CC BY-SA 4.0 · Medium · 50K. Catalog only; files stay with the original provider.
Korean Summarization Dataset · Source: Hugging Face · License: Apache 2.0 · 44.4MB · 27,400 articles-. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: License unknown · 8.4GB · 10K. Catalog only; files stay with the original provider.
Korean Toxicity Dataset · Source: GitHub · License: MIT · Small · 5.8K sentences. Catalog only; files stay with the original provider.
Korean Translation Dataset · Source: AI Hub · License: AI Hub terms · 55MB · 300K. Catalog only; files stay with the original provider.
Korean Translation Dataset · Source: AI Hub · License: AI Hub terms · 1.8GB · 1.1M. Catalog only; files stay with the original provider.
Korean Translation Dataset · Source: Hugging Face (cfpark00) · License: CC BY-NC-SA 4.0 · 5.5MB · 4.5K -. Catalog only; files stay with the original provider.
Korean Translation Dataset · Source: GitHub · License: CC0 · Medium · 12K. Catalog only; files stay with the original provider.
Korean Translation Dataset · Source: GitHub / AI Hub · License: Apache 2.0 · Large · 40K. Catalog only; files stay with the original provider.
Korean Translation Dataset · Source: GitHub · License: MIT · Small · 6K. Catalog only; files stay with the original provider.
Korean Translation Dataset · Source: AI Hub · License: AI Hub terms · 140MB · 100K. Catalog only; files stay with the original provider.
Korean Translation Dataset · Source: AI Hub · License: AI Hub terms · 320MB · 180K. Catalog only; files stay with the original provider.
Korean Translation Dataset · Source: AI Hub · License: AI Hub terms · 120MB · 110K. Catalog only; files stay with the original provider.
Korean Translation Dataset · Source: AI Hub · License: AI Hub terms · 410MB · 140K docs. Catalog only; files stay with the original provider.
Korean Translation Dataset · Source: AI Hub · License: AI Hub terms · 620MB · 250K articles. Catalog only; files stay with the original provider.
Korean Translation Dataset · Source: GitHub · License: License unknown · 18MB · 65K. Catalog only; files stay with the original provider.
Korean Translation Dataset · Source: AI Hub · License: AI Hub terms · 2.2GB · 400K. Catalog only; files stay with the original provider.
Korean Translation Dataset · Source: AI Hub · License: AI Hub terms · 180MB · 75K. Catalog only; files stay with the original provider.
Korean Translation Dataset · Source: AI Hub · License: AI Hub terms · 240MB · 55K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: License unknown · 4.2GB · 9,900. Catalog only; files stay with the original provider.
Korean QA Dataset · Source: Hugging Face · License: CC BY-NC 4.0 · 2.1MB · 7,489 items. Catalog only; files stay with the original provider.
Korean QA Dataset · Source: Hugging Face · License: CC BY-NC 4.0 · 2.6MB · 10,007 items. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub (kakaobrain) · License: CC BY-SA 4.0 · Large · 950K sentences. Catalog only; files stay with the original provider.
Korean QA Dataset · Source: KorQuAD · License: CC BY-ND 4.0 · Large ( ) · 100K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub (kakaobrain) · License: CC BY-SA 4.0 · Small · 8.6K sentences. Catalog only; files stay with the original provider.
Korean QA Dataset · Source: GitHub · License: CC BY-SA 4.0 · 70K QA. Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: Hugging Face · License: CC BY 4.0 · 15MB · 5K items. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: MIT · Large · 68K sentences. Catalog only; files stay with the original provider.
Korean QA Dataset · Source: Hugging Face · License: License unknown · 301KB · 1,000 items. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face (iontail) · License: CC BY-NC 4.0 · 45GB · 700. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face (searle-j) · License: MIT · Medium · 50K comments. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: MIT · 4.2MB · 6.5K. Catalog only; files stay with the original provider.
Korean QA Dataset · Source: Hugging Face · License: License unknown · 125MB · 15 2.58K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face · License: Apache 2.0 · 180MB · 152,630. Catalog only; files stay with the original provider.
Korean Instruction Tuning Dataset · Source: Hugging Face · License: CC BY-NC-SA 4.0 · Small · 1K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face · License: MIT · Medium · 30K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: MIT · Small · 5K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face (nvidia) · License: CC BY 4.0 · 120MB · 9.4K. Catalog only; files stay with the original provider.
Korean NER Dataset · Source: GitHub (Naver-NLP) · License: CC0 · Medium · 80K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 1.5TB · 450K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 290GB · 150K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: License unknown · 1,150hours · 1,107 speakers. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face · License: Apache 2.0 · 16.7MB · 10,364 dialogues. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face · License: MIT · 21.8MB · 21,632. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face · License: Apache 2.0 · Medium · 29K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: CC BY-NC-ND 4.0 · ~3hours · 41 speakers. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: CC BY-SA 4.0 · 10,000 (v1). Catalog only; files stay with the original provider.
Korean Instruction Tuning Dataset · Source: Hugging Face · License: MIT · Medium · 33K. Catalog only; files stay with the original provider.
Korean Instruction Tuning Dataset · Source: Hugging Face · License: Apache 2.0 · Small · 15K. Catalog only; files stay with the original provider.
Korean Instruction Tuning Dataset · Source: Hugging Face (huggingface-KREW) · License: Apache 2.0 · 28MB · 15K dialogues sessions. Catalog only; files stay with the original provider.
Korean QA Dataset · Source: Hugging Face · License: Korean VQA License · 100,445 · -Q&A. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: CC BY-NC-ND 4.0 · Small · 18K sentences. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub (smilegate-ai) · License: MIT · 35MB · 21K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: MIT · Small · 12K. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 980MB · 1.2M turns. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face (thunder-research-group) · License: MIT · Small · images. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: License unknown · 3.2GB. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: MIT · Medium · 40K sentences. Catalog only; files stay with the original provider.
Korean QA Dataset · Source: AI Hub · License: AI Hub terms · 230MB · 40K. Catalog only; files stay with the original provider.
Korean QA Dataset · Source: Hugging Face · License: Apache 2.0 · 185MB · LLaVA. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: CC BY 2.0 · 3.4MB. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: GitHub · License: Apache 2.0 · 8.5MB · 8K. Catalog only; files stay with the original provider.
Korean LLM Benchmark Dataset · Source: Hugging Face (thunder-research-group) · License: CC BY-NC-SA 4.0 · 5.4MB · items. Catalog only; files stay with the original provider.
Korean Translation Dataset · Source: Hugging Face · License: MIT · Large · 175K docs. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 300hours · 210K speakers. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 90hours · 35K dialogues. Catalog only; files stay with the original provider.
Korean Pretraining Corpus · Source: Hugging Face (opendatalab) · License: CC BY 4.0 · 124 GB. Catalog only; files stay with the original provider.
Korean NLP Dataset · Source: Hugging Face (SKT-brain) · License: CC BY-SA 4.0 · Small · 3K items. Catalog only; files stay with the original provider.
Korean Instruction Tuning Dataset · Source: Hugging Face · License: Apache 2.0 · 4,642,468rows. Catalog only; files stay with the original provider.
External Data Resources
Beyond the GearDel catalog, these public lists on GitHub and blogs are useful references.
Open a speech, OCR, or benchmark collection first. Then read the license mark and leave for the original host.
ASR is not TTS. Receipt OCR is not scene text. KLUE-family rows are for scores, not pretraining.
Commercial-allowed on GearDel is a catalog hint. AI Hub terms are not MIT.
Check provider, source URL, and size text, then download on the host. GearDel does not keep files.
Full guide: how to choose a dataset →·License comparison →·FAQ →·About GearDel →