AI Dataset DiscoveryLicense · Source · Korean datasets
Korean AI datasets · license, size, original download

286 Korean AI datasets,
license, size, and original download

KoBEST, KMMLU, OCR, TTS, instruction tuning — filter by license, size, and format, then open the original download.GearDel is a Korean NLP, speech, and vision discovery catalog. We do not host files.

Try
Dataset Catalog

Browse AI datasets

Korean AI datasets from AI Hub, Hugging Face, GitHub, and NIKL—compare source, license, and size. GearDel does not host files.

Showing 286 of 286

Verification badges show catalog status per field — not every dataset is fully verified.

🔍
TextKoreanTextbenchmarkcommonsense

KoBEST

Korean LLM Benchmark Dataset · Source: Hugging Face · License: CC BY-SA 4.0 · 5.1MB · 4.8K items. Catalog only; files stay with the original provider.


SourceHugging Face
Size5.1MB
FormatText
LicenseCC BY-SA 4.0
Source verifiedLicense verified
📣
SpeechKoreanAudioASRKsponSpeech

KsponSpeech

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 1000hours · 2,000 · WAV/. Catalog only; files stay with the original provider.


SourceAI Hub
Size1000hours
FormatAudio
LicenseAI Hub terms
Source verifiedLicense verified
💬
TextKoreanTextbenchmarkLLM

KMMLU

Korean LLM Benchmark Dataset · Source: Hugging Face · License: CC BY-ND 4.0 · 70.7 MB · 35,030 items. Catalog only; files stay with the original provider.


SourceHugging Face
Size70.7 MB
FormatText
LicenseCC BY-ND 4.0
Source verifiedLicense verified
⚖
TextKoreanTextbenchmarkNLU

KLUE Benchmark

Korean LLM Benchmark Dataset · Source: Hugging Face · License: CC BY-SA 4.0 · 62.3 MB · 8. Catalog only; files stay with the original provider.


SourceHugging Face
Size62.3 MB
FormatText
LicenseCC BY-SA 4.0
Source verifiedLicense verified
🚏
ImageKoreanImageOCR

Korean Receipts OCR

Korean NLP Dataset · Source: Hugging Face · License: CC BY 4.0 · 95MB · 12K. Catalog only; files stay with the original provider.


SourceHugging Face
Size95MB
FormatImage
LicenseCC BY 4.0
Source verifiedLicense verified
✦
TextKoreanTextbenchmarkNLU

KoGEM

Korean LLM Benchmark Dataset · Source: Hugging Face · License: MIT · 1.2MB · 1,524 QA pairs. Catalog only; files stay with the original provider.


SourceHugging Face
Size1.2MB
FormatText
LicenseMIT
Source verifiedLicense needs review
📻
SpeechKoreanAudioTTSTTS

KSS Dataset

Korean NLP Dataset · Source: Hugging Face · License: CC BY-NC-SA 4.0 · 12hours+ · WAV 12,853. Catalog only; files stay with the original provider.


SourceHugging Face
Size12hours+
FormatAudio
LicenseCC BY-NC-SA 4.0
Source verifiedLicense verified
🔍
TextKoreanTextbenchmarkreasoning

KMMLU-Pro

Korean LLM Benchmark Dataset · Source: Hugging Face (LGAI-EXAONE) · License: CC BY-NC 4.0 · 12MB · 2,822 items. Catalog only; files stay with the original provider.


SourceHugging Face (LGAI-EXAONE)
Size12MB
FormatText
LicenseCC BY-NC 4.0
Source verifiedLicense verified
🔍
TextKoreanTextmovie-reviewssentiment

NSMC

Korean Sentiment Analysis Dataset · Source: GitHub (e9t) · License: CC BY 2.0 · Small · 200K reviews. Catalog only; files stay with the original provider.


SourceGitHub (e9t)
SizeSmall
FormatText
LicenseCC BY 2.0
Source verifiedLicense verified
⚖
TextKoreanTextMRCQA

KorQuAD

Korean QA Dataset · Source: KorQuAD · License: CC BY-ND 4.0 · Medium · 70K items. Catalog only; files stay with the original provider.


SourceKorQuAD
SizeMedium
FormatText
LicenseCC BY-ND 4.0
Source verifiedLicense verified
✍
TextKoreanTextSFT

Korean Instruction Tuning Dataset

Korean Instruction Tuning Dataset · Source: Hugging Face · License: MIT · Large · 250K -. Catalog only; files stay with the original provider.


SourceHugging Face
SizeLarge
FormatText
LicenseMIT
Source verifiedLicense verified
📰
TextKoreanTexthate-speechsafety

APEACH

Korean Toxicity Dataset · Source: Hugging Face · License: CC BY-SA 4.0 · 1.1MB · 11,666 sentences. Catalog only; files stay with the original provider.


SourceHugging Face
Size1.1MB
FormatText
LicenseCC BY-SA 4.0
Source verifiedLicense needs review
📻
SpeechKoreanAudioASRASR

Zeroth

Korean NLP Dataset · Source: GitHub · License: Apache 2.0 · Large · 51.6hours (22.2K ). Catalog only; files stay with the original provider.


SourceGitHub
SizeLarge
FormatAudio
LicenseApache 2.0
Source verifiedLicense verified
📖
TextKoreanTexthate-speechsafety

BEEP! Korean Hate Speech

Korean Toxicity Dataset · Source: GitHub · License: CC BY-SA 4.0 · 1.2MB · 9.4K comments. Catalog only; files stay with the original provider.


SourceGitHub
Size1.2MB
FormatText
LicenseCC BY-SA 4.0
Source verifiedLicense verified
🔍
TextKoreanTextBoolQMRC

BoolQ

Korean QA Dataset · Source: Hugging Face (SKT-brain) · License: CC BY-SA 4.0 · Small · 1.4K items. Catalog only; files stay with the original provider.


SourceHugging Face (SKT-brain)
SizeSmall
FormatText
LicenseCC BY-SA 4.0
Source verifiedLicense verified
🏥
ImageKoreanImage

CCTV

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 5TB · 85K. Catalog only; files stay with the original provider.


SourceAI Hub
Size5TB
FormatImage
LicenseAI Hub terms
Source not verifiedLicense verified
🔍
TextKoreanTextbenchmark

CLIcK (Cultural & Legal Knowledge)

Korean LLM Benchmark Dataset · Source: GitHub · License: License unknown · 5.5MB · 2K. Catalog only; files stay with the original provider.


SourceGitHub
Size5.5MB
FormatText
LicenseLicense unknown
Source verifiedLicense not verified
🎵
SpeechKoreanAudioASR

ClovaCall

Korean NLP Dataset · Source: GitHub · License: MIT · 11,000. Catalog only; files stay with the original provider.


SourceGitHub
Size11,000
FormatAudio
LicenseMIT
Source verifiedLicense verified
📝
TextKoreanTextKoBEST

COPA

Korean LLM Benchmark Dataset · Source: Hugging Face (SKT-brain) · License: CC BY-SA 4.0 · Small · 1K items. Catalog only; files stay with the original provider.


SourceHugging Face (SKT-brain)
SizeSmall
FormatText
LicenseCC BY-SA 4.0
Source verifiedLicense verified
👁
ImageKoreanImageCLIPfeatured

COYO-700M

Korean NLP Dataset · Source: Hugging Face · License: CC BY 4.0 · Large (7 ). Catalog only; files stay with the original provider.


SourceHugging Face
SizeLarge (7 )
FormatImage
LicenseCC BY 4.0
Source verifiedLicense verified
🦙
TextKoreanTextbenchmark

CSAT-KOREAN-2025

Korean LLM Benchmark Dataset · Source: Hugging Face (KKACHI-HUB) · License: MIT · 3.2MB. Catalog only; files stay with the original provider.


SourceHugging Face (KKACHI-HUB)
Size3.2MB
FormatText
LicenseMIT
Source verifiedLicense verified
📖
TextKoreanText

DKTC (Dataset of Korean Threat Dialogue)

Korean Dialogue Dataset · Source: GitHub · License: CC BY-NC-SA 4.0 · 1.8MB · 4.5K dialogues. Catalog only; files stay with the original provider.


SourceGitHub
Size1.8MB
FormatText
LicenseCC BY-NC-SA 4.0
Source verifiedLicense verified
💡
TextKoreanTextNLI

Entailment

Korean NLP Dataset · Source: GitHub · License: Apache 2.0 · 6.2MB · 15K. Catalog only; files stay with the original provider.


SourceGitHub
Size6.2MB
FormatText
LicenseApache 2.0
Source not verifiedLicense needs review
📝
TextKoreanTextbenchmark

Eval

Korean LLM Benchmark Dataset · Source: Hugging Face · License: MIT · 178KB · 522 items. Catalog only; files stay with the original provider.


SourceHugging Face
Size178KB
FormatText
LicenseMIT
Source verifiedLicense needs review
🔍
TextKoreanTextFineWeb

FineWeb-Edu

Korean Pretraining Corpus · Source: Hugging Face (eliceai) · License: Apache 2.0 · Medium · 30GB. Catalog only; files stay with the original provider.


SourceHugging Face (eliceai)
SizeMedium
FormatText
LicenseApache 2.0
Source verifiedLicense verified
📰
TextKoreanTextAPIToolUse

Function Calling

Korean Instruction Tuning Dataset · Source: Hugging Face (gyung) · License: Apache 2.0 · Small · 1.5K. Catalog only; files stay with the original provider.


SourceHugging Face (gyung)
SizeSmall
FormatText
LicenseApache 2.0
Source verifiedLicense verified
⚖
TextKoreanTextbenchmark

FunctionChat-Bench

Korean LLM Benchmark Dataset · Source: GitHub · License: Apache 2.0 · 500 + 45. Catalog only; files stay with the original provider.


SourceGitHub
SizeUnknown
FormatText
LicenseApache 2.0
Source verifiedLicense needs review
✍
TextKoreanTextNER

Gold Standard

Korean NER Dataset · Source: GitHub · License: MIT · Small · 10K. Catalog only; files stay with the original provider.


SourceGitHub
SizeSmall
FormatText
LicenseMIT
Source not verifiedLicense needs review
✦
TextKoreanTextQA

GovOn Legal Response Dataset

Korean QA Dataset · Source: Hugging Face · License: CC BY 4.0 · 267 MB · 269,837. Catalog only; files stay with the original provider.


SourceHugging Face
Size267 MB
FormatText
LicenseCC BY 4.0
Source verifiedLicense verified
🦙
TextKoreanTextmathLLM

GSM8K

Korean Translation Dataset · Source: Hugging Face · License: MIT · 4.8MB · 8.5K items. Catalog only; files stay with the original provider.


SourceHugging Face
Size4.8MB
FormatText
LicenseMIT
Source verifiedLicense verified
🔍
TextKoreanTextbenchmarkLLM

HAE-RAE Bench 1.1

Korean LLM Benchmark Dataset · Source: Hugging Face · License: CC BY-NC-ND 4.0 · 2.9MB · 4,900 items, 13. Catalog only; files stay with the original provider.


SourceHugging Face
Size2.9MB
FormatText
LicenseCC BY-NC-ND 4.0
Source verifiedLicense needs review
⚖
TextKoreanTextbenchmarkreasoning

HAE-RAE Bench 2.0

Korean LLM Benchmark Dataset · Source: Hugging Face · License: MIT · 1.0MB · 3,841 items. Catalog only; files stay with the original provider.


SourceHugging Face
Size1.0MB
FormatText
LicenseMIT
Source verifiedLicense needs review
⚖
TextKoreanTextLLMCoT

HAE-RAE CoT 1.5M

Korean NLP Dataset · Source: Hugging Face · License: CC BY 4.0 · 1.05GB · 1,586,688. Catalog only; files stay with the original provider.


SourceHugging Face
Size1.05GB
FormatText
LicenseCC BY 4.0
Source verifiedLicense needs review
🦙
TextKoreanText

HateScore

Korean Toxicity Dataset · Source: GitHub · License: Apache 2.0 · Medium · 11K. Catalog only; files stay with the original provider.


SourceGitHub
SizeMedium
FormatText
LicenseApache 2.0
Source verifiedLicense verified
💡
TextKoreanTextcommonsenseKoBEST

HellaSwag

Korean LLM Benchmark Dataset · Source: Hugging Face (SKT-brain) · License: CC BY-SA 4.0 · Small · 2.5K. Catalog only; files stay with the original provider.


SourceHugging Face (SKT-brain)
SizeSmall
FormatText
LicenseCC BY-SA 4.0
Source verifiedLicense verified
💬
TextKoreanTextmathbenchmark

HRM8K (HAE-RAE Math 8K)

Korean LLM Benchmark Dataset · Source: Hugging Face · License: MIT · 4.9MB · 8,011 items. Catalog only; files stay with the original provider.


SourceHugging Face
Size4.9MB
FormatText
LicenseMIT
Source verifiedLicense needs review
💬
TextKoreanTextbenchmarkcommonsense

HRMCR

Korean LLM Benchmark Dataset · Source: Hugging Face · License: Apache 2.0 · 106KB · 100 items. Catalog only; files stay with the original provider.


SourceHugging Face
Size106KB
FormatText
LicenseApache 2.0
Source verifiedLicense needs review
📞
SpeechKoreanAudiosmart-home

IoT

Korean NLP Dataset · Source: GitHub · License: Apache 2.0 · Medium · 10K words. Catalog only; files stay with the original provider.


SourceGitHub
SizeMedium
FormatAudio
LicenseApache 2.0
Source not verifiedLicense needs review
📝
TextKoreanTexttranslationTTS

Jejueo JIT JSS

Korean Translation Dataset · Source: GitHub · License: Apache 2.0 · 170K+ sentences, 10K. Catalog only; files stay with the original provider.


SourceGitHub
SizeUnknown
FormatText
LicenseApache 2.0
Source verifiedLicense needs review
✦
TextKoreanTextbenchmark

K-BrowseComp

Korean LLM Benchmark Dataset · Source: Hugging Face (prometheus-eval) · License: MIT · 3.5MB · 300 items. Catalog only; files stay with the original provider.


SourceHugging Face (prometheus-eval)
Size3.5MB
FormatText
LicenseMIT
Source verifiedLicense verified
💡
TextKoreanTexthate-speechsafety

K-HATERS

Korean Toxicity Dataset · Source: Hugging Face · License: CC BY 4.0 · 50.5MB · 192,158 comments. Catalog only; files stay with the original provider.


SourceHugging Face
Size50.5MB
FormatText
LicenseCC BY 4.0
Source verifiedLicense needs review
🦙
TextKoreanTexthate-speechmulti-label

K-MHaS

Korean Toxicity Dataset · Source: Hugging Face · License: CC BY-SA 4.0 · 9.1MB · 109,692 comments. Catalog only; files stay with the original provider.


SourceHugging Face
Size9.1MB
FormatText
LicenseCC BY-SA 4.0
Source verifiedLicense needs review
🏥
ImageKoreanImageVLMbenchmark

K-MMBench

Korean QA Dataset · Source: Hugging Face · License: CC BY-NC 4.0 · 92.4MB · 4,329 items. Catalog only; files stay with the original provider.


SourceHugging Face
Size92.4MB
FormatImage
LicenseCC BY-NC 4.0
Source verifiedLicense needs review
📰
TextKoreanTextKPOP

K-POP

Korean NLP Dataset · Source: GitHub · License: License unknown · 2.5MB · 3.5K. Catalog only; files stay with the original provider.


SourceGitHub
Size2.5MB
FormatText
LicenseLicense unknown
Source not verifiedLicense not verified
⚖
TextKoreanTextlegalbenchmark

KBL

Korean QA Dataset · Source: Hugging Face · License: CC BY-NC 4.0 · 8.1MB · 3,456 items. Catalog only; files stay with the original provider.


SourceHugging Face
Size8.1MB
FormatText
LicenseCC BY-NC 4.0
Source verifiedLicense needs review
⚖
TextKoreanTextmedicalNER

KBMC NER

Korean NER Dataset · Source: Hugging Face · License: Apache 2.0 · 1.4MB · 6,150. Catalog only; files stay with the original provider.


SourceHugging Face
Size1.4MB
FormatText
LicenseApache 2.0
Source verifiedLicense needs review
📝
TextKoreanTextbroadcastdebate-summary

KBS MBC

Korean Summarization Dataset · Source: AI Hub · License: AI Hub terms · 440MB · 35K. Catalog only; files stay with the original provider.


SourceAI Hub
Size440MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
💬
TextKoreanTextpretrainingspoken

KcBERT

Korean Pretraining Corpus · Source: Hugging Face · License: CC BY-SA 4.0 · 7.90 GiB · 86,246,285 sentences. Catalog only; files stay with the original provider.


SourceHugging Face
Size7.90 GiB
FormatText
LicenseCC BY-SA 4.0
Source verifiedLicense needs review
💡
TextKoreanTextLLMNLU

KIT-19

Korean NLP Dataset · Source: Hugging Face · License: CC BY 4.0 · 85MB · 100K. Catalog only; files stay with the original provider.


SourceHugging Face
Size85MB
FormatText
LicenseCC BY 4.0
Source verifiedLicense needs review
🔍
TextKoreanTextKLUE

KLUE

Korean NLP Dataset · Source: Hugging Face (KLUE) · License: CC BY-SA 4.0 · Small · 12K. Catalog only; files stay with the original provider.


SourceHugging Face (KLUE)
SizeSmall
FormatText
LicenseCC BY-SA 4.0
Source verifiedLicense verified
⚖
TextKoreanTextNLIKLUE

KLUE

Korean NLP Dataset · Source: Hugging Face (KLUE) · License: CC BY-SA 4.0 · Small · 25K sentences. Catalog only; files stay with the original provider.


SourceHugging Face (KLUE)
SizeSmall
FormatText
LicenseCC BY-SA 4.0
Source verifiedLicense verified
📰
TextKoreanTextSTSKLUE

KLUE

Korean NLP Dataset · Source: Hugging Face (KLUE) · License: CC BY-SA 4.0 · Small · 11K sentences. Catalog only; files stay with the original provider.


SourceHugging Face (KLUE)
SizeSmall
FormatText
LicenseCC BY-SA 4.0
Source verifiedLicense verified
✍
TextKoreanTextNERKLUE

KLUE

Korean NER Dataset · Source: Hugging Face (KLUE) · License: CC BY-SA 4.0 · Medium · 21K sentences. Catalog only; files stay with the original provider.


SourceHugging Face (KLUE)
SizeMedium
FormatText
LicenseCC BY-SA 4.0
Source verifiedLicense verified
✍
TextKoreanTextREKLUE

KLUE

Korean NLP Dataset · Source: Hugging Face (KLUE) · License: CC BY-SA 4.0 · Small · 32K sentences. Catalog only; files stay with the original provider.


SourceHugging Face (KLUE)
SizeSmall
FormatText
LicenseCC BY-SA 4.0
Source verifiedLicense verified
📖
TextKoreanTextMRCMRC

KLUE

Korean QA Dataset · Source: Hugging Face (KLUE) · License: CC BY-SA 4.0 · Medium · 29K items. Catalog only; files stay with the original provider.


SourceHugging Face (KLUE)
SizeMedium
FormatText
LicenseCC BY-SA 4.0
Source verifiedLicense verified
📝
TextKoreanText

KMRE

Korean NLP Dataset · Source: GitHub · License: Apache 2.0 · 15MB · 12K. Catalog only; files stay with the original provider.


SourceGitHub
Size15MB
FormatText
LicenseApache 2.0
Source not verifiedLicense needs review
💡
TextKoreanText

ko

Korean NLP Dataset · Source: GitHub · License: Other · 5.1MB · 5K. Catalog only; files stay with the original provider.


SourceGitHub
Size5.1MB
FormatText
LicenseOther
Source verifiedLicense verified
📖
TextKoreanText

ko

Korean NLP Dataset · Source: GitHub · License: Apache 2.0 · 18.5MB · 11K. Catalog only; files stay with the original provider.


SourceGitHub
Size18.5MB
FormatText
LicenseApache 2.0
Source verifiedLicense verified
📰
TextKoreanTextDPO

Ko-Agent-Trajectories-1.0

Korean NLP Dataset · Source: Hugging Face · License: Apache 2.0 · Large · 561K trajectories. Catalog only; files stay with the original provider.


SourceHugging Face
SizeLarge
FormatText
LicenseApache 2.0
Source verifiedLicense needs review
💡
TextKoreanTextbenchmark

Ko-ARC Challenge

Korean LLM Benchmark Dataset · Source: Hugging Face · License: CC BY-SA 4.0 · 8MB · 2,590 items. Catalog only; files stay with the original provider.


SourceHugging Face
Size8MB
FormatText
LicenseCC BY-SA 4.0
Source verifiedLicense needs review
📝
TextKoreanTextspoken

ko-en

Korean NLP Dataset · Source: GitHub · License: License unknown · Large. Catalog only; files stay with the original provider.


SourceGitHub
SizeLarge
FormatText
LicenseLicense unknown
Source not verifiedLicense not verified
🦙
TextKoreanTextLlamaSFT

Ko-Llama

Korean Instruction Tuning Dataset · Source: Hugging Face · License: Apache 2.0 · Large · 80K. Catalog only; files stay with the original provider.


SourceHugging Face
SizeLarge
FormatText
LicenseApache 2.0
Source not verifiedLicense needs review
📰
TextKoreanTextbenchmarkMMLU

Ko-MMLU-Pro

Korean LLM Benchmark Dataset · Source: Hugging Face · License: MIT · 45MB · 12K items. Catalog only; files stay with the original provider.


SourceHugging Face
Size45MB
FormatText
LicenseMIT
Source verifiedLicense needs review
🔍
TextKoreanTextbenchmarkcommonsense

Ko-PIQA

Korean QA Dataset · Source: Hugging Face · License: License unknown · 204KB · 441 items. Catalog only; files stay with the original provider.


SourceHugging Face
Size204KB
FormatText
LicenseLicense unknown
Source verifiedLicense not verified
📰
TextKoreanTextICML2026

Ko-WinoGrande

Korean NLP Dataset · Source: Hugging Face (thunder-research-group) · License: CC BY 4.0 · 4.2MB · WinoGrande. Catalog only; files stay with the original provider.


SourceHugging Face (thunder-research-group)
Size4.2MB
FormatText
LicenseCC BY 4.0
Source verifiedLicense verified
💡
TextKoreanTextLLMSFT

KoAlpaca v1.0 Alpaca

Korean Instruction Tuning Dataset · Source: Hugging Face · License: CC BY-NC 4.0 · 8.1MB · 49,620. Catalog only; files stay with the original provider.


SourceHugging Face
Size8.1MB
FormatText
LicenseCC BY-NC 4.0
Source verifiedLicense needs review
✦
TextKoreanTextLLM

KoAlpaca-v1.1a

Korean Instruction Tuning Dataset · Source: Hugging Face · License: License unknown · Large · 21,155. Catalog only; files stay with the original provider.


SourceHugging Face
SizeLarge
FormatText
LicenseLicense unknown
Source verifiedLicense not verified
✍
TextKoreanTextbenchmarkNLU

KoBALT-700

Korean LLM Benchmark Dataset · Source: Hugging Face · License: CC BY-NC 4.0 · 644KB · 700 items. Catalog only; files stay with the original provider.


SourceHugging Face
Size644KB
FormatText
LicenseCC BY-NC 4.0
Source verifiedLicense needs review
💡
TextKoreanTextsafetybenchmark

KoBBQ QA

Korean QA Dataset · Source: Hugging Face · License: MIT · 2.4MB · 81,128 (HF). Catalog only; files stay with the original provider.


SourceHugging Face
Size2.4MB
FormatText
LicenseMIT
Source verifiedLicense needs review
📝
TextKoreanTextbenchmarkLLM

KoBigBench Lite

Korean LLM Benchmark Dataset · Source: Hugging Face · License: License unknown · 12MB · 35. Catalog only; files stay with the original provider.


SourceHugging Face
Size12MB
FormatText
LicenseLicense unknown
Source not verifiedLicense not verified
✦
TextKoreanTextbenchmarkcommonsense

KoCommonGEN v2

Korean NER Dataset · Source: Hugging Face · License: License unknown · 183KB · 852 public samples. Catalog only; files stay with the original provider.


SourceHugging Face
Size183KB
FormatText
LicenseLicense unknown
Source verifiedLicense not verified
📝
TextKoreanText

KoCoSa

Korean NLP Dataset · Source: GitHub · License: Apache 2.0 · 18MB · 25K. Catalog only; files stay with the original provider.


SourceGitHub
Size18MB
FormatText
LicenseApache 2.0
Source not verifiedLicense needs review
💬
TextKoreanTextdialoguebenchmark

KoDialogBench

Korean LLM Benchmark Dataset · Source: Hugging Face · License: CC BY-NC-SA 4.0 · 82,962 , 21. Catalog only; files stay with the original provider.


SourceHugging Face
SizeUnknown
FormatText
LicenseCC BY-NC-SA 4.0
Source verifiedLicense needs review
💡
TextKoreanText

KoELECTRA

Korean NLP Dataset · Source: GitHub (monologg) · License: Apache 2.0 · Large (34GB). Catalog only; files stay with the original provider.


SourceGitHub (monologg)
SizeLarge (34GB)
FormatText
LicenseApache 2.0
Source verifiedLicense verified
📞
SpeechKoreanAudioASR

KoelLabs

Korean NLP Dataset · Source: Hugging Face (KoelLabs) · License: CC BY-NC 4.0 · 18 GB · 4.5K. Catalog only; files stay with the original provider.


SourceHugging Face (KoelLabs)
Size18 GB
FormatAudio
LicenseCC BY-NC 4.0
Source verifiedLicense verified
📖
TextKoreanTextpretrainingGPT2

KoGPT2

Korean Pretraining Corpus · Source: Hugging Face · License: License unknown · 3.2GB · 40M tokens. Catalog only; files stay with the original provider.


SourceHugging Face
Size3.2GB
FormatText
LicenseLicense unknown
Source not verifiedLicense not verified
🎨
ImageKoreanImage

KoIn

Korean NLP Dataset · Source: GitHub (dukong1) · License: MIT · 140 GB · 100K. Catalog only; files stay with the original provider.


SourceGitHub (dukong1)
Size140 GB
FormatImage
LicenseMIT
Source verifiedLicense verified
📰
TextKoreanTextbenchmarkeval

KoInFoBench

Korean Instruction Tuning Dataset · Source: Hugging Face · License: MIT · Small · 233 items (60 ). Catalog only; files stay with the original provider.


SourceHugging Face
SizeSmall
FormatText
LicenseMIT
Source verifiedLicense verified
📰
TextKoreanTexthate-speechsafety

KOLD

Korean Toxicity Dataset · Source: GitHub · License: License unknown · 40,429 comments. Catalog only; files stay with the original provider.


SourceGitHub
SizeUnknown
FormatText
LicenseLicense unknown
Source verifiedLicense not verified
💡
TextKoreanTextcodinglogic

KOpen

Korean Translation Dataset · Source: Hugging Face · License: MIT · Medium · 60K. Catalog only; files stay with the original provider.


SourceHugging Face
SizeMedium
FormatText
LicenseMIT
Source verifiedLicense verified
🔍
TextKoreanTextLLMSTEM

KOpen-Platypus

Korean NLP Dataset · Source: Hugging Face · License: CC BY 4.0 · 15.9MB · 24,926. Catalog only; files stay with the original provider.


SourceHugging Face
Size15.9MB
FormatText
LicenseCC BY 4.0
Source verifiedLicense needs review
🧮
ImageKoreanImage

Korean Computer Vision Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 12TB · 5M rows. Catalog only; files stay with the original provider.


SourceAI Hub
Size12TB
FormatImage
LicenseAI Hub terms
Source not verifiedLicense verified
🎨
ImageKoreanImage

Korean Computer Vision Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 850GB · 1.2M. Catalog only; files stay with the original provider.


SourceAI Hub
Size850GB
FormatImage
LicenseAI Hub terms
Source not verifiedLicense verified
📦
ImageKoreanImage

Korean Computer Vision Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 2.4TB · 800K. Catalog only; files stay with the original provider.


SourceAI Hub
Size2.4TB
FormatImage
LicenseAI Hub terms
Source not verifiedLicense verified
🎨
ImageKoreanImage

Korean Computer Vision Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 420GB · 220K. Catalog only; files stay with the original provider.


SourceAI Hub
Size420GB
FormatImage
LicenseAI Hub terms
Source not verifiedLicense verified
🖼
ImageKoreanImage

Korean Computer Vision Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 380GB · 95K. Catalog only; files stay with the original provider.


SourceAI Hub
Size380GB
FormatImage
LicenseAI Hub terms
Source not verifiedLicense verified
🎨
ImageKoreanImagepests

Korean Computer Vision Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 190GB · 120K. Catalog only; files stay with the original provider.


SourceAI Hub
Size190GB
FormatImage
LicenseAI Hub terms
Source not verifiedLicense verified
🏔
ImageKoreanImage

Korean Computer Vision Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 8TB · 140K. Catalog only; files stay with the original provider.


SourceAI Hub
Size8TB
FormatImage
LicenseAI Hub terms
Source not verifiedLicense verified
👁
ImageKoreanImage

Korean Computer Vision Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 640GB · 110K. Catalog only; files stay with the original provider.


SourceAI Hub
Size640GB
FormatImage
LicenseAI Hub terms
Source not verifiedLicense verified
🖼
ImageKoreanImage

Korean Computer Vision Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 910GB · 350K. Catalog only; files stay with the original provider.


SourceAI Hub
Size910GB
FormatImage
LicenseAI Hub terms
Source not verifiedLicense verified
👁
ImageKoreanImage

Korean Computer Vision Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 430GB · 65K. Catalog only; files stay with the original provider.


SourceAI Hub
Size430GB
FormatImage
LicenseAI Hub terms
Source not verifiedLicense verified
🎨
ImageKoreanImage

Korean Computer Vision Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 150GB · 45K. Catalog only; files stay with the original provider.


SourceAI Hub
Size150GB
FormatImage
LicenseAI Hub terms
Source not verifiedLicense verified
🚏
ImageKoreanImage

Korean Computer Vision Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 210GB · 55K. Catalog only; files stay with the original provider.


SourceAI Hub
Size210GB
FormatImage
LicenseAI Hub terms
Source not verifiedLicense verified
🧮
ImageKoreanImage

Korean Computer Vision Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 1.2TB · 400K. Catalog only; files stay with the original provider.


SourceAI Hub
Size1.2TB
FormatImage
LicenseAI Hub terms
Source not verifiedLicense verified
📦
ImageKoreanImage

Korean Computer Vision Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 750GB · 180K. Catalog only; files stay with the original provider.


SourceAI Hub
Size750GB
FormatImage
LicenseAI Hub terms
Source not verifiedLicense verified
🏔
ImageKoreanImagetourism

Korean Computer Vision Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 1.1TB · 600K. Catalog only; files stay with the original provider.


SourceAI Hub
Size1.1TB
FormatImage
LicenseAI Hub terms
Source not verifiedLicense verified
🎨
ImageKoreanImagePap smear

Korean Computer Vision Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 1.6TB · 15K. Catalog only; files stay with the original provider.


SourceAI Hub
Size1.6TB
FormatImage
LicenseAI Hub terms
Source not verifiedLicense verified
👁
ImageKoreanImagehealthcare

Korean Computer Vision Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 320GB · 75K. Catalog only; files stay with the original provider.


SourceAI Hub
Size320GB
FormatImage
LicenseAI Hub terms
Source not verifiedLicense verified
🖼
ImageKoreanImage

Korean Computer Vision Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 480GB · 210K. Catalog only; files stay with the original provider.


SourceAI Hub
Size480GB
FormatImage
LicenseAI Hub terms
Source not verifiedLicense verified
🚏
ImageKoreanImage

Korean Computer Vision Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 520GB · 130K. Catalog only; files stay with the original provider.


SourceAI Hub
Size520GB
FormatImage
LicenseAI Hub terms
Source not verifiedLicense verified
🧮
ImageKoreanImage

Korean Computer Vision Dataset

Korean NLP Dataset · Source: Hugging Face (Nagase-Kotono) · License: Apache 2.0 · Large (7.85 GB) · 120K -. Catalog only; files stay with the original provider.


SourceHugging Face (Nagase-Kotono)
SizeLarge (7.85 GB)
FormatImage
LicenseApache 2.0
Source verifiedLicense verified
✍
TextKoreanTextpublicstandard-Korean

Korean Dialogue Dataset

Korean Dialogue Dataset · Source: NIKL · License: NIKL terms · Large · 978,342. Catalog only; files stay with the original provider.


SourceNIKL
SizeLarge
FormatText
LicenseNIKL terms
Source not verifiedLicense verified
✍
TextKoreanTextpublic

Korean Dialogue Dataset

Korean Dialogue Dataset · Source: NIKL · License: NIKL terms · 3.5MB · 691,535. Catalog only; files stay with the original provider.


SourceNIKL
Size3.5MB
FormatText
LicenseNIKL terms
Source not verifiedLicense verified
📰
TextKoreanTextpublicdaily-chat

Korean Dialogue Dataset

Korean Dialogue Dataset · Source: NIKL · License: NIKL terms · 24.5MB · 12K sessions. Catalog only; files stay with the original provider.


SourceNIKL
Size24.5MB
FormatText
LicenseNIKL terms
Source not verifiedLicense verified
⚖
TextKoreanTextpublicweb

Korean Dialogue Dataset

Korean Dialogue Dataset · Source: NIKL · License: NIKL terms · 1.2GB · 5.5M words. Catalog only; files stay with the original provider.


SourceNIKL
Size1.2GB
FormatText
LicenseNIKL terms
Source not verifiedLicense verified
📰
TextKoreanTextpublicspoken

Korean Dialogue Dataset

Korean Dialogue Dataset · Source: NIKL · License: NIKL terms · 520MB · 1.5M sentences. Catalog only; files stay with the original provider.


SourceNIKL
Size520MB
FormatText
LicenseNIKL terms
Source not verifiedLicense verified
🔍
TextKoreanTextpublicwritten

Korean Dialogue Dataset

Korean Dialogue Dataset · Source: NIKL · License: NIKL terms · 850MB · 3.2M sentences. Catalog only; files stay with the original provider.


SourceNIKL
Size850MB
FormatText
LicenseNIKL terms
Source not verifiedLicense verified
💬
TextKoreanTextpublicSNS

Korean Dialogue Dataset

Korean Dialogue Dataset · Source: NIKL · License: NIKL terms · 120MB · 1.1M. Catalog only; files stay with the original provider.


SourceNIKL
Size120MB
FormatText
LicenseNIKL terms
Source not verifiedLicense verified
✍
TextKoreanTextnewspretraining

Korean Dialogue Dataset

Korean Dialogue Dataset · Source: AI Hub · License: AI Hub terms · 2.4GB · 1.2M articles. Catalog only; files stay with the original provider.


SourceAI Hub
Size2.4GB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
✦
TextKoreanTextdomainpretraining

Korean Dialogue Dataset

Korean Dialogue Dataset · Source: AI Hub · License: AI Hub terms · 3.5GB · 2.1M. Catalog only; files stay with the original provider.


SourceAI Hub
Size3.5GB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
📖
TextKoreanTextpersonachatbot

Korean Dialogue Dataset

Korean Dialogue Dataset · Source: AI Hub · License: AI Hub terms · 310MB · 75K. Catalog only; files stay with the original provider.


SourceAI Hub
Size310MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
📰
TextKoreanTextspokendialogue

Korean Dialogue Dataset

Korean Dialogue Dataset · Source: Hugging Face · License: Apache 2.0 · Large · 80K dialogues sessions. Catalog only; files stay with the original provider.


SourceHugging Face
SizeLarge
FormatText
LicenseApache 2.0
Source verifiedLicense verified
🦙
TextKoreanTextwikicleaned-text

Korean Dialogue Dataset

Korean Dialogue Dataset · Source: GitHub · License: CC BY-SA 4.0 · Medium · 1.2M sentences. Catalog only; files stay with the original provider.


SourceGitHub
SizeMedium
FormatText
LicenseCC BY-SA 4.0
Source not verifiedLicense needs review
🦙
TextKoreanText

Korean Dialogue Dataset

Korean Dialogue Dataset · Source: GitHub · License: MIT · Small · 5K. Catalog only; files stay with the original provider.


SourceGitHub
SizeSmall
FormatText
LicenseMIT
Source not verifiedLicense needs review
✍
TextKoreanTextcommercialinstruction

Korean Instruction Tuning Dataset

Korean Instruction Tuning Dataset · Source: Hugging Face · License: MIT · Large · 100K. Catalog only; files stay with the original provider.


SourceHugging Face
SizeLarge
FormatText
LicenseMIT
Source verifiedLicense verified
🔍
TextKoreanTextCarrotAIassistant

Korean Instruction Tuning Dataset

Korean Instruction Tuning Dataset · Source: Hugging Face · License: Apache 2.0 · Small · 4.8K. Catalog only; files stay with the original provider.


SourceHugging Face
SizeSmall
FormatText
LicenseApache 2.0
Source verifiedLicense verified
📖
TextKoreanTextfinanceAlpaca

Korean Instruction Tuning Dataset

Korean Instruction Tuning Dataset · Source: Hugging Face · License: MIT · Medium · 49K. Catalog only; files stay with the original provider.


SourceHugging Face
SizeMedium
FormatText
LicenseMIT
Source verifiedLicense verified
✦
TextKoreanText

Korean Instruction Tuning Dataset

Korean Instruction Tuning Dataset · Source: GitHub · License: Apache 2.0 · Small · 5K. Catalog only; files stay with the original provider.


SourceGitHub
SizeSmall
FormatText
LicenseApache 2.0
Source verifiedLicense verified
💡
TextKoreanTextSFT

Korean Instruction Tuning Dataset

Korean Instruction Tuning Dataset · Source: Hugging Face (devngho) · License: MIT · Large · 15.8 GB. Catalog only; files stay with the original provider.


SourceHugging Face (devngho)
SizeLarge
FormatText
LicenseMIT
Source verifiedLicense verified
🔍
TextKoreanTextinstruction

Korean Instruction Tuning Dataset

Korean Instruction Tuning Dataset · Source: Hugging Face (coastral) · License: Apache 2.0 · Medium. Catalog only; files stay with the original provider.


SourceHugging Face (coastral)
SizeMedium
FormatText
LicenseApache 2.0
Source verifiedLicense verified
📖
TextKoreanTextbenchmark

Korean LLM Benchmark Dataset

Korean LLM Benchmark Dataset · Source: Hugging Face (SOGANG-ISDS) · License: Apache 2.0 · Small · 530 items. Catalog only; files stay with the original provider.


SourceHugging Face (SOGANG-ISDS)
SizeSmall
FormatText
LicenseApache 2.0
Source verifiedLicense verified
📖
TextKoreanTextproverbs

Korean LLM Benchmark Dataset

Korean LLM Benchmark Dataset · Source: GitHub · License: MIT · Small · 3K. Catalog only; files stay with the original provider.


SourceGitHub
SizeSmall
FormatText
LicenseMIT
Source not verifiedLicense needs review
📝
TextKoreanTextcommonsenseeval

Korean LLM Benchmark Dataset

Korean LLM Benchmark Dataset · Source: Hugging Face / GitHub · License: Apache 2.0 · Medium · 18K. Catalog only; files stay with the original provider.


SourceHugging Face / GitHub
SizeMedium
FormatText
LicenseApache 2.0
Source not verifiedLicense needs review
✍
TextKoreanTextcommonsense

Korean LLM Benchmark Dataset

Korean LLM Benchmark Dataset · Source: Hugging Face · License: Apache 2.0 · Small · 5K items. Catalog only; files stay with the original provider.


SourceHugging Face
SizeSmall
FormatText
LicenseApache 2.0
Source not verifiedLicense needs review
✍
TextKoreanTextcustomer-supportsmall-business

Korean NLP Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 450MB · 110K. Catalog only; files stay with the original provider.


SourceAI Hub
Size450MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
⚖
TextKoreanTextpatentsdomain

Korean NLP Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 890MB · 120K. Catalog only; files stay with the original provider.


SourceAI Hub
Size890MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
✍
TextKoreanTextdialectstandard-Korean

Korean NLP Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 320MB · 250K sentences. Catalog only; files stay with the original provider.


SourceAI Hub
Size320MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
🦙
TextKoreanTextcase-lawlegal

Korean NLP Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 1.1GB · 45K. Catalog only; files stay with the original provider.


SourceAI Hub
Size1.1GB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
📰
TextKoreanTexttaxlegal

Korean NLP Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 180MB · 45K QA pairs. Catalog only; files stay with the original provider.


SourceAI Hub
Size180MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
💡
TextKoreanTextreal-estatelistings

Korean NLP Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 340MB · 110K. Catalog only; files stay with the original provider.


SourceAI Hub
Size340MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
🔍
TextKoreanTexteducationtextbook

Korean NLP Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 190MB · 35K. Catalog only; files stay with the original provider.


SourceAI Hub
Size190MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
💡
TextKoreanTextmedicalhealthcare

Korean NLP Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 480MB · 90K. Catalog only; files stay with the original provider.


SourceAI Hub
Size480MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
💬
TextKoreanTextscreenplayscript

Korean NLP Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 520MB · 14K. Catalog only; files stay with the original provider.


SourceAI Hub
Size520MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
💡
TextKoreanTextmarketingads

Korean NLP Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 90MB · 120K. Catalog only; files stay with the original provider.


SourceAI Hub
Size90MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
🦙
TextKoreanTextpatentsengineering

Korean NLP Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 1.5GB · 210K. Catalog only; files stay with the original provider.


SourceAI Hub
Size1.5GB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
📰
TextKoreanTextagriculturepests

Korean NLP Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 110MB · 30K. Catalog only; files stay with the original provider.


SourceAI Hub
Size110MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
🦙
TextKoreanTextSMEtech-support

Korean NLP Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 130MB · 28K QA pairs. Catalog only; files stay with the original provider.


SourceAI Hub
Size130MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
📝
TextKoreanText

Korean NLP Dataset

Korean NLP Dataset · Source: Hugging Face · License: MIT · Large · 400K. Catalog only; files stay with the original provider.


SourceHugging Face
SizeLarge
FormatText
LicenseMIT
Source verifiedLicense verified
🔍
TextKoreanTextmath

Korean NLP Dataset

Korean NLP Dataset · Source: GitHub · License: MIT · Small · 3K items. Catalog only; files stay with the original provider.


SourceGitHub
SizeSmall
FormatText
LicenseMIT
Source verifiedLicense verified
💬
TextKoreanTextproverbsidioms

Korean NLP Dataset

Korean NLP Dataset · Source: GitHub · License: MIT · Small · 5K sentences. Catalog only; files stay with the original provider.


SourceGitHub
SizeSmall
FormatText
LicenseMIT
Source not verifiedLicense needs review
💡
TextKoreanText

Korean NLP Dataset

Korean NLP Dataset · Source: Hugging Face (AFOS-Analytics1) · License: CC BY 4.0 · 45MB · 22K. Catalog only; files stay with the original provider.


SourceHugging Face (AFOS-Analytics1)
Size45MB
FormatText
LicenseCC BY 4.0
Source verifiedLicense verified
💬
TextKoreanText

Korean NLP Dataset

Korean NLP Dataset · Source: Hugging Face · License: Apache 2.0 · Medium · 15K. Catalog only; files stay with the original provider.


SourceHugging Face
SizeMedium
FormatText
LicenseApache 2.0
Source not verifiedLicense needs review
💬
TextKoreanText

Korean NLP Dataset

Korean NLP Dataset · Source: Hugging Face · License: Apache 2.0 · Medium · 20K. Catalog only; files stay with the original provider.


SourceHugging Face
SizeMedium
FormatText
LicenseApache 2.0
Source not verifiedLicense needs review
✍
TextKoreanTextG2P

Korean NLP Dataset

Korean NLP Dataset · Source: GitHub · License: MIT · Small · 150K words. Catalog only; files stay with the original provider.


SourceGitHub
SizeSmall
FormatText
LicenseMIT
Source not verifiedLicense needs review
🔍
TextKoreanText

Korean NLP Dataset

Korean NLP Dataset · Source: GitHub · License: MIT · Small · 8K. Catalog only; files stay with the original provider.


SourceGitHub
SizeSmall
FormatText
LicenseMIT
Source not verifiedLicense needs review
📖
TextKoreanText

Korean NLP Dataset

Korean NLP Dataset · Source: GitHub · License: MIT · Small · 8K. Catalog only; files stay with the original provider.


SourceGitHub
SizeSmall
FormatText
LicenseMIT
Source not verifiedLicense needs review
📝
TextKoreanText

Korean NLP Dataset

Korean NLP Dataset · Source: GitHub · License: MIT · Medium · 15K. Catalog only; files stay with the original provider.


SourceGitHub
SizeMedium
FormatText
LicenseMIT
Source not verifiedLicense needs review
✍
TextKoreanText

Korean NLP Dataset

Korean NLP Dataset · Source: GitHub · License: Apache 2.0 · Medium · 15K. Catalog only; files stay with the original provider.


SourceGitHub
SizeMedium
FormatText
LicenseApache 2.0
Source not verifiedLicense needs review
✦
TextKoreanText

Korean NLP Dataset

Korean NLP Dataset · Source: GitHub · License: MIT · Small · 10K. Catalog only; files stay with the original provider.


SourceGitHub
SizeSmall
FormatText
LicenseMIT
Source not verifiedLicense needs review
📖
TextKoreanTextcustomer-support

Korean NLP Dataset

Korean NLP Dataset · Source: GitHub · License: MIT · Small · 12K. Catalog only; files stay with the original provider.


SourceGitHub
SizeSmall
FormatText
LicenseMIT
Source not verifiedLicense needs review
💡
TextKoreanTextchatbot

Korean NLP Dataset

Korean Dialogue Dataset · Source: Hugging Face · License: Apache 2.0 · Medium · 20K. Catalog only; files stay with the original provider.


SourceHugging Face
SizeMedium
FormatText
LicenseApache 2.0
Source not verifiedLicense needs review
✦
TextKoreanText

Korean NLP Dataset

Korean NLP Dataset · Source: GitHub · License: MIT · Small · 5K. Catalog only; files stay with the original provider.


SourceGitHub
SizeSmall
FormatText
LicenseMIT
Source not verifiedLicense needs review
✍
TextKoreanText

Korean NLP Dataset

Korean NLP Dataset · Source: GitHub · License: MIT · Small · 6K. Catalog only; files stay with the original provider.


SourceGitHub
SizeSmall
FormatText
LicenseMIT
Source not verifiedLicense needs review
📝
TextKoreanTextalignment

Korean NLP Dataset

Korean NLP Dataset · Source: Hugging Face · License: Apache 2.0 · Small · 4K items. Catalog only; files stay with the original provider.


SourceHugging Face
SizeSmall
FormatText
LicenseApache 2.0
Source not verifiedLicense needs review
💬
TextKoreanTextmathreasoning

Korean NLP Dataset

Korean NLP Dataset · Source: Hugging Face · License: License unknown · 120MB · 50K. Catalog only; files stay with the original provider.


SourceHugging Face
Size120MB
FormatText
LicenseLicense unknown
Source verifiedLicense not verified
📦
ImageKoreanImage

Korean OCR Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 140GB · 1.2M. Catalog only; files stay with the original provider.


SourceAI Hub
Size140GB
FormatImage
LicenseAI Hub terms
Source not verifiedLicense verified
📰
TextKoreanTextLLMDPO

Korean Open-ORPO SFT+DPO

Korean Instruction Tuning Dataset · Source: Hugging Face · License: Apache 2.0 · 65MB · 40K pairs. Catalog only; files stay with the original provider.


SourceHugging Face
Size65MB
FormatText
LicenseApache 2.0
Source verifiedLicense needs review
💬
TextKoreanTextpretrainingWikipedia

Korean Pretraining Corpus

Korean Pretraining Corpus · Source: Hugging Face · License: MIT · Medium · 5.8K. Catalog only; files stay with the original provider.


SourceHugging Face
SizeMedium
FormatText
LicenseMIT
Source verifiedLicense verified
🔍
TextKoreanTexttextbookeducation

Korean Pretraining Corpus

Korean Pretraining Corpus · Source: Hugging Face · License: Apache 2.0 · Large · 21K. Catalog only; files stay with the original provider.


SourceHugging Face
SizeLarge
FormatText
LicenseApache 2.0
Source verifiedLicense verified
✍
TextKoreanTextpretrainingeducation

Korean Pretraining Corpus

Korean Pretraining Corpus · Source: Hugging Face (devngho) · License: MIT · Large · docs. Catalog only; files stay with the original provider.


SourceHugging Face (devngho)
SizeLarge
FormatText
LicenseMIT
Source verifiedLicense verified
💬
TextKoreanText

Korean Pretraining Corpus

Korean Pretraining Corpus · Source: Hugging Face · License: MIT · Small · 8K. Catalog only; files stay with the original provider.


SourceHugging Face
SizeSmall
FormatText
LicenseMIT
Source not verifiedLicense needs review
🔍
TextKoreanText

Korean Pretraining Corpus

Korean Pretraining Corpus · Source: GitHub · License: MIT · Medium · 45K. Catalog only; files stay with the original provider.


SourceGitHub
SizeMedium
FormatText
LicenseMIT
Source not verifiedLicense needs review
✦
TextKoreanText

Korean Pretraining Corpus

Korean Pretraining Corpus · Source: GitHub · License: MIT · Small · 7K. Catalog only; files stay with the original provider.


SourceGitHub
SizeSmall
FormatText
LicenseMIT
Source not verifiedLicense needs review
📖
TextKoreanText

Korean Pretraining Corpus

Korean Pretraining Corpus · Source: Hugging Face · License: Apache 2.0 · Small · 8K. Catalog only; files stay with the original provider.


SourceHugging Face
SizeSmall
FormatText
LicenseApache 2.0
Source not verifiedLicense needs review
⚖
TextKoreanText

Korean Pretraining Corpus

Korean Pretraining Corpus · Source: GitHub · License: Apache 2.0 · Medium · 15K. Catalog only; files stay with the original provider.


SourceGitHub
SizeMedium
FormatText
LicenseApache 2.0
Source not verifiedLicense needs review
⚖
TextKoreanTextpretrainingeducation

Korean Pretraining Corpus

Korean Pretraining Corpus · Source: Hugging Face · License: CC BY-SA 4.0 · 764MB. Catalog only; files stay with the original provider.


SourceHugging Face
Size764MB
FormatText
LicenseCC BY-SA 4.0
Source verifiedLicense needs review
✍
TextKoreanTextMRCgovernment

Korean QA Dataset

Korean QA Dataset · Source: AI Hub · License: AI Hub terms · 180MB · 80K. Catalog only; files stay with the original provider.


SourceAI Hub
Size180MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
✍
TextKoreanTextfinanceMRC

Korean QA Dataset

Korean QA Dataset · Source: AI Hub · License: AI Hub terms · 140MB · 55K QA pairs. Catalog only; files stay with the original provider.


SourceAI Hub
Size140MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
✦
TextKoreanTextcommonsenseQA

Korean QA Dataset

Korean QA Dataset · Source: AI Hub · License: AI Hub terms · 45MB · 100K. Catalog only; files stay with the original provider.


SourceAI Hub
Size45MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
🔍
TextKoreanTextwellnesscounseling

Korean QA Dataset

Korean QA Dataset · Source: AI Hub · License: AI Hub terms · 8.5MB · 5.2K dialogues. Catalog only; files stay with the original provider.


SourceAI Hub
Size8.5MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
🔍
TextKoreanTextretail

Korean QA Dataset

Korean QA Dataset · Source: AI Hub · License: AI Hub terms · 120MB · 65K QA pairs. Catalog only; files stay with the original provider.


SourceAI Hub
Size120MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
✦
TextKoreanTextgamesQA

Korean QA Dataset

Korean QA Dataset · Source: AI Hub · License: AI Hub terms · 85MB · 22K. Catalog only; files stay with the original provider.


SourceAI Hub
Size85MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
🦙
TextKoreanTexteval

Korean QA Dataset

Korean QA Dataset · Source: Hugging Face · License: Apache 2.0 · Small · 2.3K. Catalog only; files stay with the original provider.


SourceHugging Face
SizeSmall
FormatText
LicenseApache 2.0
Source verifiedLicense verified
📝
TextKoreanText

Korean QA Dataset

Korean QA Dataset · Source: Hugging Face · License: MIT · Medium · 40K. Catalog only; files stay with the original provider.


SourceHugging Face
SizeMedium
FormatText
LicenseMIT
Source not verifiedLicense needs review
⚖
TextKoreanTextmath

Korean SAT Math

Korean NLP Dataset · Source: Hugging Face (Quadyun) · License: MIT · 8.5MB · 2K. Catalog only; files stay with the original provider.


SourceHugging Face (Quadyun)
Size8.5MB
FormatText
LicenseMIT
Source verifiedLicense verified
📝
TextKoreanTextfinancestocks

Korean Sentiment Analysis Dataset

Korean Sentiment Analysis Dataset · Source: AI Hub · License: AI Hub terms · 40MB · 150K. Catalog only; files stay with the original provider.


SourceAI Hub
Size40MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
⚖
TextKoreanTextchatbot

Korean Sentiment Analysis Dataset

Korean Sentiment Analysis Dataset · Source: GitHub · License: MIT · Small · 11.8K QA pairs. Catalog only; files stay with the original provider.


SourceGitHub
SizeSmall
FormatText
LicenseMIT
Source verifiedLicense verified
✍
TextKoreanText

Korean Sentiment Analysis Dataset

Korean Sentiment Analysis Dataset · Source: GitHub (bab2min) · License: CC BY-SA 4.0 · Small · 200K reviews. Catalog only; files stay with the original provider.


SourceGitHub (bab2min)
SizeSmall
FormatText
LicenseCC BY-SA 4.0
Source verifiedLicense verified
✦
TextKoreanTextspoken

Korean Sentiment Analysis Dataset

Korean Sentiment Analysis Dataset · Source: GitHub (bab2min) · License: CC BY-SA 4.0 · Small · 100K. Catalog only; files stay with the original provider.


SourceGitHub (bab2min)
SizeSmall
FormatText
LicenseCC BY-SA 4.0
Source verifiedLicense verified
⚖
TextKoreanText

Korean Sentiment Analysis Dataset

Korean Sentiment Analysis Dataset · Source: Hugging Face · License: MIT · Small · 25K. Catalog only; files stay with the original provider.


SourceHugging Face
SizeSmall
FormatText
LicenseMIT
Source not verifiedLicense needs review
✦
TextKoreanText

Korean Sentiment Analysis Dataset

Korean Sentiment Analysis Dataset · Source: GitHub · License: MIT · Small · 8K. Catalog only; files stay with the original provider.


SourceGitHub
SizeSmall
FormatText
LicenseMIT
Source not verifiedLicense needs review
💬
TextKoreanTextsentiment

Korean Sentiment Analysis Dataset

Korean Sentiment Analysis Dataset · Source: GitHub · License: Apache 2.0 · Small · 10K articles. Catalog only; files stay with the original provider.


SourceGitHub
SizeSmall
FormatText
LicenseApache 2.0
Source not verifiedLicense needs review
📝
TextKoreanTextchatbot

Korean Sentiment Analysis Dataset

Korean Sentiment Analysis Dataset · Source: GitHub · License: MIT · Medium · 30K dialogues. Catalog only; files stay with the original provider.


SourceGitHub
SizeMedium
FormatText
LicenseMIT
Source not verifiedLicense needs review
📖
TextKoreanTextmulti-labelsentiment

Korean Sentiment Analysis Dataset

Korean Sentiment Analysis Dataset · Source: GitHub · License: MIT · Medium · 30K reviews. Catalog only; files stay with the original provider.


SourceGitHub
SizeMedium
FormatText
LicenseMIT
Source not verifiedLicense needs review
🎙
SpeechKoreanAudiodialect-speechtranscripts

Korean Speech Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 500hours · 300K clips. Catalog only; files stay with the original provider.


SourceAI Hub
Size500hours
FormatAudio
LicenseAI Hub terms
Source not verifiedLicense verified
🎙
SpeechKoreanAudioin-cardenoising

Korean Speech Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 200hours · 110K clips. Catalog only; files stay with the original provider.


SourceAI Hub
Size200hours
FormatAudio
LicenseAI Hub terms
Source not verifiedLicense verified
🎙
SpeechKoreanAudiosmart-homeshort-commands

Korean Speech Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 80hours · 150K. Catalog only; files stay with the original provider.


SourceAI Hub
Size80hours
FormatAudio
LicenseAI Hub terms
Source not verifiedLicense verified
🎧
SpeechKoreanAudiocall-centercustomer-support

Korean Speech Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 800hours · 500K. Catalog only; files stay with the original provider.


SourceAI Hub
Size800hours
FormatAudio
LicenseAI Hub terms
Source not verifiedLicense verified
📞
SpeechKoreanAudiomeetingsdiarization

Korean Speech Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 600hours · 15K. Catalog only; files stay with the original provider.


SourceAI Hub
Size600hours
FormatAudio
LicenseAI Hub terms
Source not verifiedLicense verified
📣
SpeechKoreanAudioambient-noiseaudio-classification

Korean Speech Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 240GB · 35K. Catalog only; files stay with the original provider.


SourceAI Hub
Size240GB
FormatAudio
LicenseAI Hub terms
Source not verifiedLicense verified
🎵
SpeechKoreanAudiolecturesacademic-speech

Korean Speech Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 400hours · 30K. Catalog only; files stay with the original provider.


SourceAI Hub
Size400hours
FormatAudio
LicenseAI Hub terms
Source not verifiedLicense verified
📻
SpeechKoreanAudiosign-languagevideo

Korean Speech Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 4TB · 12K. Catalog only; files stay with the original provider.


SourceAI Hub
Size4TB
FormatAudio
LicenseAI Hub terms
Source not verifiedLicense verified
🎙
SpeechKoreanAudio

Korean Speech Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 50hours · 3.5K. Catalog only; files stay with the original provider.


SourceAI Hub
Size50hours
FormatAudio
LicenseAI Hub terms
Source not verifiedLicense verified
📞
SpeechKoreanAudio

Korean Speech Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 40GB · 22K. Catalog only; files stay with the original provider.


SourceAI Hub
Size40GB
FormatAudio
LicenseAI Hub terms
Source not verifiedLicense verified
🎙
SpeechKoreanAudiohealthcare

Korean Speech Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 15GB · 11K. Catalog only; files stay with the original provider.


SourceAI Hub
Size15GB
FormatAudio
LicenseAI Hub terms
Source not verifiedLicense verified
📣
SpeechKoreanAudio

Korean Speech Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 30GB · 8.5K. Catalog only; files stay with the original provider.


SourceAI Hub
Size30GB
FormatAudio
LicenseAI Hub terms
Source not verifiedLicense verified
📞
SpeechKoreanAudio

Korean Speech Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 85GB · 18K clips. Catalog only; files stay with the original provider.


SourceAI Hub
Size85GB
FormatAudio
LicenseAI Hub terms
Source not verifiedLicense verified
🎵
SpeechKoreanAudio

Korean Speech Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 120GB · 45K. Catalog only; files stay with the original provider.


SourceAI Hub
Size120GB
FormatAudio
LicenseAI Hub terms
Source not verifiedLicense verified
🎵
SpeechKoreanAudio

Korean Speech Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 60hours · 12K. Catalog only; files stay with the original provider.


SourceAI Hub
Size60hours
FormatAudio
LicenseAI Hub terms
Source not verifiedLicense verified
🗣
SpeechKoreanAudio

Korean Speech Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 150GB · 5.5K clips. Catalog only; files stay with the original provider.


SourceAI Hub
Size150GB
FormatAudio
LicenseAI Hub terms
Source not verifiedLicense verified
🎭
SpeechKoreanAudio

Korean Speech Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 8GB · 4.2K. Catalog only; files stay with the original provider.


SourceAI Hub
Size8GB
FormatAudio
LicenseAI Hub terms
Source not verifiedLicense verified
🎵
SpeechKoreanAudio

Korean Speech Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 220hours · 14K. Catalog only; files stay with the original provider.


SourceAI Hub
Size220hours
FormatAudio
LicenseAI Hub terms
Source not verifiedLicense verified
🎭
SpeechKoreanAudio

Korean Speech Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 140hours · 60K files. Catalog only; files stay with the original provider.


SourceAI Hub
Size140hours
FormatAudio
LicenseAI Hub terms
Source not verifiedLicense verified
🎭
SpeechKoreanAudiodialect

Korean Speech Dataset

Korean NLP Dataset · Source: AI Hub / GitHub · License: Apache 2.0 · Large · 150hours. Catalog only; files stay with the original provider.


SourceAI Hub / GitHub
SizeLarge
FormatAudio
LicenseApache 2.0
Source not verifiedLicense needs review
🗣
SpeechKoreanAudiochild-speechASR

Korean Speech Recognition Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 120hours · 80K files. Catalog only; files stay with the original provider.


SourceAI Hub
Size120hours
FormatAudio
LicenseAI Hub terms
Source not verifiedLicense verified
🎵
SpeechKoreanAudioelderly-speechASR

Korean Speech Recognition Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 150hours · 90K files. Catalog only; files stay with the original provider.


SourceAI Hub
Size150hours
FormatAudio
LicenseAI Hub terms
Source not verifiedLicense verified
🎧
SpeechKoreanAudiostandard-KoreanTTS

Korean Speech Synthesis Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 100hours · 45K. Catalog only; files stay with the original provider.


SourceAI Hub
Size100hours
FormatAudio
LicenseAI Hub terms
Source not verifiedLicense verified
🗣
SpeechKoreanAudioaudiobookexpressive-TTS

Korean Speech Synthesis Dataset

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 180hours · 55K. Catalog only; files stay with the original provider.


SourceAI Hub
Size180hours
FormatAudio
LicenseAI Hub terms
Source not verifiedLicense verified
🎭
SpeechKoreanAudioexpressive-TTS

Korean Speech Synthesis Dataset

Korean NLP Dataset · Source: AI Hub / GitHub · License: Apache 2.0 · Large · 15K. Catalog only; files stay with the original provider.


SourceAI Hub / GitHub
SizeLarge
FormatAudio
LicenseApache 2.0
Source not verifiedLicense needs review
🔍
TextKoreanTextfeatureddialogue

Korean Summarization Dataset

Korean Summarization Dataset · Source: AI Hub · License: AI Hub terms · Large · 350K dialogues. Catalog only; files stay with the original provider.


SourceAI Hub
SizeLarge
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
📖
TextKoreanTextsummarizationbooks

Korean Summarization Dataset

Korean Summarization Dataset · Source: AI Hub · License: AI Hub terms · 620MB · 200K. Catalog only; files stay with the original provider.


SourceAI Hub
Size620MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
🔍
TextKoreanTextsummarizationnews

Korean Summarization Dataset

Korean Summarization Dataset · Source: AI Hub · License: AI Hub terms · 540MB · 300K. Catalog only; files stay with the original provider.


SourceAI Hub
Size540MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
✦
TextKoreanTextsummarizationacademic

Korean Summarization Dataset

Korean Summarization Dataset · Source: AI Hub · License: AI Hub terms · 750MB · 150K. Catalog only; files stay with the original provider.


SourceAI Hub
Size750MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
📰
TextKoreanTextgovernmentsummarization

Korean Summarization Dataset

Korean Summarization Dataset · Source: AI Hub · License: AI Hub terms · 410MB · 80K rows. Catalog only; files stay with the original provider.


SourceAI Hub
Size410MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
✦
TextKoreanTextcall-centersummarization

Korean Summarization Dataset

Korean Summarization Dataset · Source: AI Hub · License: AI Hub terms · 780MB · 220K. Catalog only; files stay with the original provider.


SourceAI Hub
Size780MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
⚖
TextKoreanTexttourismpromo

Korean Summarization Dataset

Korean Summarization Dataset · Source: AI Hub · License: AI Hub terms · 85MB · 55K. Catalog only; files stay with the original provider.


SourceAI Hub
Size85MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
💬
TextKoreanTextnews

Korean Summarization Dataset

Korean Summarization Dataset · Source: Hugging Face / AI Hub · License: Apache 2.0 · Large · 50K. Catalog only; files stay with the original provider.


SourceHugging Face / AI Hub
SizeLarge
FormatText
LicenseApache 2.0
Source not verifiedLicense needs review
📝
TextKoreanText

Korean Summarization Dataset

Korean Summarization Dataset · Source: Hugging Face · License: CC BY-SA 4.0 · Medium · 50K. Catalog only; files stay with the original provider.


SourceHugging Face
SizeMedium
FormatText
LicenseCC BY-SA 4.0
Source not verifiedLicense needs review
✍
TextKoreanTextsummarizationnews

Korean Summarization Dataset

Korean Summarization Dataset · Source: Hugging Face · License: Apache 2.0 · 44.4MB · 27,400 articles-. Catalog only; files stay with the original provider.


SourceHugging Face
Size44.4MB
FormatText
LicenseApache 2.0
Source verifiedLicense needs review
📸
ImageKoreanImagetourism

Korean Tourist Spot Dataset

Korean NLP Dataset · Source: GitHub · License: License unknown · 8.4GB · 10K. Catalog only; files stay with the original provider.


SourceGitHub
Size8.4GB
FormatImage
LicenseLicense unknown
Source not verifiedLicense not verified
💡
TextKoreanTextclassification

Korean Toxicity Dataset

Korean Toxicity Dataset · Source: GitHub · License: MIT · Small · 5.8K sentences. Catalog only; files stay with the original provider.


SourceGitHub
SizeSmall
FormatText
LicenseMIT
Source verifiedLicense verified
📝
TextKoreanTextloanwordspronunciation

Korean Translation Dataset

Korean Translation Dataset · Source: AI Hub · License: AI Hub terms · 55MB · 300K. Catalog only; files stay with the original provider.


SourceAI Hub
Size55MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
⚖
TextKoreanTextclassicsclassical-Chinese

Korean Translation Dataset

Korean Translation Dataset · Source: AI Hub · License: AI Hub terms · 1.8GB · 1.1M. Catalog only; files stay with the original provider.


SourceAI Hub
Size1.8GB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
🦙
TextKoreanTexteducation

Korean Translation Dataset

Korean Translation Dataset · Source: Hugging Face (cfpark00) · License: CC BY-NC-SA 4.0 · 5.5MB · 4.5K -. Catalog only; files stay with the original provider.


SourceHugging Face (cfpark00)
Size5.5MB
FormatText
LicenseCC BY-NC-SA 4.0
Source verifiedLicense verified
⚖
TextKoreanText

Korean Translation Dataset

Korean Translation Dataset · Source: GitHub · License: CC0 · Medium · 12K. Catalog only; files stay with the original provider.


SourceGitHub
SizeMedium
FormatText
LicenseCC0
Source not verifiedLicense needs review
⚖
TextKoreanText

Korean Translation Dataset

Korean Translation Dataset · Source: GitHub / AI Hub · License: Apache 2.0 · Large · 40K. Catalog only; files stay with the original provider.


SourceGitHub / AI Hub
SizeLarge
FormatText
LicenseApache 2.0
Source not verifiedLicense needs review
💡
TextKoreanTextset-phrases

Korean Translation Dataset

Korean Translation Dataset · Source: GitHub · License: MIT · Small · 6K. Catalog only; files stay with the original provider.


SourceGitHub
SizeSmall
FormatText
LicenseMIT
Source not verifiedLicense needs review
📖
TextKoreanText

Korean Translation Dataset

Korean Translation Dataset · Source: AI Hub · License: AI Hub terms · 140MB · 100K. Catalog only; files stay with the original provider.


SourceAI Hub
Size140MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
📖
TextKoreanText

Korean Translation Dataset

Korean Translation Dataset · Source: AI Hub · License: AI Hub terms · 320MB · 180K. Catalog only; files stay with the original provider.


SourceAI Hub
Size320MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
💬
TextKoreanText

Korean Translation Dataset

Korean Translation Dataset · Source: AI Hub · License: AI Hub terms · 120MB · 110K. Catalog only; files stay with the original provider.


SourceAI Hub
Size120MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
📖
TextKoreanText

Korean Translation Dataset

Korean Translation Dataset · Source: AI Hub · License: AI Hub terms · 410MB · 140K docs. Catalog only; files stay with the original provider.


SourceAI Hub
Size410MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
📝
TextKoreanText

Korean Translation Dataset

Korean Translation Dataset · Source: AI Hub · License: AI Hub terms · 620MB · 250K articles. Catalog only; files stay with the original provider.


SourceAI Hub
Size620MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
🔍
TextKoreanText

Korean Translation Dataset

Korean Translation Dataset · Source: GitHub · License: License unknown · 18MB · 65K. Catalog only; files stay with the original provider.


SourceGitHub
Size18MB
FormatText
LicenseLicense unknown
Source not verifiedLicense not verified
⚖
TextKoreanTextpatents

Korean Translation Dataset

Korean Translation Dataset · Source: AI Hub · License: AI Hub terms · 2.2GB · 400K. Catalog only; files stay with the original provider.


SourceAI Hub
Size2.2GB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
📝
TextKoreanText

Korean Translation Dataset

Korean Translation Dataset · Source: AI Hub · License: AI Hub terms · 180MB · 75K. Catalog only; files stay with the original provider.


SourceAI Hub
Size180MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
✍
TextKoreanText

Korean Translation Dataset

Korean Translation Dataset · Source: AI Hub · License: AI Hub terms · 240MB · 55K. Catalog only; files stay with the original provider.


SourceAI Hub
Size240MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
🖼
ImageKoreanImageMIT

Korean-Light-OCR

Korean NLP Dataset · Source: GitHub · License: License unknown · 4.2GB · 9,900. Catalog only; files stay with the original provider.


SourceGitHub
Size4.2GB
FormatImage
LicenseLicense unknown
Source not verifiedLicense not verified
🔍
TextKoreanTextmedicalbenchmark

KorMedMCQA QA

Korean QA Dataset · Source: Hugging Face · License: CC BY-NC 4.0 · 2.1MB · 7,489 items. Catalog only; files stay with the original provider.


SourceHugging Face
Size2.1MB
FormatText
LicenseCC BY-NC 4.0
Source verifiedLicense needs review
✦
TextKoreanTextbenchmarkcommonsense

KorNAT QA

Korean QA Dataset · Source: Hugging Face · License: CC BY-NC 4.0 · 2.6MB · 10,007 items. Catalog only; files stay with the original provider.


SourceHugging Face
Size2.6MB
FormatText
LicenseCC BY-NC 4.0
Source verifiedLicense needs review
✍
TextKoreanTextNLI

KorNLI

Korean NLP Dataset · Source: GitHub (kakaobrain) · License: CC BY-SA 4.0 · Large · 950K sentences. Catalog only; files stay with the original provider.


SourceGitHub (kakaobrain)
SizeLarge
FormatText
LicenseCC BY-SA 4.0
Source verifiedLicense verified
📖
TextKoreanTextlong-context-MRCtable-QA

KorQuAD

Korean QA Dataset · Source: KorQuAD · License: CC BY-ND 4.0 · Large ( ) · 100K. Catalog only; files stay with the original provider.


SourceKorQuAD
SizeLarge ( )
FormatText
LicenseCC BY-ND 4.0
Source not verifiedLicense needs review
📰
TextKoreanTextSTS

KorSTS

Korean NLP Dataset · Source: GitHub (kakaobrain) · License: CC BY-SA 4.0 · Small · 8.6K sentences. Catalog only; files stay with the original provider.


SourceGitHub (kakaobrain)
SizeSmall
FormatText
LicenseCC BY-SA 4.0
Source verifiedLicense verified
🔍
TextKoreanTextQA

KorWikiTableQuestions

Korean QA Dataset · Source: GitHub · License: CC BY-SA 4.0 · 70K QA. Catalog only; files stay with the original provider.


SourceGitHub
SizeUnknown
FormatText
LicenseCC BY-SA 4.0
Source verifiedLicense needs review
🔍
TextKoreanTextsafetybenchmark

KoSafetyBench

Korean LLM Benchmark Dataset · Source: Hugging Face · License: CC BY 4.0 · 15MB · 5K items. Catalog only; files stay with the original provider.


SourceHugging Face
Size15MB
FormatText
LicenseCC BY 4.0
Source verifiedLicense needs review
📝
TextKoreanTextsafety

KoSBi

Korean NLP Dataset · Source: GitHub · License: MIT · Large · 68K sentences. Catalog only; files stay with the original provider.


SourceGitHub
SizeLarge
FormatText
LicenseMIT
Source verifiedLicense verified
💬
TextKoreanTextbenchmarkQA

KoSimpleQA QA

Korean QA Dataset · Source: Hugging Face · License: License unknown · 301KB · 1,000 items. Catalog only; files stay with the original provider.


SourceHugging Face
Size301KB
FormatText
LicenseLicense unknown
Source verifiedLicense not verified
🎵
SpeechKoreanAudio

KoSum

Korean NLP Dataset · Source: Hugging Face (iontail) · License: CC BY-NC 4.0 · 45GB · 700. Catalog only; files stay with the original provider.


SourceHugging Face (iontail)
Size45GB
FormatAudio
LicenseCC BY-NC 4.0
Source verifiedLicense verified
📰
TextKoreanTextemotionaffect

KOTE

Korean NLP Dataset · Source: Hugging Face (searle-j) · License: MIT · Medium · 50K comments. Catalog only; files stay with the original provider.


SourceHugging Face (searle-j)
SizeMedium
FormatText
LicenseMIT
Source verifiedLicense verified
🔍
TextKoreanText

KPoEM

Korean NLP Dataset · Source: GitHub · License: MIT · 4.2MB · 6.5K. Catalog only; files stay with the original provider.


SourceGitHub
Size4.2MB
FormatText
LicenseMIT
Source not verifiedLicense needs review
🧾
ImageKoreanImageRichTextOCR-VQA

KRETA (Korean Rich Text Visual QA)

Korean QA Dataset · Source: Hugging Face · License: License unknown · 125MB · 15 2.58K. Catalog only; files stay with the original provider.


SourceHugging Face
Size125MB
FormatImage
LicenseLicense unknown
Source verifiedLicense not verified
🔍
TextKoreanTextLLM

KULLM-v2

Korean NLP Dataset · Source: Hugging Face · License: Apache 2.0 · 180MB · 152,630. Catalog only; files stay with the original provider.


SourceHugging Face
Size180MB
FormatText
LicenseApache 2.0
Source verifiedLicense verified
📰
TextKoreanTextLIMASFT

LIMA

Korean Instruction Tuning Dataset · Source: Hugging Face · License: CC BY-NC-SA 4.0 · Small · 1K. Catalog only; files stay with the original provider.


SourceHugging Face
SizeSmall
FormatText
LicenseCC BY-NC-SA 4.0
Source verifiedLicense verified
🎙
SpeechKoreanAudioTTSTTS

MeloTTS

Korean NLP Dataset · Source: Hugging Face · License: MIT · Medium · 30K. Catalog only; files stay with the original provider.


SourceHugging Face
SizeMedium
FormatAudio
LicenseMIT
Source verifiedLicense verified
📰
TextKoreanText

Meme

Korean NLP Dataset · Source: GitHub · License: MIT · Small · 5K. Catalog only; files stay with the original provider.


SourceGitHub
SizeSmall
FormatText
LicenseMIT
Source not verifiedLicense needs review
🦙
TextKoreanTextpersona

Nemotron

Korean NLP Dataset · Source: Hugging Face (nvidia) · License: CC BY 4.0 · 120MB · 9.4K. Catalog only; files stay with the original provider.


SourceHugging Face (nvidia)
Size120MB
FormatText
LicenseCC BY 4.0
Source verifiedLicense verified
🔍
TextKoreanTextNER

NLP

Korean NER Dataset · Source: GitHub (Naver-NLP) · License: CC0 · Medium · 80K. Catalog only; files stay with the original provider.


SourceGitHub (Naver-NLP)
SizeMedium
FormatText
LicenseCC0
Source not verifiedLicense needs review
📦
ImageKoreanImageOCR

OCR

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 1.5TB · 450K. Catalog only; files stay with the original provider.


SourceAI Hub
Size1.5TB
FormatImage
LicenseAI Hub terms
Source not verifiedLicense verified
🧮
ImageKoreanImage

OCR

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 290GB · 150K. Catalog only; files stay with the original provider.


SourceAI Hub
Size290GB
FormatImage
LicenseAI Hub terms
Source not verifiedLicense verified
🔊
SpeechKoreanAudioASR

OLKAVS

Korean NLP Dataset · Source: GitHub · License: License unknown · 1,150hours · 1,107 speakers. Catalog only; files stay with the original provider.


SourceGitHub
Size1,150hours
FormatAudio
LicenseLicense unknown
Source verifiedLicense not verified
🔍
TextKoreanTextLLM

OpenAssistant Guanaco

Korean NLP Dataset · Source: Hugging Face · License: Apache 2.0 · 16.7MB · 10,364 dialogues. Catalog only; files stay with the original provider.


SourceHugging Face
Size16.7MB
FormatText
LicenseApache 2.0
Source verifiedLicense needs review
📝
TextKoreanTextLLMreasoning

OpenOrca-KO

Korean NLP Dataset · Source: Hugging Face · License: MIT · 21.8MB · 21,632. Catalog only; files stay with the original provider.


SourceHugging Face
Size21.8MB
FormatText
LicenseMIT
Source verifiedLicense needs review
📖
TextKoreanTextpreferenceDPO

Orca DPO

Korean NLP Dataset · Source: Hugging Face · License: Apache 2.0 · Medium · 29K. Catalog only; files stay with the original provider.


SourceHugging Face
SizeMedium
FormatText
LicenseApache 2.0
Source verifiedLicense verified
📣
SpeechKoreanAudioASRASR

Pansori TEDxKR

Korean NLP Dataset · Source: GitHub · License: CC BY-NC-ND 4.0 · ~3hours · 41 speakers. Catalog only; files stay with the original provider.


SourceGitHub
Size~3hours
FormatAudio
LicenseCC BY-NC-ND 4.0
Source verifiedLicense needs review
✍
TextKoreanText

ParaKQC

Korean NLP Dataset · Source: GitHub · License: CC BY-SA 4.0 · 10,000 (v1). Catalog only; files stay with the original provider.


SourceGitHub
SizeUnknown
FormatText
LicenseCC BY-SA 4.0
Source verifiedLicense needs review
📰
TextKoreanTextcommonsenseSFT

QARV

Korean Instruction Tuning Dataset · Source: Hugging Face · License: MIT · Medium · 33K. Catalog only; files stay with the original provider.


SourceHugging Face
SizeMedium
FormatText
LicenseMIT
Source verifiedLicense verified
📰
TextKoreanTextMRCinstruction

QG

Korean Instruction Tuning Dataset · Source: Hugging Face · License: Apache 2.0 · Small · 15K. Catalog only; files stay with the original provider.


SourceHugging Face
SizeSmall
FormatText
LicenseApache 2.0
Source not verifiedLicense needs review
🔍
TextKoreanTextpersonaSFT

Role-Playing

Korean Instruction Tuning Dataset · Source: Hugging Face (huggingface-KREW) · License: Apache 2.0 · 28MB · 15K dialogues sessions. Catalog only; files stay with the original provider.


SourceHugging Face (huggingface-KREW)
Size28MB
FormatText
LicenseApache 2.0
Source verifiedLicense verified
📸
ImageKoreanImageVQA

skt KVQA

Korean QA Dataset · Source: Hugging Face · License: Korean VQA License · 100,445 · -Q&A. Catalog only; files stay with the original provider.


SourceHugging Face
Size100,445
FormatImage
LicenseKorean VQA License
Source verifiedLicense verified
🦙
TextKoreanText

Smilegate UnSmile

Korean NLP Dataset · Source: GitHub · License: CC BY-NC-ND 4.0 · Small · 18K sentences. Catalog only; files stay with the original provider.


SourceGitHub
SizeSmall
FormatText
LicenseCC BY-NC-ND 4.0
Source verifiedLicense verified
🔍
TextKoreanText

SmileStyle

Korean NLP Dataset · Source: GitHub (smilegate-ai) · License: MIT · 35MB · 21K. Catalog only; files stay with the original provider.


SourceGitHub (smilegate-ai)
Size35MB
FormatText
LicenseMIT
Source verifiedLicense verified
🔍
TextKoreanText

SMS

Korean NLP Dataset · Source: GitHub · License: MIT · Small · 12K. Catalog only; files stay with the original provider.


SourceGitHub
SizeSmall
FormatText
LicenseMIT
Source not verifiedLicense needs review
✍
TextKoreanTextSNSdeidentified

SNS

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 980MB · 1.2M turns. Catalog only; files stay with the original provider.


SourceAI Hub
Size980MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
🦙
TextKoreanTextlogiceval

SNU Ko-MuSR

Korean NLP Dataset · Source: Hugging Face (thunder-research-group) · License: MIT · Small · images. Catalog only; files stay with the original provider.


SourceHugging Face (thunder-research-group)
SizeSmall
FormatText
LicenseMIT
Source verifiedLicense verified
🎧
SpeechKoreanAudioMIT

Speech Emotion Dataset (ko)

Korean NLP Dataset · Source: GitHub · License: License unknown · 3.2GB. Catalog only; files stay with the original provider.


SourceGitHub
Size3.2GB
FormatAudio
LicenseLicense unknown
Source not verifiedLicense not verified
✦
TextKoreanTextsafety

SQuARe

Korean NLP Dataset · Source: GitHub · License: MIT · Medium · 40K sentences. Catalog only; files stay with the original provider.


SourceGitHub
SizeMedium
FormatText
LicenseMIT
Source verifiedLicense verified
📰
TextKoreanTextautomotivemanual

SUV QA

Korean QA Dataset · Source: AI Hub · License: AI Hub terms · 230MB · 40K. Catalog only; files stay with the original provider.


SourceAI Hub
Size230MB
FormatText
LicenseAI Hub terms
Source not verifiedLicense verified
📸
ImageKoreanImageVQA

Table-VQA-ko

Korean QA Dataset · Source: Hugging Face · License: Apache 2.0 · 185MB · LLaVA. Catalog only; files stay with the original provider.


SourceHugging Face
Size185MB
FormatImage
LicenseApache 2.0
Source verifiedLicense verified
⚖
TextKoreanText

Tatoeba (ko-en)

Korean NLP Dataset · Source: GitHub · License: CC BY 2.0 · 3.4MB. Catalog only; files stay with the original provider.


SourceGitHub
Size3.4MB
FormatText
LicenseCC BY 2.0
Source verifiedLicense verified
💡
TextKoreanTextreasoning

Theory of Mind

Korean NLP Dataset · Source: GitHub · License: Apache 2.0 · 8.5MB · 8K. Catalog only; files stay with the original provider.


SourceGitHub
Size8.5MB
FormatText
LicenseApache 2.0
Source not verifiedLicense needs review
🦙
TextKoreanTextbenchmarkACL2026

Thunder-KoNUBench

Korean LLM Benchmark Dataset · Source: Hugging Face (thunder-research-group) · License: CC BY-NC-SA 4.0 · 5.4MB · items. Catalog only; files stay with the original provider.


SourceHugging Face (thunder-research-group)
Size5.4MB
FormatText
LicenseCC BY-NC-SA 4.0
Source verifiedLicense verified
💬
TextKoreanTextpretrainingtranslation

TinyStories

Korean Translation Dataset · Source: Hugging Face · License: MIT · Large · 175K docs. Catalog only; files stay with the original provider.


SourceHugging Face
SizeLarge
FormatText
LicenseMIT
Source verifiedLicense verified
🎙
SpeechKoreanAudioTTSmulti-speaker

TTS

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 300hours · 210K speakers. Catalog only; files stay with the original provider.


SourceAI Hub
Size300hours
FormatAudio
LicenseAI Hub terms
Source not verifiedLicense verified
🔊
SpeechKoreanAudioTTS

TTS

Korean NLP Dataset · Source: AI Hub · License: AI Hub terms · 90hours · 35K dialogues. Catalog only; files stay with the original provider.


SourceAI Hub
Size90hours
FormatAudio
LicenseAI Hub terms
Source not verifiedLicense verified
✍
TextKoreanTextpretrainingcorpus

WanJuanSiLu

Korean Pretraining Corpus · Source: Hugging Face (opendatalab) · License: CC BY 4.0 · 124 GB. Catalog only; files stay with the original provider.


SourceHugging Face (opendatalab)
Size124 GB
FormatText
LicenseCC BY 4.0
Source verifiedLicense verified
🔍
TextKoreanTextWiCKoBEST

WiC

Korean NLP Dataset · Source: Hugging Face (SKT-brain) · License: CC BY-SA 4.0 · Small · 3K items. Catalog only; files stay with the original provider.


SourceHugging Face (SKT-brain)
SizeSmall
FormatText
LicenseCC BY-SA 4.0
Source verifiedLicense verified
📰
TextKoreanTextInstruction

xP3x

Korean Instruction Tuning Dataset · Source: Hugging Face · License: Apache 2.0 · 4,642,468rows. Catalog only; files stay with the original provider.


SourceHugging Face
SizeUnknown
FormatText
LicenseApache 2.0
Source verifiedLicense needs review

External Data Resources

External Korean NLP corpus lists

Beyond the GearDel catalog, these public lists on GitHub and blogs are useful references.

Suggested
How to start

Korean AI datasets:
task, license, original page.

Open a speech, OCR, or benchmark collection first. Then read the license mark and leave for the original host.

01

Korean speech / OCR / benchmarks

ASR is not TTS. Receipt OCR is not scene text. KLUE-family rows are for scores, not pretraining.

02

Korean dataset licenses

Commercial-allowed on GearDel is a catalog hint. AI Hub terms are not MIT.

03

Original download

Check provider, source URL, and size text, then download on the host. GearDel does not keep files.

Full guide: how to choose a dataset →·License comparison →·FAQ →·About GearDel →